Senior GPU Systems & Fabric Engineer
bitdeer
Job description
About the role
Bitdeer is seeking a Senior GPU Systems & Fabric Engineer to lead the design, integration and optimisation of GPU‑accelerated infrastructure for large‑scale AI training and inference. You will work across hardware, networking and container orchestration layers to ensure high‑performance, reliable compute fabric in our global datacenters.
Key responsibilities
- Architect and maintain NVIDIA/AMD GPU device plugins and Kubernetes Operators to expose hardware capabilities to the control plane.
- Configure and optimise high‑performance networking stacks (RDMA, SR‑IOV, RoCEv2, InfiniBand) for line‑rate throughput in distributed AI workloads.
- Build automated remediation pipelines using DCGM telemetry to detect and reset degraded GPU or NIC components before they affect production jobs.
- Implement and manage GPU slicing technologies such as MIG and vGPU to support multi‑tenant inference workloads and maximise cluster utilisation.
- Profile and tune kernel parameters, device drivers and runtime libraries (CUDA, NCCL) to eliminate bottlenecks in containerised AI workloads.
- Collaborate with scheduling and storage teams to ensure topology‑aware placement and efficient data movement across the fabric.
- Define and enforce standards for bare‑metal provisioning, BIOS/firmware updates and OS hardening in containerised environments.
- Lead investigations of complex performance issues spanning hardware, fabric and software, and mentor junior engineers.
Required profile
- Bachelor’s or Master’s degree in Computer Science, Electrical Engineering or a related field.
- 5+ years of systems engineering experience with deep knowledge of Linux kernel internals.
- Hands‑on experience with modern GPU architectures (e.g., NVIDIA H100/A100) and CUDA runtimes.
- Proven track record operating, debugging and scaling bare‑metal or HPC systems in production.
- Strong problem‑solving skills and ability to work in ambiguous, high‑performance environments.
Required skills
- NVIDIA/AMD GPU device plugins
- Kubernetes Operators and device plugin architecture
- RDMA, SR‑IOV, RoCEv2, InfiniBand networking
- DCGM telemetry and GPU health monitoring
- MIG, vGPU slicing technologies
- CUDA, NCCL runtime libraries
- Linux kernel development, C, Go
- Terraform, Ansible, CI/CD pipelines
- Bare‑metal provisioning, BIOS/firmware management
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Singapore.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 month ago
Expires 3 weeks from now
31 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
bitdeer