Jobiglo

No results.

GPU Compute & Bare Metal / DPU Engineer

bitdeer

Mid 🇬🇧 English
NVIDIA HGX CUDA Linux PXE IPMI Redfish BMC BIOS NIC DPU SmartNIC NVIDIA BlueField Ansible Terraform Python Go

Job description

About the role

Bitdeer is scaling its AI‑cloud platform to a fleet of 10,000+ GPUs across multiple regions. As a GPU Compute & Bare Metal / DPU Engineer you will own the end‑to‑end lifecycle of bare‑metal GPU nodes, ensuring they are provisioned, operated, and decommissioned reliably and efficiently.

Key responsibilities

  • Manage the full lifecycle of bare‑metal GPU nodes – provisioning, onboarding, in‑service operation, break‑fix, and decommissioning across several data‑center regions.
  • Build and maintain automated, repeatable node‑delivery pipelines to eliminate backlogs and support fleet growth.
  • Control DPU/SmartNIC and server firmware (BMC, BIOS, NIC, GPU) version baselines, upgrades, and validation.
  • Drive fleet reliability by reducing MTTR, leading incident response, performing root‑cause analysis, and improving hardware‑health monitoring.
  • Participate in a 24/7 multi‑region on‑call rotation, creating runbooks and tooling to reduce manual toil.
  • Collaborate with Storage/Image and Network teams to streamline provisioning‑to‑handoff processes.
  • Define bring‑up, rack, capacity, and acceptance standards for new GPU SKUs and new data‑center regions.

Required profile

  • 3+ years (senior candidates 6+ years) in large‑scale bare‑metal/server‑fleet operations, HPC, or cloud infrastructure.
  • Hands‑on experience with GPU servers at scale (e.g., NVIDIA HGX/DGX), including driver, CUDA, and firmware management.
  • Strong Linux systems expertise, including PXE, IPMI, Redfish, OS imaging, and automated provisioning.
  • Familiarity with DPU/SmartNIC technologies such as NVIDIA BlueField.
  • Experience with infrastructure automation tools (Ansible, Terraform) and scripting/programming languages (Python, Go).
  • Comfortable owning on‑call duties, incident management, and operational runbooks.
  • Multi‑region or large‑fleet operational experience is a strong plus.

Required skills

  • NVIDIA HGX/DGX GPU servers
  • CUDA driver management
  • Linux (PXE, IPMI, Redfish)
  • Server firmware (BMC, BIOS, NIC)
  • DPU/SmartNIC (NVIDIA BlueField)
  • Ansible
  • Terraform
  • Python
  • Go

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec bitdeer.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:breezy

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in Singapore.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 4 weeks ago

Expires 1 month from now

28 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

bitdeer