Jobiglo

No results.

AI Infrastructure Engineer

THE SUPREME HR ADVISORY PTE. LTD. · Singapore

Senior 5,000 - 7,000 SGD/month 🇬🇧 English
Linux systems administration Kernel tuning Bash Python GPU hardware architectures CUDA cuDNN PCIe NVLink Kubernetes GPU operator Device plugins Slurm Run:ai Ray RDMA RoCE v2 InfiniBand Priority Flow Control ECN Lustre GPFS IBM Spectrum Scale Ceph MinIO Terraform Ansible Helm Pulumi Prometheus Grafana DCGM exporter NVIDIA System Management Interface

Job description

About the role

The AI Infrastructure Engineer will design, build and maintain high‑density GPU compute clusters that power AI/ML workloads across the organisation. You will work closely with AI/ML teams to ensure optimal performance, reliability and scalability of the underlying infrastructure.

Key responsibilities

  • Architect, configure and operate multi‑GPU compute clusters (e.g., NVIDIA HGX/DGX) and associated container orchestration platforms such as Kubernetes, Slurm or Ray.
  • Monitor GPU health, telemetry and utilization, and implement strategies to minimise idle time and bottlenecks.
  • Design and optimise low‑latency network fabrics (InfiniBand, RoCE v2, NVLink) and high‑throughput storage solutions (Lustre, GPFS, Ceph, MinIO, NVMe‑oF) for distributed training.
  • Develop automated deployment pipelines using Infrastructure as Code tools (Terraform, Ansible, Helm, Pulumi) and maintain golden images and firmware updates.
  • Implement end‑to‑end observability with Prometheus, Grafana, DCGM exporter and NVIDIA System Management Interface, and troubleshoot performance issues with AI/ML teams.

Required profile

  • Bachelor’s degree in Computer Science, IT, Computer Engineering or equivalent practical experience.
  • 3–6+ years of hands‑on experience in infrastructure engineering, high‑performance computing, DevOps or cloud infrastructure.
  • Relevant certifications (e.g., CKA/CKAD, NVIDIA Certified Associate/Professional, AWS/Azure/GCP Solutions Architect) are a plus.

Required skills

  • Linux systems administration, kernel tuning, Bash and Python scripting.
  • GPU hardware architectures, CUDA, cuDNN, PCIe/NVLink.
  • Kubernetes (GPU operator, device plugins), Slurm, Run:ai, Ray.
  • RDMA, RoCE v2, InfiniBand, Priority Flow Control (PFC), ECN.
  • High‑IOPS storage systems: Lustre, GPFS/IBM Spectrum Scale, Ceph, MinIO, NVMe‑oF.
  • Infrastructure as Code: Terraform, Ansible, Helm, Pulumi.
  • Monitoring and observability: Prometheus, Grafana, DCGM exporter, NVIDIA SMI.

Questions fréquentes

Le salaire proposé pour ce poste est de 5-7k SGD par mois. Le détail figure dans l'annonce.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 1 month ago

Expires 3 weeks from now

36 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

THE SUPREME HR ADVISORY PTE. LTD.

Singapore