Lead MLOps/SRE Engineer – Developer Experience (LLM)
HTX (Home Team Science & Technology Agency) · Singapour
Job description
About the role
HTX is looking for a Lead Engineer to own the developer experience for its AI infrastructure. You will be responsible for deploying, operating, and continuously improving a production large‑language‑model (LLM) system within a secure, high‑performance environment.
Key responsibilities
- Deploy and manage LLM models using vLLM, TensorRT‑LLM or similar on GPU clusters, optimizing throughput, latency and GPU utilization.
- Provision and maintain supporting services such as vector databases for retrieval‑augmented generation, orchestration platforms, Redis/queue systems and API gateways.
- Profile and tune inference performance, experimenting with batching, context caching and quantisation to maximise efficiency.
- Implement observability using Prometheus, Grafana, DCGM exporters and the Elastic Stack to monitor latency, token throughput, cache hit rates and overall system health.
- Define SLOs/SLIs, set up auto‑scaling policies, design failure‑recovery mechanisms and conduct chaos‑engineering exercises to ensure high availability.
- Monitor GPU utilisation and inference costs, identify optimisation opportunities and implement measures to reduce compute spend.
- Ensure security and compliance by managing secrets, maintaining audit logs and operating within secure network boundaries.
- Participate in on‑call rotation, troubleshoot incidents, perform root‑cause analysis and drive preventive actions.
Required profile
- 4+ years of experience in MLOps, SRE or DevOps roles, with at least 1 year focused on ML/AI systems.
- Hands‑on experience deploying and operating LLMs in production environments.
- Strong understanding of GPU architecture and inference optimisation techniques.
- Proven ability to design, implement and maintain large‑scale, reliable infrastructure.
Required skills
- Python
- Bash scripting
- Kubernetes (operators, StatefulSets, GPU scheduling)
- vLLM, TensorRT‑LLM, TGI or equivalent LLM serving frameworks
- GPU inference optimisation (batching, quantisation, context caching)
- Prometheus, Grafana, Elastic Stack (ELK)
- Vector databases (e.g., Milvus, Pinecone)
- Redis, queue systems, API gateways
- SLO/SLI definition, auto‑scaling, chaos engineering
Questions fréquentes
Why are you reporting this job?
Explore further
Salaries, guides and searches in Singapore.
Salaries by job title
Apply in 30 seconds
Enter your email to apply. An account will be created automatically.
By continuing, you accept our terms of use.
Already have an account? Login
Published 1 month ago
Expires 1 day from now
32 views · 0 interested
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
HTX (Home Team Science & Technology Agency)
Singapour