📢 New: get today's jobs on our WhatsApp Channel
Jobiglo

No results.

Senior Inference Runtime Engineer

bitdeer

Senior 🇬🇧 English
vLLM Dynamo SGLang TensorRT-LLM TGI Triton CUDA NCCL Batching Streaming Distributed inference Go Python

Job description

About the role

Bitdeer is looking for a Senior Inference Runtime Engineer to lead the performance‑critical serving layer of its MaaS platform. You will make self‑hosted large language models faster, cheaper, and more stable by optimizing the runtime stack behind OpenAI‑ and Anthropic‑compatible APIs.

Key responsibilities

  • Optimize prefill/decode scheduling, continuous batching, KV‑cache behavior, speculative decoding, long‑context serving, and streaming smoothness.
  • Tune and operate vLLM/Dynamo/SGLang/TensorRT‑LLM‑style runtimes for latency, throughput, GPU utilization, and cost efficiency.
  • Profile bottlenecks across GPU memory, HBM bandwidth, NCCL/network, tokenizer, frontend/proxy, and model worker paths.
  • Lead high‑value model onboarding, selecting runtimes, parallelism strategies, quantization, context length, and rollback plans.
  • Define runtime playbooks and safe defaults for reasoning, tool calling, multimodal, prompt cache, and provider‑specific parameters.
  • Collaborate with SRE and performance engineers to turn benchmark findings into production improvements.

Required profile

  • 6+ years of systems, ML infrastructure, or high‑performance backend engineering experience.
  • Hands‑on experience with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT‑LLM, TGI, or Triton.
  • Strong understanding of GPU memory, CUDA/NCCL basics, KV‑cache, batching, streaming, and distributed inference trade‑offs.
  • Proficient in Go or Python and comfortable reading runtime source code, profiling traces, and production metrics.
  • Experience operating production inference services with strict latency, availability, and cost targets.

Required skills

  • vLLM
  • Dynamo
  • SGLang
  • TensorRT‑LLM
  • TGI
  • Triton
  • GPU memory management
  • CUDA
  • NCCL
  • KV‑cache handling
  • Batching and streaming techniques
  • Distributed inference
  • Go
  • Python
  • Performance profiling
  • Production metrics analysis

What we offer

  • A culture that values authenticity and diverse perspectives.
  • An inclusive environment with open workspaces and a startup spirit.
  • Fast‑growing company providing networking opportunities with industry pioneers.
  • Direct impact on cutting‑edge AI cloud and Bitcoin mining technologies.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec bitdeer.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.
Source : ats:breezy

Why are you reporting this job?

Thank you for your report. We will review this job.

Explore further

Salaries, guides and searches in Singapore.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 3 weeks ago

Expires 1 month from now

29 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

bitdeer