Sr LLMOps Engineer — LLM Inference & Model Optimization
- Posted
- Employment
- unknown
- Work mode
- unknown
- Experience
- 6–10 years
Skills explicitly mentioned: LLM inference, CUDA, Triton, Kubernetes, vLLM, Model optimization
Gnani.ai is hiring a senior engineer in Bengaluru to lead production inference for its own language and speech models. The role combines GPU serving architecture, model efficiency, custom kernel work and operational reliability. Candidates need 6–10 years of experience, including at least three years owning large-scale LLM or speech inference in production. Relevant strengths include CUDA or Triton, distributed profiling, mixture-of-experts serving and evaluation-led optimization. Responsibilities also cover capacity planning, safe releases, incident response and mentoring engineers. The employer does not specify a salary, degree requirement or work arrangement in this listing.