LLMOps Engineer — LLM Inference & Model Optimization
- Posted
- Employment
- unknown
- Work mode
- unknown
- Experience
- 3-6 years; 2+ production LLM inference
Skills explicitly mentioned: Python, PyTorch, CUDA, Kubernetes, Docker, vLLM, TensorRT-LLM
Gnani.ai is hiring an LLMOps Engineer in Bengaluru to operate and optimize its in-house language models for real-time voice applications. The role covers GPU inference services, model compression, deployment reliability and performance measurement. Applicants need three to six years in ML infrastructure, model serving or performance engineering, including at least two years of production LLM inference. Relevant skills include Python, PyTorch, CUDA fundamentals, Kubernetes and practical experience with inference engines such as vLLM, SGLang or TensorRT-LLM. Responsibilities include benchmarking, quantization, incident response and reducing inference costs while protecting model quality.