Senior AI Research Engineer
- Posted
- Employment
- full time
- Work mode
- unknown
Skills explicitly mentioned: LLM, Machine Learning, Hugging Face, PEFT, Model Evaluation, Fine-Tuning, Python, Quantization, vLLM
Official role: Senior AI Research Engineer Trianz requisition: TR25064 Location: Bangalore, India Experience: 14-18 years Posted: 07-Jul-2026
Role Overview
ABOUT THIS ROLE
You produce the evidence base that justifies every AI architectural decision in the company. When the team decides to deploy a 7B model on CPU rather than a 70B model via API, that decision rests on your benchmark data. You run systematic evaluations across parameter sizes and model families on domain-specific enterprise tasks -- not just general benchmarks. You compare open-source models against closed models and produce rigorous, reproducible results that drive real architectural bets.
WHAT YOU WILL DO
Design and implement the LLM evaluation framework: benchmark suite, evaluation harness, reproducibility standards- Run systematic evaluations across 5B, 10B, and 30B parameter models on enterprise-specific task types- Benchmark open-source models (Llama, Mistral, Qwen, Phi) vs closed models (GPT-4, Claude, Gemini) on accuracy, latency, and cost- Evaluate model performance on CPU-only inference vs GPU inference to validate hardware routing decisions- Produce model selection recommendations with supporting data for every AI architectural decision- Design and run fine-tuning experiments for smaller models on domain-specific tasks- Maintain the model leaderboard: continuously updated benchmarks as new models are released- Evaluate quantization strategies (GPTQ, AWQ, GGUF) and their accuracy vs performance trade-offs
MUST HAVE
Designed and run LLM evaluation frameworks in a production or research context -- not just existing benchmarks
* Hands-on with HuggingFace Transformers, PEFT, and evaluation libraries (lm-eval-harness, HELM, or equivalent)
* Evaluated models on domain-specific tasks, not just general benchmarks
* Fine-tuned open-source LLMs for specific downstream tasks
* Strong Python -- can build a complete evaluation pipeline from scratch
* 5+ years in ML/AI with at least 2 years focused on LLM evaluation or fine-tuning
GOOD TO HAVE
Inference optimisation experience: vLLM, TensorRT-LLM, or Triton- CPU inference frameworks: llama.cpp, Intel OpenVINO, or AMD ROCm- Published research or technical writing on LLM evaluation methodology
Company Overview
Trianz is an applied AI solutions company that accelerates customer business transformation through AI powered "Transformation Services as a Software Model" . With 25+ years of transforming enterprises, we've evolved to a product-led, platform-driven organization serving global enterprises across Financial Services, Insurance, Healthcare, Hi-Tech, Manufacturing, and other industries.
With global presence across 4 continents, our platform portfolio under the unified Concierto brand delivers end-to-end transformations including solutions for Migrate, Manage, Maximize, Modernize, Insights & Agentic AI, and SecOps - delivered through strategic partnerships with leading hyperscalers.
We're building the premier innovation-led organization in the digital transformation space through AI-first methodologies and data-driven excellence - RevolutionAIzing Transformations.