algoleap
MLOps / LLM Infra Engineers
Mumbai, Maharashtra, India
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Mid level
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 26 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Job Description – MLOps / LLM Infrastructure Engineer
Experience: 7–12 years
Role Overview
Responsible for hosting, deploying, and operating open-weight LLMs within a sovereign cloud environment. The role focuses on GPU infrastructure, model serving, performance optimization, and reliable model lifecycle management.
Key Responsibilities
- Deploy and operate models such as GPT, LLaMA, Gemma, Mistral, and other product/open-weight models within sovereign cloud.
- Manage GPU provisioning, capacity planning, utilization, and performance optimization.
- Implement LLM inference serving using platforms such as vLLM, Triton, or similar frameworks.
- Build model deployment, versioning, rollback, and lifecycle management processes.
- Develop MLOps pipelines for model packaging, testing, deployment, and monitoring.
- Monitor latency, throughput, GPU utilization, availability, and inference costs.
- Implement scalable and highly available model-serving infrastructure using Kubernetes and containers.
- Work closely with platform, security, and gateway teams to ensure secure model access and governance.
- Troubleshoot production issues across GPU, inference, Kubernetes, networking, and model-serving layers.
Required Skills
- Strong experience in MLOps / LLM infrastructure / model serving.
- Hands-on experience with GPU-based inference and Kubernetes.
- Experience with vLLM, NVIDIA Triton, TensorRT-LLM, or equivalent.
- Experience deploying and managing LLaMA, Gemma, Mistral, GPT or similar LLMs.
- Knowledge of Docker, Kubernetes, CI/CD, model registries, and observability.
- Understanding of LLM quantization, batching, caching, GPU memory management, and inference optimization.
- Experience operating ML/LLM workloads in private, on-premises, or sovereign-cloud environments.
Preferred: Experience with NVIDIA GPUs, CUDA, Helm, Prometheus/Grafana, MLflow, and automated model deployment pipelines .
Similar jobs
- AI / Machine Learning Engineer – MLOpsUmanist Staffing LLC · IndiaFirst seen today
- DevOpsMLOpsPythonMLInfosys Limited · Bangalore, Karnataka, IndiaFirst seen today
- DevOpsMLOpsPythonML DeveloperInfosys Limited · Bangalore, Karnataka, IndiaFirst seen today
- MLOps EngineerInfosys · Bangalore, IndiaFirst seen today
- DevOps & MLOps Engineer – Automation, Analytics & GenAIHpe · Bengaluru, Karnātaka, IndiaFirst seen 2d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job