hirly

algoleap

MLOps / LLM Infra Engineers

Mumbai, Maharashtra, India

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at algoleap first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Seniority
Mid level
Country
IN
Work mode
On-site / unstated
First seen by hirly
26 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Job Description – MLOps / LLM Infrastructure Engineer

Experience: 7–12 years

Role Overview

Responsible for hosting, deploying, and operating open-weight LLMs within a sovereign cloud environment. The role focuses on GPU infrastructure, model serving, performance optimization, and reliable model lifecycle management.

Key Responsibilities

  • Deploy and operate models such as GPT, LLaMA, Gemma, Mistral, and other product/open-weight models within sovereign cloud.
  • Manage GPU provisioning, capacity planning, utilization, and performance optimization.
  • Implement LLM inference serving using platforms such as vLLM, Triton, or similar frameworks.
  • Build model deployment, versioning, rollback, and lifecycle management processes.
  • Develop MLOps pipelines for model packaging, testing, deployment, and monitoring.
  • Monitor latency, throughput, GPU utilization, availability, and inference costs.
  • Implement scalable and highly available model-serving infrastructure using Kubernetes and containers.
  • Work closely with platform, security, and gateway teams to ensure secure model access and governance.
  • Troubleshoot production issues across GPU, inference, Kubernetes, networking, and model-serving layers.

Required Skills

  • Strong experience in MLOps / LLM infrastructure / model serving.
  • Hands-on experience with GPU-based inference and Kubernetes.
  • Experience with vLLM, NVIDIA Triton, TensorRT-LLM, or equivalent.
  • Experience deploying and managing LLaMA, Gemma, Mistral, GPT or similar LLMs.
  • Knowledge of Docker, Kubernetes, CI/CD, model registries, and observability.
  • Understanding of LLM quantization, batching, caching, GPU memory management, and inference optimization.
  • Experience operating ML/LLM workloads in private, on-premises, or sovereign-cloud environments.

Preferred: Experience with NVIDIA GPUs, CUDA, Helm, Prometheus/Grafana, MLflow, and automated model deployment pipelines .

Original posting on algoleap's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job