hirly

Umanist Staffing LLC

Senior Site Reliability Engineer (SRE) Engineer

India

Apply through hirly

Upload your resume and get a version tailored to this job, plus a cover letter, in about thirty seconds — before you create an account.

Apply with hirly

hirly's read of this role

Role family
Engineering
Seniority
Senior
Country
IN
Work mode
On-site / unstated
First seen by hirly
26 Sept 2026

Derived automatically from the posting. Sign up to see how the role scores against your own resume.

the posting

  • Senior Site Reliability Engineer (SRE) Engineer
  • ONLY PUNE, NAGPUR,KOLHAPUR PROFILES WILL BE CONSIDERED FOR INTERVIEW, A BIG "NO" FOR ANY OTHER LOCATIONS EVEN FOR MUMBAI.

"Microsoft Azure/AWS/GCP, Kubernetes, Terraform, Datadog, OpenTelemetry, Golden Signals monitoring( Latency,Traffic, Errors, Saturation) and modern SRE practices( SLIs, SLOs, SLAs, and Error Budgets.)" Please match your work experience with these mentioned skillset for a quick right-fit check.

  • Location: Viman Nagar, Pune – Work From Office
  • Experience Overall(must have): 8 Years
  • CTC: Up to ₹25 LPA
  • Notice Period: Immediate Joiners Only within 15d or (if serving max 30days)
  • Working Hours: 3:00 PM – 12:00 AM, Monday to Friday
  • On-Call: 24/7 Production Support – On-Call Rotation Required
  • Employment Type: Full-Time

About the Role

We are looking for an experienced Senior Site Reliability Engineer (SRE) / DevOps Engineer to manage and improve the reliability, scalability, performance, security, and observability of mission-critical production environments.

The role requires strong hands-on expertise in Cloud, Kubernetes, DevOps automation, Monitoring & Observability, Incident Management, and SRE practices . The ideal candidate should be comfortable handling production incidents while also driving long-term initiatives around reliability, automation, scalability, and reduction of operational toil.

Must-Have Skills & Experience1. SRE & Production Operations

Relevant 7+ years of relevant experience in SRE / DevOps / Cloud Infrastructure / Production Engineering .

Hands-on experience with 24/7 production support and on-call operations .

Strong experience in incident management, troubleshooting, RCA, and post-mortems .

Good understanding of SLI, SLO, SLA, Error Budgets, MTTR, and reliability engineering .

Experience with toil reduction, capacity planning, high availability, disaster recovery, and failover strategies .

Ability to improve system availability, performance, scalability, and operational reliability.

2. Cloud & Infrastructure

Strong hands-on experience with Microsoft Azure, AWS, and/or GCP .

Strong understanding of cloud infrastructure, networking, IAM, storage, compute, and cloud-native services.

Hands-on experience with at least one major cloud platform and good exposure to multi-cloud environments.

Experience with:

Azure: VMs, Networking, Storage, IAM, Azure Monitor, AKS

AWS: EC2, S3, RDS, IAM, VPC, CloudWatch, EKS

GCP: Compute Engine, Cloud Storage, IAM, VPC, GKE, Cloud Monitoring

3. Kubernetes & Containerization

Strong hands-on experience with Kubernetes and containerized workloads.

Experience with AKS / EKS / GKE or equivalent Kubernetes environments.

Hands-on experience with Helm deployments.

Understanding of Kubernetes troubleshooting, scaling, networking, and workload management.

4. Infrastructure as Code & DevOps

Hands-on experience with Terraform / Infrastructure as Code (IaC) .

Experience with Git-based workflows using GitHub, GitLab, or Azure Repos .

Strong DevOps automation and CI/CD understanding.

Strong scripting skills in Python and/or Bash .

5. Monitoring & Observability

Strong hands-on experience with OpenTelemetry .

Experience with monitoring and observability tools such as:

Prometheus

Grafana

Datadog

Azure Monitor

AWS CloudWatch

GCP Cloud Monitoring

Strong understanding of metrics, logs, distributed tracing, and alerting .

Experience implementing monitoring based on Golden Signals :

Latency

Traffic

Errors

Saturation

Ability to develop symptom-based, user-impact-focused alerting.

6. Linux & Networking

Strong knowledge of Linux system administration .

Strong understanding of:

DNS

TCP/IP

Load Balancing

SSL/TLS

Networking fundamentals

Experience supporting highly available production environments.

7. Incident & Reliability Engineering

Ability to rapidly diagnose and resolve high-severity production incidents .

Experience driving MTTR reduction .

Strong debugging and analytical problem-solving skills.

Ability to identify recurring issues and implement permanent corrective/preventive solutions.

Good-to-Have Skills

Experience working across Azure + AWS + GCP in a multi-cloud environment.

Knowledge of Go (Golang) .

Experience with OpenSearch / ELK Stack .

Experience supporting AI/ML workloads in production.

Exposure to Azure AI Services and Azure AI Foundry .

Experience supporting RAG (Retrieval-Augmented Generation) workloads.

Experience designing infrastructure for AI/ML platforms.

Experience building enterprise-wide OpenTelemetry observability frameworks .

Strong understanding of distributed systems architecture.

Exposure to advanced cloud-native architectures and reliability patterns.

Experience with security, compliance, vulnerability remediation, secrets management, and network segmentation.

Key ResponsibilitiesProduction & Incident Management

Participate in the 24/7 on-call rotation .

Diagnose, mitigate, and resolve production incidents.

Lead RCA and post-incident reviews.

Implement corrective and preventive actions.

Continuously improve MTTR and production stability.

Reliability Engineering

Define and improve SLIs, SLOs, SLAs, and Error Budgets .

Identify and eliminate operational toil.

Conduct reliability and capacity reviews.

Improve redundancy, failover, disaster recovery, and system resilience.

Cloud & Infrastructure

Manage and optimize cloud infrastructure across Azure, AWS, and/or GCP.

Manage Kubernetes clusters and containerized applications.

Implement and maintain Infrastructure as Code using Terraform.

Support CI/CD and Git-based development workflows.

Observability & Performance

Build and improve monitoring, logging, metrics, and tracing.

Implement OpenTelemetry and distributed tracing .

Establish Golden Signals-based monitoring and alerting.

Identify and resolve infrastructure and application performance bottlenecks.

Security

Implement cloud security best practices around IAM, network segmentation, and secrets management .

Support vulnerability remediation and compliance initiatives.

Collaborate with Development, Security, and Infrastructure teams.

Ideal Candidate

We are looking for someone with:

Strong SRE mindset and production ownership .

Excellent troubleshooting and incident-management skills.

Hands-on expertise in Cloud + Kubernetes + Terraform + Observability .

Strong understanding of OpenTelemetry and Golden Signals .

Experience working in highly available, production-critical environments.

Ability to remain calm and make effective decisions during critical incidents.

Strong communication and cross-functional collaboration skills.

Passion for automation, scalability, reliability, and continuous improvement .

Important Hiring Criteria

Must be:

7+ years relevant experience

Immediate joiner

Willing to work from office in Viman Nagar, Pune

Comfortable with 3:00 PM – 12:00 AM shift

Comfortable with 24/7 on-call rotation

Strong hands-on SRE/DevOps experience

Strong Cloud + Kubernetes + Observability experience

Strong production incident management experience

Good to have:

Multi-cloud: Azure + AWS + GCP

Original posting on Umanist Staffing LLC's site ↗

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job