hirly

ZEISS

Senior Site Reliability Engineer

Bangalore

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at ZEISS first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.7M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Senior
Country
IN
Work mode
On-site / unstated
First seen by hirly
3 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

ZEISS in India

ZEISS in India is headquartered in Bengaluru and present in the fields of Industrial Quality Solutions, Research Microscopy Solutions, Medical Technology, Vision Care and Sports & Cine Optics.

ZEISS India has 3 production facilities, R&D center, Global IT services and about 40 Sales & Service offices in almost all Tier I and Tier II cities in India. With 2200+ employees and continued investments over 25 years in India, ZEISS’ success story in India is continuing at a rapid pace.

Further information at ZEISS India .

As a Site Reliability Engineer, you will bridge the gap between Development, Cloud Platform Engineering Team and Product Owners of different Digital Offerings. Define the SRE-concepts, align the service quality with the business objectives and user expectations:

What you will do:

Define and measures the reliability of the service using SLI, SLOs and consider the risk minimization of service degradation.

Enable the development team to bring new software or new features to production as quickly as possible, while also ensuring an agreed-upon acceptable level of IT operations performance and error risk in line with the service level agreements (SLAs) agreed.

Define and drive observability for self-developed software and the managed cloud components by collecting appropriate observability data for insights and alerting including setting up proper alerting for critical components.

Ensure availability and responsiveness of application by setting up and maintaining the required documentation method and tools. Building Playbooks, Runbooks for troubleshooting techniques to effectively identify and investigate issues.

Setup automation and CI/CD Pipelines for continuously deploying software applications using the best practices and review them for improvement.

Monitor systems for performance and reliability, respond to incidents, and conduct post-incident reviews to identify root causes and improve system resilience.

Create and maintain comprehensive documentation for systems, processes, disaster recovery plans, and procedures.

  • In-depth knowledge of system architecture, networking, and distributed systems.
  • Expertise in designing and implementing reliable, scalable, and fault-tolerant systems.
  • Proficiency in setting up and managing monitoring, alerting, and logging systems for early detection and resolution of issues for container orchestrators like Kubernetes using Tools like Prometheus, Grafana, Open Telemetry Collector or similar tools.
  • Hands-on experience in incident management, including incident response, troubleshooting, and post-mortem analysis.
  • Proficiency in coding/scripting languages commonly used in infrastructure automation and monitoring (such as Terraform).
  • Familiar with deployment process and strategies.
  • Knowledge of best practices in disaster recovery planning and execution for cloud based Systems.
  • Capability to advocate for SRE best practices and principles within the organization and drive cultural changes as needed.
  • Willingness to stay updated with the latest trends, tools, and technologies in the field of site reliability engineering.
  • Strong communication skills to effectively collaborate with cross-functional teams, including Software Developers, Product Owners, and Cloud Platform Engineers.

Your ZEISS Recruiting Team:

Saptarshi Chowdhury, Sultana Mehzabin

Original posting on ZEISS's site ↗

Listed on hirly, a job board. hirly is not the employer: ZEISS is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job