Amgen
Manager, Site Reliability Engineer - Data Platforms
India - Hyderabad
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Lead / management
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 24 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Career Category
Engineering
Job Description
Join Amgen’s Mission of Serving Patients
At Amgen, if you feel like you’re part of something bigger, it’s because you are. Our shared mission-to serve patients living with serious illnesses-drives all that we do.
Since 1980, we’ve helped pioneer the world of biotech in our fight against the world’s toughest diseases. With our focus on four therapeutic areas -Oncology, Inflammation, General Medicine, and Rare Disease- we reach millions of patients each year. As a member of the Amgen team, you’ll help make a lasting impact on the lives of patients as we research, manufacture, and deliver innovative medicines to help people live longer, fuller happier lives.
Our award-winning culture is collaborative, innovative, and science based. If you have a passion for challenges and the opportunities that lay within them, you’ll thrive as part of the Amgen team. Join us and transform the lives of patients while transforming your career.
Manager, Site Reliability Engineer - Data Platforms
About Amgen
Amgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today.
About the Role
AWS Site Reliability Engineer with strong cloud architecture expertise to design, build, automate, and operate secure, resilient, scalable, and cost-efficient AWS platforms. You will combine AWS architecture leadership with practical Site Reliability Engineering-writing Infrastructure as Code, developing automation, improving observability, solving complex production problems, and strengthening operational excellence.
This role also includes responsibility for leading an SRE pod and managing assigned engineers. However, it is primarily a hands-on technical role. The successful candidate will remain a key technical contributor who leads through architecture, implementation, production ownership, incident response, coaching, and example.
What you will do
Roles & Responsibilities:
AWS Architecture and Platform Engineering
- Design secure, highly available, scalable, and cost-efficient AWS architectures for EDSE applications and data platforms.
- Define reusable platform services, Infrastructure as Code modules, reference architectures, and engineering guardrails.
- Make and document architectural decisions across networking, identity, compute, containers, storage, security, observability, and resilience.
- Lead architecture and production-readiness reviews, addressing reliability, security, performance, operability, and cost risks.
Reliability and Production Operations
- Establish and improve SRE practices, including SLIs, SLOs, error budgets, availability targets, capacity planning, and operational-readiness criteria.
- Use operational data to identify recurring failures, performance bottlenecks, capacity risks, and opportunities to reduce manual toil.
- Participate in the on-call rotation and lead the technical response to complex or high-severity production incidents.
- Conduct blameless post-incident reviews and ensure corrective actions produce durable engineering improvements.
Automation, Delivery, and Observability
- Build reusable, tested infrastructure and operational automation using Terraform and programming languages such as Python.
- Automate provisioning, configuration, validation, deployment, recovery, compliance checks, and routine operational activities.
- Strengthen CI/CD and GitOps practices through automated testing, deployment controls, progressive delivery, and reliable rollback mechanisms.
- Standardize metrics, logs, traces, dashboards, and actionable alerting to improve issue detection, diagnosis, and recovery.
Resilience, Security, and Cost Efficiency
- Design and validate backup, high-availability, and disaster-recovery strategies aligned with defined business and recovery objectives.
- Conduct restore tests, failover exercises, resilience reviews, and controlled game-day scenarios.
- Embed least-privilege access, encryption, network segmentation, secrets management, vulnerability remediation, and policy-based controls into platform designs.
- Improve AWS cost efficiency through right-sizing, resource lifecycle management, tagging, usage analysis, and architecture optimization without compromising reliability or security.
Pod and People Leadership
- Lead the SRE pod by setting technical direction, prioritizing work, managing operational commitments, and driving delivery to completion.
- Serve as a senior AWS and SRE advisor, partnering with application engineering, data engineering, cybersecurity, architecture, product, and FinOps teams.
- Manage and mentor assigned engineers through regular feedback, one-on-one discussions, technical coaching, design and code reviews, and career-development support.
The expected allocation of responsibilities is:
- Approximately 75-80% hands-on technical work , including AWS architecture, coding, Infrastructure as Code, automation, design reviews, production troubleshooting, observability, performance improvement, and incident response.
- Approximately 20-25% pod and people leadership , including prioritization, work allocation, delivery coordination, mentoring, one-on-one discussions, performance feedback, career development, and removing team blockers.
The individual will be expected to move comfortably between architecture decisions, hands-on implementation, production operations, incident leadership, stakeholder communication, and people development. The balance may vary temporarily during major incidents, critical releases, or important delivery milestones.
What we expect of you
Basic Qualifications and Experience:
- Master’s or Bachelor’s degree in computer science or engineering field and 9 to 12 years of relevant experience, including substantial hands-on experience in AWS cloud engineering, platform engineering, infrastructure engineering, DevOps, or Site Reliability Engineering.
- Prior experience leading a technical pod, engineering squad, or small team while continuing to contribute hands-on.
Must-Have Skills:
- Significant hands-on experience designing, implementing, and operating business-critical production workloads on AWS, including experience with EKS, Sagemaker, Bedrock, VPC, PrivateLink, S3, EC2, KMS, CloudWatch, CloudTrail, Secrets Manager, Lambda, and RDS.
- Deep knowledge of AWS architecture across networking, identity and access management, security, compute, containers, storage, monitoring, and resilience.
- Strong experience with Infrastructure as Code using Terraform, CloudFormation, AWS CDK, or comparable technologies.
- Proficiency in at least one programming or scripting language, such as Python, Go, Java, TypeScript, Bash, or PowerShell.
- Ability to write maintainable, production-quality infrastructure code, automation, and operational tooling.
- Practical experience applying SRE principles, including SLIs and SLOs, actionable alerting, incident response, root-cause analysis, post-incident improvement, and toil reduction.
- Experience building or operating containerized platforms using Docker and Kubernetes, preferably Amazon EKS.
- Experience with CI/CD, GitOps, source control, automated infrastructure testing, deployment controls, and safe production-change practices.
- Experience implementing observability using AWS CloudWatch, OpenTelemetry, Prometheus, Grafana, Splunk, Datadog, or comparable platforms.
- Strong knowledge of Linux, networking, distributed systems, performance analysis, high-availability design, and disaster recovery.
- Demonstrated ability to troubleshoot complex production systems and lead technical decision-making during high-severity incid
Similar jobs
- Planlaufmanager (m/w/d) in Voll- oder Teilzeitquattron · Berlin, GermanyFirst seen today
- (Junior) BIM Manager / BIM Gesamtkoordinator (Remote) (m/w/d) in Vollzeitquattron · Frankfurt am Main, Hessen, GermanyFirst seen todayremote
- Qualitätsverantwortlicher / Qualitätsmanager (all genders)Viega · Bad Sulza, Thüringen, GermanyFirst seen today
- BIM-Manager (w/m/d)Die Autobahn GmbH des Bundes · Hohen Neuendorf, Brandenburg, GermanyFirst seen today
- Technischer Sales Manager (m/w/d) Messtechnik Region NordPassion for People GmbH · Lübeck, Schleswig-Holstein, GermanyFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job