hirly

Sysco

Senior Site Reliability Engineer

Sysco LABS - Sri Lanka

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Sysco first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Role family
Engineering
Seniority
Senior
Country
LK
Work mode
On-site / unstated
First seen by hirly
30 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

JOB DESCRIPTION

Senior Software Site Reliability Engineer

About Sysco LABS:

Sysco LABS is the Global In-House Center of Sysco Corporation (NYSE: SYY), the world’s largest foodservice company. Sysco ranks 55th in the Fortune 500 list and is the global leader in the trillion-dollar foodservice industry.

Sysco operates 333 distribution centers across 10 countries, with 75,000 colleagues serving approximately 670,000 customer locations, including restaurants, healthcare and educational facilities, lodging establishments, entertainment venues and more. For fiscal year 2026, which ended June 27, 2026, the company generated sales of more than $84 billion.

Sysco LABS Sri Lanka delivers the technology that powers Sysco’s end-to-end operations.

Sysco LABS’ enterprise technology is present in the end-to-end foodservice journey, enabling the sourcing of food products, merchandising, storage and warehouse operations, order placement and pricing algorithms, the delivery of food and supplies to Sysco’s global network and the in-restaurant dining experience of the end-customer.

The Opportunity

Join the Sysco Commercial Technology (CT) Site Reliability Engineering team as a Senior Software Site Reliability Engineer , where you will improve the reliability, scalability, performance, security, and operational excellence of Sysco's Commercial Technology ecosystem. Our team is responsible for delivering end-to-end reliability across the technology stack from cloud infrastructure and platform services to APIs, applications, and customer-facing digital experiences supporting enterprise API platforms, digital commerce, B2B integrations, and other business-critical systems across the CT landscape.

This is a Software SRE role focused on applying software engineering principles to solve reliability challenges at scale. You will design and build automation, develop internal tools and platforms, enhance observability, reduce operational toil, and deliver engineering solutions that strengthen the reliability and resilience of distributed systems.

As a Senior Software Site Reliability Engineer , you will combine software engineering, systems thinking, and production reliability expertise to identify and address reliability gaps, improve incident response, and drive long-term reliability initiatives. Working closely with product, platform, and infrastructure engineering teams, you will help build highly available, scalable, and resilient services while embedding reliability throughout the software development lifecycle.

Responsibilities:

Own the reliability, scalability, performance, and operational excellence of critical platforms and services across the Commercial Technology ecosystem.

Apply software engineering principles to design and build systems that improve reliability, resilience, and engineering productivity.

Design, develop, and maintain automation, internal platforms, self-service capabilities, and reliability engineering tools that eliminate operational toil.

Define and evolve service reliability through SLIs, SLOs, error budgets, observability standards, and actionable operational metrics.

Partner with product, platform, and infrastructure engineering teams to design reliable, scalable, and operable systems from inception through production.

Build engineering solutions that improve deployment safety, release automation, progressive delivery, and production readiness.

Investigate complex production issues using application code, distributed systems knowledge, logs, metrics, traces, and infrastructure telemetry to identify systemic failures and drive permanent solutions.

Lead incident reviews and postmortems, ensuring root causes are understood and translated into engineering improvements rather than operational workarounds.

Drive continuous reliability improvements by identifying recurring operational patterns and solving them through software, automation, or platform capabilities.

Improve the resilience of distributed systems through capacity planning, performance engineering, disaster recovery, and resilience testing.

Participate in architecture and design reviews to ensure systems are reliable, scalable, observable, secure, and operationally efficient.

Contribute to shared engineering libraries, frameworks, and platform capabilities that enable development teams to build and operate reliable services.

Participate in an on-call rotation, using operational experience to continuously improve system reliability, automate manual work, and reduce future operational burden.

Mentor engineers and champion Software SRE principles, engineering excellence, and a culture of reliability across the organization.

Requirements:

4+ years of experience in Site Reliability Engineering, Software Engineering, Platform Engineering, Production Engineering, or a related engineering role.

Strong software engineering skills, with experience designing, building, testing, and operating reliable production systems.

Proficiency in one or more programming languages, such as Java, Go, Python, or JavaScript/TypeScript , with a focus on writing maintainable, production-quality code.

Ability to read, understand, and debug application code to identify systemic issues and implement long-term engineering solutions.

Solid understanding of distributed systems, cloud-native architectures, microservices, APIs, databases, messaging systems, caching, networking, and system design.

Experience operating large-scale distributed systems, enterprise SaaS solutions, digital commerce platforms, or other high-traffic environments.

Experience applying Site Reliability Engineering principles, including SLIs, SLOs, error budgets, incident management, blameless postmortems, toil reduction, and automation .

Experience debugging complex production issues across applications, distributed systems, APIs, infrastructure, and cloud platforms.

Hands-on experience with observability practices, including metrics, logs, traces, and profiling, using platforms such as Datadog, Prometheus, Grafana, OpenTelemetry, Splunk, or ELK .

Experience with AWS, Azure, or GCP and cloud-native technologies such as containers and Kubernetes.

Experience with CI/CD, Infrastructure as Code, and platform engineering tools such as Terraform, Helm, Argo CD, Jenkins, or GitHub Actions .

Experience building internal developer platforms, automation frameworks, reliability engineering tools, or operational tooling that improves reliability and engineering productivity.

Experience with service meshes, API gateways, messaging systems, caching, database reliability, or resilience engineering.

Experience with AIOps, AI-assisted operations, or modern observability platforms.

Strong analytical and systems-thinking skills, with the ability to translate operational challenges into scalable engineering improvements.

Excellent communication and collaboration skills, with experience working across Product, Software Engineering, Platform Engineering, and Infrastructure teams.

Demonstrated ownership, curiosity, and a continuous-improvement mindset, with the ability to independently drive reliability initiatives.

Benefits:

Performance-based annual bonus

Performance rewards and recognition

Agile Benefits - special allowances for Health, Wellness & Academic purposes

Paid birthday leave

Team engagement allowance

Comprehensive Health & Life Insurance Cover - extendable to parents and in-laws

Hybrid work arrangement

Sysco LABS is an Equal Opportunity Employer.

Original posting on Sysco's site ↗

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job