hirly

Otppb

Site Reliability Engineering Lead, Product Enablement

Toronto, Canada

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Otppb first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.7M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Lead / management
Country
CA
Work mode
On-site / unstated
First seen by hirly
6 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

The opportunity

The Site Reliability Engineering Lead (SREL) is accountable for the reliability, health, and operational sustainability of a broader portfolio of applications, data pipelines, platforms, and services. The role proactively identifies systemic and recurring risks and leads or directs technical initiatives, including the design and implementation of monitoring, automation, operational support tooling, and selected code to reduce incidents, manual intervention, cost, and operational burden. The SREL serves as a senior technical advisor and escalation point, translating complex issues, options, risks, and trade-offs into clear business language for stakeholders.

Who you'll work with

  • Reports to: Director, Product / Data Enablement
  • Works closely with: Site Reliability Engineering Specialists, Product Engineering Leads (PEL), Data Solutions Leads, Data Platform Leads, Enterprise Architecture, Security, Technology Services, business stakeholders, and third-party vendors.
  • Collaborates regularly with Product Engineering, Data Solutions, and Data Platform teams to support smooth transitions from build to run and to identify and resolve recurring reliability and supportability issues. Depending on the nature and complexity of the change, the Lead may implement or direct selected code, configuration, automation, or process changes through development, testing, release, and post-implementation validation, or work with the appropriate team to secure prioritization, ownership, and completion.
  • Partners with Technology Services teams including Infrastructure, Cloud, and Platform teams to troubleshoot complex and cross-service issues and to design and implement, or direct the implementation of, observability, automation, resilience, and operational support tooling that improves service quality and reduces operational burden.
  • Frequent communication with business stakeholders is required to provide transparency into system health, incident status, systemic risks, and improvement priorities, adapting the level of technical detail to the audience, including senior business stakeholders when required.

What you'll do

Service Reliability Strategy and Ownership

  • Own and continuously improve the reliability, supportability, and operational efficiency of a broad portfolio of applications, data pipelines, platforms, and services including availability, latency, performance, resilience, and cross-service dependencies.
  • Define and govern Service Level Availability (SLAs), and related reliability metrics for the portfolio in accordance with enterprise standards and business expectations and drive changes where performance or operational risk warrants.
  • Establish and maintain a prioritized reliability roadmap incorporating application, data pipeline and service health plans balancing business criticality, operational risk, support effort, capacity, cost, and technology priorities
  • Use deep technical and trend analysis across incidents, alerts, service performance, capacity, storage, cost, and support effort to identify systemic failure modes and lead cross-service improvement initiatives through implementation and outcome validation.

Operational Excellence

  • Lead response to high-impact, complex, or cross-service incidents, coordinating technical teams and ensuring timely audience-appropriate stakeholder communication.
  • Drive Root Cause Analysis and problem-management activities beyond minimum process requirements, identify systemic technical, architectural, monitoring, process, and organizational contributing factors, and challenge recurring issues that have become normalized as operational support work. Ensure corrective or preventive actions are completed and validated through permanent remediation, automated mitigation, or appropriately authorized risk acceptance.
  • Own the end-to-end resolution of selected recurring production defects and supportability issues. Where appropriate, implement or direct code, configuration, scripting, automation, and process changes through technical review, testing, release, and post-implementation validation, and drive changes requiring Product or Data ownership through prioritization and completion.Lead portfolio-level capacity, performance, resilience, and recovery analysis and improve Mean Time to Detect (MTTD) and Mean Time to Restore through monitoring, diagnostics, runbooks, automation, and operational learning.
  • Contribute to Disaster Recovery (DR) and resilience plans and participate in recovery testing.

Observability & Automation

  • Design and implement, or direct the implementation of, portfolio-wide monitoring, observability, event-management, diagnostic, automation, self-healing, and operational support tools that improve support effectiveness and service quality.
  • Assess and rationalize existing alerts, automated emails, service checks, dashboards, automated restarts, runbooks, and integrations to reduce duplication and noise, simplify support processes, and clarify ownership.
  • Define target-state standards and reusable patterns for telemetry, dashboards, alerting, event correlation, diagnostics, and automation, and oversee adoption, technical quality, documentation, and sustainable support arrangements.
  • Measure the effectiveness of monitoring and automation improvements through reductions in non-actionable alerts, recurring incidents, manual intervention, support effort, Mean Time to Detect, Mean Time to Restore, and avoidable technology cost.

Production Readiness & Governance

  • Establish minimum portfolio standards for production readiness, supportability, monitoring, resilience, runbooks, operational documentation, service ownership, and recovery capabilities.
  • Lead or provide technical assurance for complex or high-risk production-readiness reviews, engaging Product Engineering, Data Solutions, Data Platform, Architecture, Security, and Technology Services teams, as appropriate, during design reviews to identify and address reliability, resilience, and supportability risks.
  • Ensure unresolved production-readiness gaps have documented remediation plans, compensating controls, or appropriately authorized risk acceptance before release.
  • Ensure changes implemented or led by the SRE function follow established source-control, peer-review, testing, security, change-management, release, rollback, documentation, and production-validation requirements.
  • Monitor and govern vendor service performance against defined SLAs driving escalation and service performance plans through the appropriate vendor and portfolio governance channels.

Stakeholder Management

  • Present portfolio health, reliability trends, systemic risks, operational burden, and improvement progress in team and governance forums.
  • Act as the senior technical escalation point for major or cross-service production incidents.
  • Translate complex technical risks, business impacts, solution options, costs, dependencies, and trade-offs into clear recommendations for business and technology stakeholders, including senior stakeholders when required.
  • Build trusted relationships and use operational evidence to influence Product, Data, Technology Services, Architecture, and vendor teams to prioritize corrective and preventive reliability work.
  • This role does not carry formal people management accountability but provides functional and technical leadership across the production support function, including setting reliability priorities and standards, coordinating work, and reviewing technical approaches and outcomes.
  • The SREL provides work direction, coaching, and mentoring to Site Reliability Engineering Specialists and operational support resources, strengthening diagnostic, automation, and independent problem-solving capability and assuring the technical quality and completion of reliability work performed under its direction. Formal performance management and employment decisions remain with the applicable people leader.

Th

Original posting on Otppb's site ↗

Listed on hirly, a job board. hirly is not the employer: Otppb is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job