Roku
Senior Software Engineer, SRE
Bengaluru, India
Apply through hirly
Upload your resume and get a version tailored to this job, plus a cover letter, in about thirty seconds — before you create an account.
Apply with hirlyhirly's read of this role
- Role family
- Engineering
- Seniority
- Senior
- Country
- IN
- Work mode
- Remote-friendly
- First seen by hirly
- 2 Sept 2026
Derived automatically from the posting. Sign up to see how the role scores against your own resume.
the posting
Teamwork makes the stream work.
Roku is changing how the world watches TV
Roku is the #1 TV streaming platform in the U.S., Canada, and Mexico, and we've set our sights on powering every television in the world. Roku pioneered streaming to the TV. Our mission is to be the TV streaming platform that connects the entire TV ecosystem. We connect consumers to the content they love, enable content publishers to build and monetize large audiences, and provide advertisers unique capabilities to engage consumers.
From your first day at Roku, you'll make a valuable - and valued - contribution. We're a fast-growing public company where no one is a bystander. We offer you the opportunity to delight millions of TV streamers around the world while gaining meaningful experience across a variety of disciplines.
What does the team work on?
The Platform Infrastructure team ensures that all Roku systems run smoothly. These systems support over 100M+ users and billions in transaction revenue per year. We are a group of highly skilled infrastructure and software engineers who help build and operate systems at internet scale, including Platform (Kubernetes, Istio, Envoy, operators, and more) and Observability (OSS/CNCF-supported observability projects). We engage with multiple teams to achieve company-impacting results.
What is the role?
We are seeking a talented and experienced SRE (Site Reliability Engineering) Senior Software Engineer to help architect, build, and operate large-scale systems that stay reliable, secure, and cost-effective at internet scale. The ideal candidate takes end-to-end ownership of outcomes—treating reliability, security, cost, operability, and supportability as part of the job, not just delivering code. They bring calm, decisive incident leadership, separating mitigation from root-cause investigation and running blameless reviews that produce lasting improvements. Strong judgment and prioritization are essential, balancing roadmap delivery against operational debt, security, and compliance while clearly explaining trade-offs. This engineer pairs deep technical depth in distributed systems with broad systems thinking, and turns ambiguous objectives into executable roadmaps, epics, and backlogs. Just as important is the ability to build influence through credibility and sound reasoning, coach other engineers, and raise the operational capability of the whole team. If you enjoy solving intriguing system challenges, are innovative at heart, and thrive on making a measurable impact across teams, this role might be a great fit for you.
How will I use AI at Roku?
At Roku, we don't just use AI, we work with it. AI agents and smart tools help power drafts, analysis, and repetitive workflows, while our people bring direction, judgment, and accountability.
We value curious, adaptable builders, who can show how they've used AI, agents, or automation to move faster, improve quality, and scale their impact. Strong candidates know how to frame problems, guide AI-assisted work, check the output, and learn quickly. Above all, they bring curiosity, adaptability, and sound judgment.
What are the responsibilities of the role?
Ownership & Incident Leadership
p]:inline" data-streamdown="list-item">Take responsibility for service outcomes end to end, including reliability, security, cost, operability, and supportability
p]:inline" data-streamdown="list-item">Lead major incidents with composure when information is incomplete, separating mitigation from root-cause investigation and communicating impact, status, risks, and next steps without speculation
p]:inline" data-streamdown="list-item">Facilitate comprehensive, blameless post-incident reviews that identify root causes and contributing factors, and follow corrective actions through to completion
p]:inline" data-streamdown="list-item">Track incident trends to surface systemic issues and prioritize reliability improvements
p]:inline" data-streamdown="list-item">Implement chaos engineering, game days, and disaster recovery exercises to validate resilience and build confidence in recovery procedures
SRE Process & Principles Implementation
p]:inline" data-streamdown="list-item">Establish and evolve SRE principles, frameworks, and methodologies across the organization, including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets
p]:inline" data-streamdown="list-item">Manage Error Budgets as a data-driven mechanism for balancing feature velocity against reliability, and facilitate risk-tolerance conversations between engineering and product teams
p]:inline" data-streamdown="list-item">Use SLOs, error budgets, incident data, and operational metrics to guide where the team invests
Reliability Engineering & Infrastructure
p]:inline" data-streamdown="list-item">Reduce toil by identifying repetitive operational work and eliminating it through infrastructure-as-code, automation frameworks, and intelligent tooling, aiming to keep toil below 50% of team time
p]:inline" data-streamdown="list-item">Prefer sustainable corrective action over repeated manual intervention
p]:inline" data-streamdown="list-item">Implement capacity planning that ensures adequate headroom to meet SLOs during peak traffic, load spikes, and degraded states, including predictive models and automated scaling
Observability, Monitoring & Reporting
p]:inline" data-streamdown="list-item">Build observability systems that provide deep visibility into service health, performance, and user experience using the Four Golden Signals and USE/RED methodologies
p]:inline" data-streamdown="list-item">Create SRE dashboards and reporting that give real-time visibility into SLO compliance, error budget consumption, and reliability metrics, including executive-level reporting on trends and incident impact
p]:inline" data-streamdown="list-item">Establish actionable, symptom-based alerting aligned with SLOs, tuning thresholds to reduce noise while ensuring critical issues trigger appropriate responses
p]:inline" data-streamdown="list-item">Treat an alert without ownership or a usable runbook as an incomplete control
Collaboration, Influence & Leadership
p]:inline" data-streamdown="list-item">Partner with development teams to design reliability in from the start, conducting design reviews focused on failure modes, scalability, observability, and operational concerns
p]:inline" data-streamdown="list-item">Build alignment across product, application, security, compliance, and infrastructure teams, managing disagreements with evidence, impact, and trade-offs
p]:inline" data-streamdown="list-item">Gain support through credibility, relationships, and sound reasoning rather than title or tenure, and lead cross-functional work where participants do not report to you
p]:inline" data-streamdown="list-item">Coach engineers by delegating meaningful ownership, giving specific and timely feedback, and creating opportunities for others to lead projects and incidents
p]:inline" data-streamdown="list-item">Manage project priorities using error budgets as a decision-making framework, ensuring reliability work is prioritized alongside feature development
Delivery & Execution Discipline
p]:inline" data-streamdown="list-item">Convert roadmap objectives into clear epics, stories, milestones, owners, and acceptance criteria, identifying dependencies and risks early
p]:inline" data-streamdown="list-item">Track outcomes rather than activity, and adjust scope transparently when incidents or unplanned work affect delivery
p]:inline" data-streamdown="list-item">Finish work completely, including testing, documentation, monitoring, runbooks, rollout, and operational handoff
Operational Excellence & Continuous Improvement
p]:inline" data-streamdown="list-item">Identify and eliminate performance bottlenecks through analysis of metrics, traces, and profiles, and
Similar jobs
- Senior Software EngineerWeekday AI · IndiaFirst seen today
- Software Engineer, Senior (Java React)MicroStrategy India · Chennai, Tamil Nadu, IndiaFirst seen today
- Sr Software EngineerRenesas Electronics · Indiranagar, Bangalore, Karnataka, IndiaFirst seen today
- Sr Software EngineerServiceNow · Hyderabad, Telangana, IndiaFirst seen today
- Senior Software Engineer IThe Walt Disney Company · IndiaFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job