This posting is no longer listed by Riotgames.
hirly last saw it live on 2 September 2026. Similar roles are on the live board.
Riotgames
Principal Software Engineer - DevOps / Site Reliability Engineer
Singapore
Apply through hirly
hirly scores this role against your resume, shows its reasoning, then writes a resume and cover letter for it and fills the application with you. Free to start — no card required.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Lead / management
- Country
- SG
- Work mode
- Remote-friendly
- First seen by hirly
- 2 Sept 2026
Derived automatically from the posting. Sign up to see how the role scores against your own resume.
the posting
Riot Games was established in 2006 by entrepreneurial gamers who believe that player-focused game development can result in great games. In 2009, Riot released its debut title League of Legends to critical and player acclaim. As the most played PC game in the world, over 100 million play every month. Players form the foundation of our community and it’s for them that we continue to evolve and improve the League of Legends experience.
We’re looking for humble but ambitious, razor-sharp professionals who can teach us a thing or two. We promise to return the favor. Like us, you take play seriously; you’re passionate about games. We embrace those who see things differently, aren’t afraid to experiment, and who have a healthy disregard for constraints.
That's where you come in.
The AI Efficiency team at Riot Games builds the platforms, tools, and technical foundations that help Rioters safely and effectively use AI to accelerate how we work. As these systems become increasingly important to creative, product, and development workflows across Riot, we need dedicated engineering leadership to ensure these systems remain stable, scalable, secure, and dependable in production.
As a Principal DevOps / Site Reliability Enginee r on the AI Efficiency team, you will own and evolve the operational foundations that allow the AI Efficiency team’s tech platform and the tools deployed within it to run reliably at growing scale. You will establish the systems, standards, automation, and support practices required to move quickly without compromising availability, deployment safety, maintainability, or user trust.
You will partner closely with software engineers, ML platform engineers, technical artists, data scientists, and Riot’s infrastructure and security teams to improve developer experience, production readiness, observability, incident response, capacity planning, and service resilience. You will also help evaluate and operationalize AI-native engineering workflows such as agent-assisted code review, automated bug triage, AI-driven performance and security analysis, and browser-based UI validation. This role ensures the broader platform and its services are safely operated, supported, and continuously improved in production.
You’re right for this role if you enjoy making complex systems reliable, reducing operational toil, improving how engineers build and ship software, and anticipating how systems will fail before those failures affect users. You are comfortable taking ownership of production health, leading through incidents, building sustainable operational practices, and creating paved roads that help teams move quickly and safely. You are also energized by the opportunity to responsibly bring new AI-native automation patterns into real engineering workflows, thoughtfully applying emerging capabilities to reduce friction, improve reliability, and enhance how engineers interact with production systems without compromising safety or control.
Responsibilities:
Own and continuously improve the reliability, availability, scalability, performance, and operational health of the Efficiency team’s (web) platform and the tools deployed within it
Design, build, and maintain the infrastructure, deployment systems, and operational foundations required to support a growing portfolio of production AI services and internal tools
Improve CI/CD pipelines, release engineering practices, environment management, and deployment automation so software can be shipped safely, quickly, and consistently
Establish production-readiness standards and ensure new utilities have appropriate monitoring, alerting, ownership, documentation, rollback strategies, and support plans before launch
Define and operationalize service health indicators, SLIs, SLOs, error budgets, and reliability metrics that guide engineering priorities and tradeoffs between reliability, velocity, cost, and complexity
Build comprehensive observability across applications, infrastructure, service dependencies, and user workflows using metrics, logs, traces, dashboards, synthetic monitoring, and actionable alerts
Establish sustainable incident-management and on-call practices, including escalation paths, runbooks, severity definitions, communication protocols, and clear service ownership
Lead or contribute to the diagnosis and resolution of production incidents, coordinating across teams and driving blameless post-incident reviews and durable corrective actions
Build automation that reduces operational toil, improves mean time to detect and recover, and eliminates recurring sources of failure or manual intervention
Implement safe deployment patterns such as automated validation, progressive delivery, canary releases, feature flags, health checks, rollback mechanisms, and controlled environment promotion
Perform capacity planning, load testing, performance analysis, and resource forecasting to ensure the Toolkit can support increasing adoption and usage across Riot
Design and validate resilience, backup, recovery, failover, and disaster-recovery strategies for critical services, data, configurations, and infrastructure
Identify single points of failure and systemic risks across applications, cloud infrastructure, networking, databases, queues, caches, third-party dependencies, and operational workflows
Improve developer experience by building self-service workflows, reusable infrastructure components, local development environments, test environments, deployment tooling, and clear operational documentation
Establish and maintain infrastructure-as-code, configuration-management, secrets-management, and environment-governance practices that make infrastructure changes safe, repeatable, and auditable
Partner with engineers throughout the software development lifecycle to embed reliability, operability, security, and maintainability into system design rather than addressing them only after launch
Troubleshoot complex production issues across web applications, APIs, distributed services, containerized workloads, cloud infrastructure, network boundaries, authentication systems, and external service dependencies
Partner with ML Platform Engineers to ensure model-serving and inference systems integrate cleanly with the team’s broader observability, deployment, incident-management, and reliability standards
Collaborate with Riot infrastructure, information security, IT, developer-platform, and compliance teams to ensure the team follows appropriate operational and security requirements
Evaluate and implement AI-assisted operational workflows such as automated anomaly investigation, log analysis, remediation recommendations, regression detection, and runbook automation
Define guardrails, approval requirements, auditability, and escalation paths for agentic or automated operational systems that can interact with production environments
Champion operational excellence through technical leadership, mentoring, documentation, standards, architecture reviews, and tooling that raise the reliability bar across the team
Required Qualifications:
Bachelor’s degree in Computer Science or a related field, or equivalent professional experience
5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Platform Engineering, Production Engineering, Developer Experience, or a similar role supporting production systems
Strong programming and automation skills in one or more languages such as Python, Go, JavaScript, or TypeScript
Experience designing, operating, and improving cloud-based production systems in AWS, GCP, Azure, or comparable environments
Experience building and maintaining CI/CD pipelines, release systems, deployment automation, and environment-management workflows
Strong understanding of observability practices, including metrics, logging, distributed tracing, dashboards, synthetic monitoring, and alert design
Is this role actually a fit for you?
hirly answers with a score and its reasoning, then writes the resume and cover letter if you decide to go for it.
Score it against my resume