IBM
Staff Developer - Incident Command
Multiple Cities, Canada
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Lead / management
- Country
- CA
- Work mode
- On-site / unstated
- First seen by hirly
- 27 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
At IBM Software, we transform client challenges into solutions. Building the world’s leading AI-powered, cloud-native products that shape the future of business and society. Our legacy of innovation creates endless opportunities for IBMers to learn, grow, and make an impact on a global scale. Working in Software means joining a team fueled by curiosity and collaboration. You’ll work with diverse technologies, partners, and industries to design, develop, and deliver solutions that power digital transformation. With a culture that values innovation, growth, and continuous learning, IBM Software places you at the heart of IBM’s product and technology landscape. Here, you’ll have the tools and opportunities to advance your career while creating software that changes the world. Confluent Cloud processes millions of events per second across AWS, GCP, and Azure. When incidents happen in a multi-cloud streaming platform, they happen at scale—data in motion, exactly-once semantics, and cascading failure modes that require deep systems thinking. We need an expert-level engineer who can drive proactive reliability improvements that prevent these incidents before they occur. This role combines hands-on technical work with strategic program ownership. You'll spend roughly 75% of your time on engineering: building automation, improving tooling, analyzing systemic failure patterns, and designing reliability improvements. The remaining 25% is teaching and coordination: coaching teams through post-mortems, training incident commanders, and evolving our incident response practices. You'll be part of a global team with follow-the-sun coverage, with clean handoffs that keep everyone working sustainable hours. This role sits within Cloud Architecture and Reliability - Supportability, a horizontal team that owns reliability standards and tooling across engineering. You're the person who makes us need incident management less. What You Will Do: * Analyze systemic failure patterns and design reliability improvements that prevent incident recurrence * Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack * Define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments * Own standards, practices, and continuous improvement of incident response across engineering * Edit and review customer-facing incident documents (CRCAs) to ensure quality and clarity * Develop and deliver training programs; coach teams through post-mortems * Partner with engineering leaders to elevate reliability practices org-wide * Deep experience with observability: metrics, logging, tracing * Kubernetes and container orchestration experience * Understanding of CI/CD pipelines and release processes * Strong written communication (design docs, runbooks, post-mortems) * Experience driving org-wide process and cultural changes * 10+ years of relevant experience in SRE, incident management, or reliability engineering * Cloud experience with at least one of AWS, GCP, or Azure (we run all three) * Experience navigating reliability/incident programs at 500+ engineer organizations * Deep expertise with incident management tooling (Rootly, PagerDuty, or similar) * Strong understanding of distributed systems and failure modes at scale * Kafka/event streaming expertise preferred, or demonstrated rapid mastery of complex systems • Advanced Cloud Knowledge: Experience with cloud-based infrastructure and its application in reliability and resiliency engineering. • Specialized Scripting Skills: Proficiency in scripting languages and automation tools to optimize system reliability and performance. Chez IBM Software, nous transformons les défis des clients en solutions. Construire les produits natifs du cloud propulsés par l’IA et les plus avancés au monde, qui façonnent l’avenir des affaires et de la société. Notre héritage d’innovation crée d’innombrables occasions pour les employés d’IBM d’apprendre, de grandir et d’avoir un impact à l’échelle mondiale. Travailler en logiciel, c’est rejoindre une équipe animée par la curiosité et la collaboration. Vous travaillerez avec diverses technologies, partenaires et industries pour concevoir, développer et livrer des solutions qui alimentent la transformation numérique. Avec une culture qui valorise l’innovation, la croissance et l’apprentissage continu, IBM Software vous place au cœur du paysage des produits et des technologies d’IBM. Ici, vous aurez les outils et les opportunités pour faire avancer votre carrière tout en créant des logiciels qui changent le monde. Avec Confluent, les données ne restent pas immobiles. Nous mettons l’information en mouvement, en diffusant presque en temps réel afin que les organisations puissent réagir plus rapidement, construire de façon plus intelligente et offrir des expériences aussi dynamiques que le monde qui les entoure. À propos du poste : Confluent Cloud traite des millions d’événements par seconde via AWS, GCP et Azure. Lorsque des incidents surviennent sur une plateforme de diffusion multi-nuages, ils se produisent à grande échelle — données en mouvement, sémantique d’une seule fois, et modes de défaillance en cascade qui nécessitent une réflexion systémique profonde. Nous avons besoin d’un développeur de niveau expert capable d’améliorer la fiabilité de façon proactive afin de prévenir ces incidents avant qu’ils ne surviennent. Ce rôle combine un travail technique pratique avec la gestion stratégique des programmes. Vous passerez environ 75% de votre temps à l’ingénierie : automatisation des bâtiments, amélioration des outils, analyse des schémas de défaillance systémiques et conception d’améliorations de fiabilité. Les 25% restants sont consacrés à l’enseignement et à la coordination : accompagnement des équipes lors des autopsies, formation des commandants d’incident et évolution de nos pratiques de réponse aux incidents. Vous ferez partie d’une équipe mondiale avec une couverture de suivi du soleil, avec des transferts propres qui permettent à tout le monde de travailler des heures durables. Ce rôle s’inscrit dans l’architecture du cloud et la fiabilité - Supportabilité, une équipe horizontale qui détient les normes de fiabilité et les outils à travers l’ingénierie. C’est toi qui nous fais moins besoin de gestion des incidents. Ce que vous ferez : Analyser les schémas de défaillance systémiques et les améliorations de fiabilité de conception qui empêchent la réapparition des incidents Configuration Rootly, flux de travail et intégrations avec PagerDuty, Jira, Confluence et Slack Définir et maintenir les cadres SLO/SLA; utiliser des budgets d’erreur pour guider les investissements en fiabilité Ses propres normes, pratiques et amélioration continue de la réponse aux incidents en ingénierie Modifier et examiner les documents d’incident destinés aux clients (CRCA) afin d’assurer la qualité et la clarté Développer et offrir des programmes de formation; coacher les équipes jusqu’aux autopsies Collaborez avec des leaders en ingénierie pour améliorer les pratiques de fiabilité à l’échelle de l’organisation Expérience approfondie de l’observabilité : métriques, journalisation, traçage Expérience sur Kubernetes et orchestration de conteneurs Compréhension des pipelines CI/CD et des processus de publication Communication écrite solide (documents de conception, manuels d’exécution, autopsies) Expérience de l’expérience à la conduite des changements de processus et culturels à l’échelle de l’organisation 10+ ans d’expérience pertinente en SRE, gestion d’incidents ou génie de la fiabilité Expérience infonuagique avec au moins un modèle d’AWS, GCP ou Azure (nous utilisons les trois) Expérience dans la navigation dans des programmes de fiabilité/incidents dans 500+ organisations d’ingénieurs Expertise approfondie avec les outils de gestion des incidents (Rootly, PagerDuty, ou similaires) Solide compréhension des systèmes distribués et des modes de défaillance à grande éche
Similar jobs
- Principal Software DeveloperSolace · Ottawa, OntarioFirst seen today
- Principal CRM DeveloperCondoauthorityontario · Toronto, OntarioFirst seen today
- Lead Full Stack DeveloperRbc · TORONTO, Ontario, CanadaFirst seen today
- Java Developer, Equities Trading, Vice PresidentCiti · Mississauga Ontario CanadaFirst seen today
- Associate DeveloperRux · RemoteFirst seen todayremote
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job