hirly

BJ's Wholesale Club

Principal Engineer - Major Incident Response & ITIL Platform

BJ's Club Support Center Marlborough, MA #5997

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at BJ's Wholesale Club first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Lead / management
Country
US
Work mode
On-site / unstated
First seen by hirly
27 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

A World-Class Team

BJ’s Wholesale Club is powered by more than 30,000 team members who make a real impact every day. Whether you're stocking shelves, solving problems or shaping strategy, your work helps families save on what matters most.

We’re a team built on purpose and opportunity. Join us and be part of something meaningful.

Why You’ll Love Working at BJ’s

At BJ’s Wholesale Club, our team members are at the heart of everything we do. That’s why we offer a comprehensive benefits package designed to support your health, well-being and future – both on and off the job. When you grow, we grow.

Here’s just some of what you can look forward to:

  • Weekly Pay: Get paid every week so that you can manage your money on your terms.
  • Free BJ’s Memberships: Enjoy a complimentary The Club Card Membership, plus a free Supplemental Membership for someone in your household.*
  • Generous Paid Time Off: Take the time you need with vacation, personal, sick days, holidays, bereavement, and jury duty leave.*
  • Flexible and Affordable Health Benefits: Choose from three medical plans, and access optional dental, vision, Health Savings Account (HSA), and flexible spending account options to fit your lifestyle.*
  • 401(k) Retirement Savings Plan: Build your financial future with a company match (available to team members 18 and older).*
  • Employee Stock Purchase Plan: Accumulate funds through after-tax payroll deductions that can be used to purchase shares of BJ’s common stock at a 15% discount.*

*Eligibility requirements vary by position.

Position Overview:

The Principal Engineer, Major Incident Response & ITIL Platform Lead is a senior individual contributor and program leader who combines deep technical expertise with operational discipline. This role owns the design, configuration, and continuous evolution of the ITIL practice — including Major Incident Response (MIR), Post-Incident Review (PIR), and Problem Management — while serving as a hands-on engineer within ServiceNow and adjacent tooling platforms.

Unlike a traditional SDM role, this position is explicitly technical: you will architect workflows, build automation, instrument observability, and drive platform maturity across stores, distribution centers, and digital environments. You will also lead a high-performing offshore team and act as the primary program authority during high-severity events — bridging the gap between engineering execution and executive communication.

Key Responsibilities:

Major Incident Response (MIR) Program Leadership

  • Own and operate the MIR program end-to-end — from playbook authorship to real-time bridge command — for incidents impacting stores, DCs, POS, fuel, e-commerce, and membership systems.
  • Serve as Incident Commander during P1/P2 events, driving technical triage, stakeholder communication, and escalation decisions under pressure.
  • Design and maintain a universal MIR playbook with consistent execution standards 24x7, including on-call rotations for nights, weekends, and holidays.
  • Establish leadership notification templates, technical bridge protocols, and business-facing communication cadences during major incidents.
  • Instrument incident severity classification logic, auto-routing, and escalation thresholds directly within ServiceNow.

P ost-Incident Review & Postmortem Excellence

  • Own the end-to-end PIR lifecycle — blameless, data-driven reviews completed within SLA — and enforce action-item closure rigor.
  • Build and maintain an enterprise-wide RCA library, problem signatures, and trend intelligence within ServiceNow's CMDB and Problem Management modules.
  • Partner with SRE and Software Engineering to translate RCA findings into reliability-driven design improvements and automated runbooks.
  • Configure and manage PIR workflows, SLA timers, and notification rules natively in ServiceNow — no manual handoffs.

ITIL Platform Engineering & ServiceNow Ownership

  • Act as a hands-on technical owner of ServiceNow ITSM modules: Incident, Problem, Change, and Event Management.
  • Design and build ServiceNow workflows, business rules, UI policies, Flow Designer automations, and integration spokes connecting monitoring platforms (Dynatrace, Splunk, PagerDuty/AlertOps, etc.).
  • Develop and maintain custom dashboards, real-time KPI reporting, and SLA/SLO tracking within ServiceNow Performance Analytics.

Problem Management

  • Own the Problem Management lifecycle: identification, logging, root cause investigation, routing, and verified resolution.
  • Surface recurring incident patterns from trend analysis and feed intelligence back into MIR and Service Excellence programs.
  • Ensure complete, accurate, and timely documentation of all Problems in ServiceNow with appropriate categorization and linkage to incidents and changes.

Continuous Improvement & Team Leadership

  • Lead and develop a high-performing offshore operations team, setting clear goals aligned to MIR and ITIL program objectives.
  • Drive a culture of automation-first thinking: identify manual toil and eliminate it through ServiceNow scripting, Flow Designer, and third-party integrations.
  • Conduct regular retrospectives, process audits, and tooling reviews; translate findings into prioritized improvement backlog items.
  • Present program health, metrics, and roadmap updates to senior IT and business leadership.

Key Outcomes:

  • Faster stabilization of high-severity events through structured, technically informed incident command.
  • Measurable reduction in repeat incidents via high-quality, action-tracked RCAs.
  • A mature, automated ServiceNow platform that minimizes manual effort and accelerates response and reporting.
  • Predictable, trust-building communication to business stakeholders during and after major incidents.
  • Continuous improvement embedded into operational DNA — not a periodic exercise.

KPIs & Success Metrics:

Implement and manage service level agreements (SLAs and SLOs) to meet organizational goals and user expectations.

  • Mean Time to Acknowledge (MTTA) and Mean Time to Resolve (MTTR) for P1/P2 incidents.
  • Postmortem SLA compliance (e.g., 100% PIR completion within 5 business days).
  • Action item closure rate from PIRs within agreed timelines.
  • Reduction in repeat incidents (measured quarterly).
  • Problem Management throughput (number of problems logged, analyzed, and resolved).
  • Leadership communication SLA adherence during major incidents.
  • Continuous improvement initiatives delivered (e.g., automation, process optimization).
  • Stakeholder satisfaction scores from incident and problem management processes.

Requirements:

Education & Certifications

  • Bachelor's degree in Computer Science, Information Systems, or equivalent experience.
  • ITIL 4 Strategic Leader or Managing Professional certification strongly preferred; ITIL Expert acceptable.
  • ServiceNow certifications (CSA, CIS-ITSM, or CIS-Event Management) highly desirable.

Experience

  • 8+ years in IT Service Management with a strong technical bias — hands-on platform work, not just process governance.
  • 5+ years of direct experience administering or engineering ServiceNow (workflow design, scripting, integrations, Performance Analytics).
  • Proven track record leading Major Incident Response in large-scale retail, e-commerce, or distributed digital environments.
  • Experience owning Problem Management and PIR programs with measurable outcomes (repeat incident reduction, MTTR improvement).
  • Demonstrated ability to manage and develop onshore/offshore teams in a follow-the-sun operations model.

Work Environment:

  • Hybrid working model: 3 days onsite (Tue, Wed, Thu), 2 days remote (Mon, Fri).
  • Occasional travel to company locations or industry events.
  • Flexibility in working hours to accommodate global operations and time zone differences.
  • Participation in Major Incident on-call rotation.

Technical Skills

Deep ServiceNow platform expertise understanding platform mechanics that drive configuration and administratio

Original posting on BJ's Wholesale Club's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Principal Engineer – BJ's Wholesale Club | hirly.me