Svb
Principal Infrastructure Engineer - Major Incident Manager
Bangalore, India
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Lead / management
- Country
- IN
- Work mode
- On-site / unstated
- First seen by hirly
- 27 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
FC Global Services India LLP (First Citizens India), a part of First Citizens BancShares, Inc., a top 20 U.S. financial institution, is a global capability center (GCC) based in Bengaluru. Our India-based teams benefit from the company’s over 125-year legacy of strength and stability. First Citizens India is responsible for delivering value and managing risks for our lines of business. We are particularly proud of our strong, relationship-driven culture and our long-term approach, which are deeply ingrained in our talented workforce. This is evident across all key areas of our operations, including Technology, Enterprise Operations, Finance, Cybersecurity, Risk Management, and Credit Administration. We are seeking talented individuals to join us in our mission of providing solutions fit for our clients’ greatest ambitions.
Job Description:
Value Proposition
Responsible for enhancing the reliability, resilience, and stability of enterprise IT services through effective Major Incident Management and rapid service restoration.
Leading the command, coordination, and technical response for critical production incidents impacting business-critical applications, infrastructure, cloud platforms, and customer-facing services.
Acting as the central Incident Commander during high-severity incidents, driving technical triage, root cause identification, resolution, and minimizing business impact and operational risk.
Partnering with Engineering, SRE, Infrastructure, Cybersecurity, and Application teams to strengthen platform reliability, operational resilience, monitoring, automation, and service maturity.
Job Details
Position Title: Principal infrastructure Engineer – Major Incident Manager
Career Level: P4
Job Category: Assistant Vice President
Role Type: Hybrid
Job Location: Bangalore
About the Team:
The Major Incident Management Team serves as the central command function for critical incidents across global operations, providing 24x7 coverage through teams in India and the US.
Focused on rapid recovery and business continuity, the team leads coordinated incident response, promotes industry best practices, drives continuous improvement, and strengthens operational resilience to support a world-class financial institution.
Key Deliverables (Duties and Responsibilities)
Major Incident Command & Technical Leadership
Serve as the Incident Commander for major incidents across enterprise applications, cloud platforms, infrastructure, databases, and network services.
Lead end-to-end major incident management activities including incident assessment, prioritization, escalation, responder engagement, technical coordination, recovery execution, and service restoration.
Direct technical bridge calls and facilitate collaboration among Infrastructure, Cloud, Network, Database, Middleware, Application Development, Cybersecurity, Vendor, and SRE teams.
Drive structured technical triage and ensure investigation efforts remain focused on business recovery and root cause isolation.
Challenge incomplete technical updates, validate remediation approaches, and remove operational blockers during critical situations.
Make informed decisions under pressure while maintaining clear ownership, accountability, and resolution momentum.
Ensure incidents are managed in accordance with established ITSM, operational risk, and regulatory requirements.
Service Reliability & SRE Partnership
Partner with Site Reliability Engineering (SRE), Production Support, and Engineering teams to improve service reliability and operational resilience.
Promote adoption of reliability engineering practices including:
Service Level Indicators (SLIs)
Service Level Objectives (SLOs)
Service Level Agreements (SLAs)
Error Budgets
Observability
Proactive Monitoring
Event Correlation
Capacity Planning
Operational Readiness Reviews
Identify recurring incidents, systemic weaknesses, and reliability risks across technology platforms.
Contribute to the continuous improvement of operational stability through automation, monitoring enhancements, and process optimization.
Support operational readiness for major releases, infrastructure transformations, and cloud migrations.
Incident Analysis & Continuous Improvement
Lead Post-Incident Reviews (PIRs) and Root Cause Analysis (RCA) activities.
Partner with Problem Management teams to identify corrective and preventive actions.
Ensure action items are tracked through closure and measurable service improvements are achieved.
Drive improvements to runbooks, knowledge articles, operational procedures, escalation paths, and response frameworks.
Identify opportunities to reduce Mean Time to Detect (MTTD), Mean Time to Restore (MTTR), and recurring service disruptions.
Reporting & Operational Insights
Produce executive-level incident communications and reporting.
Develop and publish operational dashboards and metrics including:
MTTD
MTTA
MTTR
Incident Volumes
Availability Metrics
Escalation Trends
SLA Compliance
Recurring Incident Analysis
Provide actionable insights to technology leadership regarding operational health and service reliability.
Stakeholder Management & Communication
Provide timely, accurate, and concise communication to executive leadership, business stakeholders, and technical teams during major incidents.
Translate complex technical issues into business impact language suitable for senior leadership.
Lead stakeholder communications throughout the incident lifecycle.
Ensure high-quality incident documentation, timelines, and executive summaries.
Governance, Risk & Compliance
Ensure adherence to ITIL, Incident Management, Problem Management, Change Management, and Operational Risk frameworks.
Support audit, compliance, regulatory, and risk management requirements.
Maintain incident governance standards, documentation, playbooks, and escalation procedures.
Participate in regulatory and operational resilience initiatives within the organization.
Skills and Qualification:
Qualifications Required
Bachelor’s degree in Computer Science, Information Technology, Engineering, or related discipline and 12+ years of experience in Technology Operations, Production Support, Incident Management, Site Reliability Engineering, Infrastructure & application support Operations or Enterprise Technology Services.
Required Experience
12-15 years of experience supporting enterprise-scale technology environments.
Hands-on Major Incident Management, Incident Command, Production Operations, SRE or Operations Engineering experience.
Proven experience leading Severity 1 / Critical production incidents within large-scale banking enterprise environments.
Demonstrated experience coordinating technical teams across infrastructure, cloud, applications, databases, networking, cybersecurity and vendor organizations.
Experience within Banking, Insurance or other regulated Financial Services environments is mandatory.
Strong experience working within ITIL-aligned Incident, Problem, and Change Management frameworks.
Experience driving operational improvements, service reliability initiatives, and incident reduction programs.
Proven ability to influence technical teams and senior stakeholders during high-pressure situations.
Technical Proficiency
The successful candidate should possess a good technical foundation and the ability to effectively engage engineering teams during incident diagnosis and recovery efforts.
Infrastructure & Platform Knowledge
Hands-on expertise and understanding of:
Network Infrastructure
Windows and Virtualization Platforms
Linux & AIX Server Platforms
Backup & Storage Technologies
Data Centers
Databases
Middleware Technologies
API Integrations
Cloud & Modern Platforms
Good working knowledge of:
Cloud Technologies (AWS/Azure/GCP)
Container Platforms (Docker/Kuberne
Similar jobs
- Cloud Infrastructure EngineerCisco · Bangalore, Karnataka, IndiaFirst seen today
- Lead Infrastructure Engineer- NetworkJPMorgan Chase · Hyderabad, Telangana, IndiaFirst seen today
- Lead Infrastructure Engineer - NetworksJPMorgan Chase · Bengaluru, Karnataka, IndiaFirst seen yesterday
- Lead Infrastructure Engineer- NetworkJPMorganChase · Hyderabad, Telangana, IndiaFirst seen 2d ago
- Azure Infrastructure EngineerInfosys · Hyderabad, IndiaFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job