TD
Senior Technology Resilience and Availability Management Analyst
Toronto, Ontario
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →hirly's read of this role
- Seniority
- Senior
- Stated salary
- $96,900 – $136,800 per year
- Country
- CA
- Work mode
- On-site / unstated
- First seen by hirly
- 30 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Work Location:
Toronto, Ontario, Canada
Hours:
37.5
Line of Business:
Technology Solutions
Pay Details:
$96,900 - $136,800 CAD
This role is eligible for a discretionary variable compensation award that considers business and individual performance.
TD is committed to providing fair and equitable compensation opportunities to all colleagues. Growth opportunities and skill development are defining features of the colleague experience at TD. Our compensation policies and practices have been designed to allow colleagues to progress through the salary range over time as they progress in their role. The base pay actually offered may vary based upon the candidate's skills and experience, job-related knowledge, geographic location, and other specific business and organizational needs.
As a candidate, you are encouraged to ask compensation related questions and have an open dialogue with your recruiter who can provide you more specific details for this role.
Job Description:
Role Summary
This first-line role supports the design, governance, assessment, and continuous improvement of Technology Resilience and Availability Management capabilities. The position works across application, infrastructure, cloud, cyber, risk, audit, and business teams to strengthen high availability, disaster recovery, backup and cyber recovery, capacity management, and operational resilience for critical technology services. The successful candidate combines strong technical depth with the ability to influence stakeholders, produce defensible evidence, and drive remediation in a complex regulated environment.
Job Responsibilities
- Lead resilience and availability assessments across applications, infrastructure, data, network, cloud, and third-party services to identify vulnerabilities, single points of failure, recovery gaps, and control weaknesses.
- Design, assess, and improve high-availability and recovery patterns, including active-active architectures, clustering, load balancing, multi-zone or multi-region deployment, automated failover, and resilient dependency design.
- Define, validate, and monitor resilience objectives and measures, including Recovery Time Objective (RTO), Recovery Point Objective (RPO), Maximum Tolerable Downtime (MTD), service-level objectives and indicators, availability targets, and capacity.
- Lead business and application impact analysis, dependency mapping, critical service mapping, and recovery prioritization to align technology capabilities with business resilience requirements.
- Plan and oversee high-availability and failover exercises, recovery-from-backup tests, tabletop scenarios, extended-duration testing, and appropriate failure-injection or chaos-testing practices.
- Assess backup, restore, and cyber-recovery capabilities, including immutable or isolated backups, point-in-time recovery, clean-room recovery, and ransomware recovery scenarios.
- Embed resilience and availability requirements into technology architecture, service operations, capacity management, change management, configuration management, incident and problem management, and the software development lifecycle.
- Use monitoring and observability data to identify availability, performance, capacity, and recovery risks, and translate findings into prioritized remediation actions and measurable improvements.
- Coordinate remediation activities, track risks and actions to closure, and provide clear reporting on capability maturity, test outcomes, control effectiveness, and residual risk.
- Produce organized, traceable, and defensible evidence for internal audit, regulatory examinations, senior management, and board or risk committee reporting.
- Serve as a trusted resilience advisor and central point of coordination across engineering, application owners, infrastructure, cyber, business continuity, technology risk, third-party risk, and operational resilience teams.
- Monitor emerging technology risks, regulatory expectations, cyber threats, and industry practices, and recommend practical enhancements to resilience standards, procedures, controls, and testing methods.
Job Requirements
Resilience, Availability and Recovery
- Demonstrated knowledge of high-availability design patterns, failover strategies, fault tolerance, redundancy, and elimination of single points of failure.
- Experience establishing and assessing RTO, RPO, MTD, availability objectives, SLOs, and recovery or availability metrics.
- Experience with business or application impact analysis, critical service mapping, technology dependency mapping, and recovery sequencing.
- Hands-on experience developing DR plans and runbooks and coordinating technical recovery exercises, failover tests, tabletop exercises, and end-to-end recovery validation.
- Knowledge of enterprise backup and restore, immutable or isolated backup, point-in-time recovery, and cyber-recovery concepts.
- Knowledge of capacity management, performance monitoring, utilization forecasting, and reporting
Platforms, Engineering and Tooling
- Experience with cloud resilience capabilities in AWS, Microsoft Azure, and/or Google Cloud, including multi-region architecture, traffic management, native backup, and disaster recovery services.
- Understanding of on-premises and hybrid technology, including VMware, SAN/NAS storage, Windows, Linux, Active Directory, DNS, and network dependencies.
- Working knowledge of container and Kubernetes high-availability patterns in cloud-native environments.
- Knowledge of database resilience methods for platforms such as Oracle, Microsoft SQL Server, and PostgreSQL, including replication, clustering, backup, restore, and point-in-time recovery.
- Experience using observability and monitoring platforms such as Splunk, Dynatrace, Datadog, or equivalent tools to assess availability, capacity, performance, and recovery outcomes.
- Experience with automation and infrastructure-as-code tools such as Python, or PowerShell.
- Experience with ServiceNow capabilities, including CMDB, incident, problem, change, and related technology risk or control workflows, is strongly preferred.
- Proficiency with Microsoft Word, Excel, and PowerPoint for analysis, evidence management, executive reporting, and program documentation.
Risk, Control and Regulatory
- Strong understanding of technology risk, operational resilience, disaster recovery, business continuity, and control assessment in a regulated environment.
- Working knowledge of relevant frameworks and guidance, including FFIEC Business Continuity Management expectations, OSFI/OCC and Federal Reserve operational-resilience guidance, NIST Cybersecurity Framework, and ITIL practices.
- Experience mapping technology applications and dependencies to important or critical business services.
- Knowledge of third-party resilience, including concentration risk, critical technology service provider testing, contingency planning, and exit strategies.
- Experience preparing evidence and written responses for internal audit, regulators, risk committees, and senior executives.
- Ability to apply major incident, problem, change, and post-incident review practices to improve resilience and reduce recurring disruption.
Competencies
- Advanced stakeholder management and relationship-building skills across engineering, application, infrastructure, cyber, risk, audit, and business teams.
- Clear and concise written and verbal communication, including executive-ready updates, root-cause analyses, postmortems, board or risk materials, and regulator-ready documentation.
- Ability to influence without direct authority and drive adoption of standards, remediation commitments, and sustainable process improvements across a matrix organization.
- Calm, structured leadership during incidents, recovery events, testing exercises, and periods of heightened scrutiny.
- Strong analytical problem-solving, including root-cause analysis, dependency analysis, scenario analysis, risk assessment, and
Similar jobs
- Technical Program Management Analyst, Launch Program 2027 - Toronto, CanadaMastercard · Toronto, CanadaFirst seen yesterday
- Model Risk Management Analyst, AVPMufgub · Toronto, ONFirst seen 3d ago
- Sr Materials Management AnalystAerospace · Mississauga, ON, CanadaFirst seen 3d ago
- Senior Software Asset Management AnalystSlihrms · CA.QC.Montréal.455 boul. René-Lévesque OuestFirst seen 6d ago
- Information Management AnalystIces · Toronto, OntarioFirst seen 11d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job