hirly

BMO

Principal Engineer, AI Platform

Toronto, ON, CAN

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at BMO first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Lead / management
Country
CA
Work mode
On-site / unstated
First seen by hirly
3 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Application Deadline:

10/29/2026

Address:

33 Dundas Street West

Job Family Group:

Technology

Principal Engineer, AI Platform & Fabrics

Description

BMO is building the platform capabilities that make enterprise AI safe, governed, and scalable. We are seeking experienced Principal/Senior engineers to build and operate the core infrastructure that governs how AI runs at BMO: the AI Gateway, Policy Engine, Identity Fabric, AI Registry, Guardrails Runtime, and AI Observability.

This is a build-and-run engineering role. You will own capabilities end to end; designing, shipping, and operating them in production, including on-call. You will not build the AI models or applications themselves (those are domain-owned); you build the governed platform they run on and the runtime evidence that proves they run within policy, across AWS, Azure, and Microsoft AI surfaces, under OSFI and OCC expectations.

You are a hands-on engineer who has built shared platform services at scale, cares deeply about operability, latency, and correctness, and understands that in a regulated bank the infrastructure must produce its own evidence . You are energized by taking real engineering assets that includes an existing developer portal, an AI registry, a body of policy-as-code, and gateway integrations, and hardening, scaling, and governing them into enterprise-grade platform capabilities. You raise the technical bar for those around you and mentor as you build.

What You'll Build & Operate

Depending on your specialization, you will own one or more of the following capability areas:

Enterprise Control Plane

Portal & Registry — a federated AI Registry (agents, models, tools, channels, evaluations across 16+ asset types) with self-service onboarding and lifecycle workflows; federation with external registries (Agent 365, AgentCore, MLflow).

Policy Engine — policy-as-code infrastructure (Cedar/OPA), a policy compilation pipeline, GitOps-based domain-scoped bundle distribution, risk-tiered approval workflows, and a policy simulation sandbox.

Observability & Audit — a multi-pipeline telemetry architecture (operational + security + compliance), OpenTelemetry GenAI conventions, cross-pipeline trace correlation, lineage-stamped traces, and a 7-year tamper-evident audit lake producing regulator-ready evidence.

Governance & Lifecycle — certification workflows, automated compliance scoring, decommission governance, and evidence generation for architecture and model-risk review.

Domain Orchestration

Gateway Runtime — domain-hub deployment across AWS and Azure; an inline enforcement engine performing request-time policy evaluation, routing, residency, budget/quota, and circuit breaking within strict tiered latency budgets (Fast Guardrails Runtime — a multi-stage safety pipeline (input moderation → prompt-injection defense → PII → output validation → hallucination detection → policy enforcement) with bilingual EN/FR parity and behavioral guardrails for agentic workloads (goal hijacking, intent drift, excessive agency).

Identity Fabric — workload identity for AI (SPIFFE/SPIRE), token-exchange bridging, per-domain trust boundaries, Entra Agent ID integration, on-behalf-of identity propagation, and cross-cloud token federation with zero-trust attestation.

What You'll Do

Own capabilities end to end — design, implement, test, ship, and operate production infrastructure, including on-call ownership of what you build (no separate run team).

Engineer for operability and defensibility from day one — instrumentation, SLOs, latency budgets, failure modes, and runtime evidence built in, not bolted on.

Build the APIs, SD'able interfaces, and integrations through which domains, DevOps pipelines, and enterprise systems consume platform capabilities.

Implement policy enforcement, guardrails, identity attestation, and audit as first-class engineering concerns — correct, performant, and provable.

Ensure every capability produces runtime evidence connecting AI activity to policy enforcement, identity, and lineage for model-risk and regulatory review (OSFI E-23, OCC).

Assess emerging AI infrastructure, foundation-model access patterns, and standards; make deliberate, cost-aware engineering choices.

Mentor and raise the bar — set engineering standards, review designs and code, and grow depth across the team.

Partner closely with AI Developer Experience (so domains can consume what you build), AI Security and AI SDLC (embedded specializations), and the Senior AI Architect (architectural coherence).

Education & Experience

Bachelor's degree in Computer Science, Software Engineering, or a related technical discipline (Master's preferred).

8+ years of software/platform engineering experience (Principal), or 5+ years (Senior), with substantial time building and operating shared platform services at enterprise scale.

Demonstrated experience operating production infrastructure with real SLOs and on-call ownership, ideally in a regulated industry (financial services strongly preferred).

Depth in one or more of: API gateways / traffic enforcement; policy-as-code and authorization; workload identity / zero-trust; observability and telemetry pipelines; audit/compliance data platforms.

Required Core Skills

Platform Engineering Depth

Strong distributed-systems and platform-engineering fundamentals: latency-sensitive request paths, resilience patterns (circuit breakers, failover), multi-tenancy, and high availability.

Strong programming skills (Python and/or Go preferred; TypeScript/Java an asset) for building performant services, APIs, and integrations.

Cloud-native architecture across AWS and Azure : containers/Kubernetes, service mesh, and Infrastructure as Code (CDK, Terraform, CloudFormation/ARM).

Robust CI/CD, GitOps, and DevSecOps practice; Git-based workflows (Bitbucket/GitHub), Jira, Confluence.

Capability-Specific Depth (one or more)

Policy/Authorization: Cedar, OPA/Rego, policy compilation and distribution, risk-tiered approval workflows.

Identity/Zero-Trust: SPIFFE/SPIRE, mTLS, token exchange, OAuth/OIDC, federated and cross-cloud identity, Entra ID/Agent ID, OBO.

Observability/Audit: OpenTelemetry (incl. GenAI conventions), distributed tracing, Dynatrace/Splunk or equivalents, tamper-evident/immutable audit stores, data lineage.

Gateway/Guardrails: API gateway internals, inline enforcement, LLM routing/abstraction, prompt-injection and PII defenses, hallucination detection, AI evaluation.

Registry/Portal: service catalogues, asset registries, lifecycle workflows, federation with external registries.

GenAI & Governance Context

Working knowledge of GenAI platform patterns: LLM/AI gateways, RAG and agentic patterns, foundation models, embeddings, and guardrails — sufficient to build the infrastructure they depend on.

Familiarity with AI/ML platforms (Bedrock, Azure OpenAI, SageMaker, MLflow) and orchestration frameworks (LangChain, LlamaIndex).

Grounding in Responsible AI, AI/data governance, privacy, cloud security, and IAM as applied to AI workloads.

Certifications (Preferred)

AWS Certified Solutions Architect (Associate/Professional) / ML – Specialty

Microsoft Certified: Azure Solutions Architect Expert / Azure AI Engineer Associate

Kubernetes (CKA/CKAD); HashiCorp Terraform Associate

Security/identity certifications (relevant to Identity Fabric roles)

Other Skills

Strong communication and collaboration across engineering, security, architecture, and domain teams.

A critical thinker with strong analytical and problem-solving skills.

Self-directed, comfortable with ambiguity and a fast-evolving mandate.

Able to deliver complex work under tight timelines; participates in on-call rotation for owned services.

Salary :

$105,000.00 - $215,000.00

Pay Type:

Salaried

The above represents BMO Financial Group’s pay range and type.

Salaries will vary based on factors su

Original posting on BMO's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job