Astrazeneca
Senior Cloud Platform Engineer (AWS), AI Infrastructure - Evinova
Canada - Mississauga
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Senior
- Country
- CA
- Work mode
- On-site / unstated
- First seen by hirly
- 2 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
- WHY JOIN US?
- Evinova is a health-tech business focused on accelerating better health outcomes by advancing digital transformation across the life sciences sector. By combining science-based expertise, evidence-led rigor, and deep human insight, we design digital solutions that enable healthcare to work better for everyone.
Operating at the intersection of healthcare, technology, data, and analytics, we are helping unlock the full potential of digital health, transforming how clinical research is conducted, how care is delivered, and how patients experience healthcare. Our solutions are built to scale, driving efficiency, improving decision-making, and ultimately delivering better outcomes for patients worldwide.
At Evinova, we are driven by a shared purpose to transform health through data and digital innovation. Our teams collaborate across disciplines to solve complex challenges, continuously learning and evolving in a fast-paced, high-impact environment.
We also recognize the importance of flexibility and balance. Our ways of working support both individual needs and team collaboration. To foster connection and collaboration, employees are expected to work from the office three days per week , creating opportunities for in-person teamwork, innovation, and meaningful connection.
This role is located in the Greater Toronto Area and follows a hybrid work model. Candidates must reside within commuting distance of the GTA or be willing to relocate for this opportunity.
Introduction to Role
The Machine Learning and Artificial Intelligence Operations team (ML/AI Ops) is the cloud platform engineering team responsible for building and operating the infrastructure that enables our AI Engineers and Data Scientists to deploy Generative AI applications reliably, securely, efficiently and at scale.
As a Senior Cloud Platform Engineer on the ML/AI Ops team, you will design, build and operate the AWS platform that runs our production Generative AI, agentic AI and conversational AI workloads. Most of the code you write will be AWS CDK (TypeScript and/or Python) that creates the infrastructure for other teams to build on.
This is a cloud and platform engineering role rather than an AI application development role. You will partner closely with the engineers and Data Scientists who build agents and models, and provide the infrastructure, deployment patterns, model access, observability and operational capabilities they need to move solutions from experimentation into reliable production environments.
You will work across AWS infrastructure, infrastructure as code (IaC), Amazon Bedrock AgentCore, Amazon ECS, CI/CD, SageMaker Unified Studio, AI gateways, observability, scalability, reliability, security, governance and cost optimization. Your work will establish reusable platform capabilities that allow teams across Evinova to deploy and operate solutions faster and more reliably while meeting the requirements of a highly regulated pharmaceutical environment.
Accountabilities
Cloud Platform Engineering
- Design, build and operate scalable AWS cloud platform capabilities for production ML/AI and Generative AI workloads.
- Create reusable infrastructure, tooling and deployment patterns that enable AI Engineers and Data Scientists to independently deploy and operate their applications.
- Write and maintain AWS IaC primarily AWS CDK in TypeScript and/or Python, including reusable CDK constructs that other teams consume.
- Build and operate containerized workloads using Amazon ECS and AgentCore.
- Develop reusable platform capabilities across compute, networking, IAM, secrets management, storage, model access and workload isolation.
- Build and maintain CI/CD and GitOps workflows that enable safe, automated and repeatable deployments across environments.
- Partner with engineers and Data Scientists to transition prototypes and research workloads into resilient, production-grade services.
- Build self-service capabilities and automation that improve developer experience and reduce operational toil.
Reliability, Scalability & Operational Excellence
- Engineer platform capabilities that improve the availability, scalability, resiliency and performance of production GenAI workloads.
- Design and implement autoscaling, load balancing, retries, timeouts, fallback, rate limiting and failure-recovery strategies.
- Establish monitoring, alerting, SLOs, runbooks and production-readiness standards for ML/AI workloads.
- Troubleshoot complex production issues across AWS infrastructure, container, application and model-provider layers.
- Automate operational processes and proactively identify opportunities to improve platform reliability and performance.
- Drive cloud and model cost optimization through capacity management, workload optimization and data-driven analysis.
AI Gateway & Model Access
- Build and operate a centralized AI gateway that routes requests across model providers.
- Provide secure, reliable and governed access to foundation models through Amazon Bedrock, OpenAI, Anthropic, Microsoft Foundry, Gemini Enterprise Agent Platform (formerly Vertex AI) and other model platforms.
- Implement model/provider routing, fallback, authentication, rate limiting, quotas and cost controls.
- Enable AI teams to evaluate and change model providers without tightly coupling their applications to individual model endpoints.
- Provide the compute, networking, storage and runtime infrastructure to reliably operate agentic AI, RAG and conversational AI workloads in production.
Observability, Governance & Cost Optimization
- Build platform-level observability for GenAI workloads, including token consumption, latency, throughput, errors, model/provider performance and cost attribution by team and application.
- Implement standardized tracing, logging, metrics and alerting capabilities that can be adopted across AI applications.
- Integrate observability technologies such as Amazon CloudWatch, OpenTelemetry, Datadog and Splunk.
- Build the telemetry and data pipelines that AI teams use to evaluate and monitor production LLM behavior.
- Build appropriate security, auditability and governance controls for AI workloads operating within a regulated environment.
- Support compliance with applicable industry standards and practices, including Good Clinical Practice and Good Machine Learning Practice and other GxP related standards and practices.
Representative Projects
- Build a library of AWS CDK constructs that gives an AI team a production-ready Amazon ECS service, with IAM, secrets, networking and observability, from a single import.
- Deploy and operate a centralized AI gateway on Amazon ECS, with multi-provider routing, fallback, quotas and rate limiting.
- Build multi-account CI/CD with CDK Pipelines or GitOps, with safe, staged deployments across environments.
- Implement token cost attribution by team and application with SLO and burn-rate alerts.
- Migrate container workloads from Amazon EKS to Amazon ECS without loss of reliability.
Essential Skills/Experience
- Minimum of 4 years of hands-on experience in a platform engineering, infrastructure, site reliability engineering (SRE) or DevOps role where infrastructure code was your main output.
- Deep hands-on AWS cloud engineering experience, including designing, deploying and operating production cloud-native infrastructure.
- Strong experience with AWS services such as Bedrock, ECS, IAM, VPC, load balancing, S3, CloudWatch, Secrets Manager and related services.
- Strong hands-on experience with IaC. AWS CDK using Python and/or TypeScript is preferred. Strong Terraform engineers who are willing to work in CDK are welcome.
- Strong experience with Docker and Amazon ECS, including workload deployment, autoscaling, resource management, health checks, networking, security and production troubleshooting.
- Experience designing and maintaining CI/CD and/or GitOps pipelines for production cloud workloads.
- Strong underst
Similar jobs
- Senior Platform EngineerManulife · 2 LocationsFirst seen 3d ago
- Senior AWS Platform EngineerFreeBalance · Ottawa, Ontario, CanadaFirst seen 4d ago
- Senior Platform Engineer, Quality Platform MaintainX (An Autodesk Company) · TorontoFirst seen 4d agoremote
- Senior Platform EngineerScalar · RemoteFirst seen 4d agoremote
- Senior DevOps / Platform Engineer — CloudDominion Dynamics · Ottawa or TorontoFirst seen 4d ago
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job