hirly

ContactMonkey

Applied AI Engineer

Toronto

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at ContactMonkey first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included
Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Mid level
Country
CA
Work mode
Remote-friendly
First seen by hirly
8 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Hey there! We're ContactMonkey 👋

Our mission? To power measurable employee engagement worldwide. And we'd love for you to join us!

About the job - Applied AI Engineer

Join the Engineering Team, where you'll help build our agentic future.

You'll work alongside senior engineers, Product, and our Chief Product & Technology Officer (CPTO) to design, prototype, and ship AI-powered capabilities quickly. This role is hands-on and iterative, focused on building production-grade agentic workflows that improve how internal communications are curated, designed, delivered, measured, and orchestrated.

This is an AI engineering role with real infrastructure ownership. You won't be handed a platform - you'll help build the one our AI features run on. If you like being close to both the model and the metal, this is that job.

Our stack, concretely: Ruby on Rails and Vue.js in a production SaaS codebase; Amazon Bedrock for model inference; AWS on EKS, provisioned with Terraform and Terragrunt across regions; Sidekiq for background work; MySQL and PostgreSQL. We're mid-migration to a GitOps deployment model with Argo CD and Karpenter. You'll touch most of this.

This is not "call an LLM and hope it works." It's careful system design, honest measurement, and shipping production systems that people rely on every working day.

This is a great fit if you've shipped LLM features in production, you're comfortable when the answer involves a Terraform plan rather than a prompt tweak, and you're ready to go deeper - with senior engineers around you to learn from.

Your impact

Agentic Design & Orchestration: Build agentic workflows and orchestration patterns (dynamic routing, tool-using agents, feedback loops) - contributing to the design and owning the implementation of well-scoped components.

Evals & Measurement: Own the eval harness for the features you ship. Build the datasets, the automated grading, and the offline regression suites that tell us whether a prompt or model change actually made things better. Raise the bar for evaluation across the team - judge design, offline and online metrics, and the judgment to know when a number is real. Turn production failures into evals that catch that class of failure next time.

Infrastructure for AI: Own work with SRE and build the infrastructure your features depend on, as code. Build the deployment path for AI workloads on EKS alongside our SRE and platform engineers.

Reliability, Cost & Safety: Keep our AI layer trustworthy as it grows. Instrument token spend, latency, and failure rates as first-class metrics. Design for the failure modes that matter - hallucination, timeout, rate limit, cost blowout - and make sure the system degrades gracefully instead of falling over. Respond when things behave unexpectedly in production.

Prompts & Model Configuration as Code: Treat prompts as versioned, reviewable, rollback-able artifacts rather than strings someone edited in a console. Own how we move a prompt or model change from idea to production safely, and how we back it out when it turns out worse.

Data Residency & Trust: We run regional deployments because our customers require it. Help make sure AI features respect those boundaries - that customer data stays in-region, model invocations are logged, and what we send to a model is what we intended to send. Think about prompt injection and PII exposure before an auditor does.

Execution & Experimentation: Take a defined problem and run with it. Contribute to our quarterly AI roadmap, run experiments to benchmark model value, and iterate systematically on prompts, retrieval, and model configurations based on what the data tells you.

Cross-Functional Partnership: Work closely with Product Management to turn product ideas into concrete technical work - both customer-facing features and internal workflow automation.

Grow With the Team: Participate actively in design discussions and code reviews. Ask questions early, escalate blockers, and share what you learn. Contribute to documentation practices (agents, markdown, Knowledge Bases) as we figure them out together.

Cultural Stewardship: Be a full participant in helping the engineering culture evolve as we grow.

About you

You hold a Bachelor's degree (or higher) in Computer Science, Statistics, Mathematics, or Engineering, or equivalent practical experience.

3+ years of professional software engineering experience, including hands-on work with AI/ML or LLM-powered features.

Strong fullstack experience - Ruby on Rails (or any other backend language/framework) and Vue.js/React (or any other framework) preferred. You're comfortable in a production SaaS codebase and can navigate unfamiliar systems without needing everything explained first.

You've written and applied infrastructure as code yourself. Terraform preferred - you've authored modules, read a plan carefully before applying it, and recovered from state that didn't match reality. Terragrunt, CloudFormation, or CDK experience transfers.

Working knowledge of AWS beyond the console: IAM (roles, policies, least privilege), VPC and networking basics, secrets management, and how a containerized workload gets deployed and observed. Kubernetes familiarity is a strong plus.

You've taken LLM work past a prototype at least once - you've thought about output quality, iterated on prompts against real usage, and dealt with the gap between "works in the notebook" and "works for customers."

You've built evaluation beyond trial and error - you can measure accuracy, safety, and latency, and you know when a benchmark win is real and when it's noise.

You understand how to structure LLM calls for reliability: using function calling and structured outputs to ground results, testing responses systematically, and designing for failure modes (hallucinations, latency, cost).

You can reason about what an AI feature costs to run - tokens, context size, retries, caching - and you treat that as an engineering constraint rather than someone else's problem.

You have familiarity with RAG architectures and vector stores, and an interest in multi-agent orchestration frameworks - you don't need to have run them at scale, but you should know the landscape.

You default to AI in your own workflow. You've already integrated modern AI tools into how you work, and you use them to move faster rather than as a crutch.

You prioritize a "question first, then answer" approach, thinking critically about problems before committing to a code path.

You bring intellectual humility, empathy, and strong listening skills to the team.

You are able to work collaboratively and independently.

You are comfortable with agile development.

You have excellent oral and written communication skills.

How you can stand out

You've built something agentic - an LLM reasoning across multiple steps, making decisions from tool outputs, adapting to feedback. Side projects count.

You've built an eval harness that changed a decision - where the numbers told you to ship, hold, or roll back, and you listened.

You've operated AI in production - you've been on the other side of a cost spike, a provider outage, a model deprecation, or a quality regression that only showed up under real traffic.

You have a bias toward simplicity - you're pleased when a well-structured baseline beats something more complex, and you know when not to reach for a model at all.

You've worked with Amazon Bedrock, or have opinions about model routing and provider abstraction earned the hard way.

You've worked within the email ecosystem or a similar high-scale communication platform.

You've built with AI coding agents and have a view on what they're good and bad at - we use them, and we're still figuring out the right practices.

You prototype fast, iterate quickly, and use data to change direction when the evidence says so.

You share what you learn with peers - writeups, de

Original posting on ContactMonkey's site ↗

Listed on hirly, a job board. hirly is not the employer: ContactMonkey is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job