hirly

Appnovation

Site Reliability Engineer, AI Observability

Toronto, Canada · New York, New York, United States

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Appnovation first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included
Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Mid level
Countries
CA, US
Work mode
On-site / unstated
First seen by hirly
6 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About us

Appnovation is the AI Product Company. Headquartered in Vancouver, Canada, with teams around the world, Appnovation builds AI products, delivers enterprise AI and digital solutions, and co-invests with clients to bring new AI products to market.

We’re looking for a Site Reliability Engineer to keep a shared observability platform for LLM-based applications running for a global life sciences client. The platform is built on Langfuse and self-hosted on Kubernetes on AWS, with ClickHouse as the analytical store, PostgreSQL for metadata and Redis as the ingestion queue, all delivered through Argo CD.

Two things need building rather than maintaining. The platform has no monitoring, alerting or defined service levels today, and you will own putting them in place. Infrastructure is defined as code throughout, in Kubernetes manifests, Helm values and Argo CD applications.

Alongside the platform itself, you will own the runbooks that make an incident survivable by someone other than the author, and the onboarding and support path for the internal teams that depend on the platform.

ROLE RESPONSIBILITIES

Platform Operations: Diagnose and resolve failures across ClickHouse, PostgreSQL, Redis, the ingestion workers and the Kubernetes layer beneath them, including ingest backpressure from queue depth and worker drain behaviour.

ClickHouse Operations: Own ClickHouse under the operator model, including Keeper quorum, replication, shard and replica topology, and S3 storage tiering.

Provable Backup and Restore: Rehearse the restore, time it, document it and test it against its failure modes, rather than assuming a successful backup job means a recoverable system.

Safe Upgrades: Plan and rehearse upgrades in a lower environment, with a rollback plan that works even if a migration is only partly complete.

Monitoring and Service Levels: Build monitoring, alerting and service levels from scratch, so problems are found here before a user reports them.

Infrastructure as Code: Maintain Kubernetes manifests, Helm values and Argo CD applications so every change goes through the delivery pipeline.

Runbooks and SOPs: Write and maintain runbooks and SOPs that let a colleague resolve an incident without the author present.

Onboarding and Support: Run the onboarding and support path for internal teams that depend on the platform, and triage what they bring.

Automation: Turn recurring operational work into automation.

Upgrade Partnership: Work with the platform engineer who owns what the platform offers. They decide what to adopt and how it is configured; you own the migration and its rollback.

QUALIFICATIONS

Hands-on experience with Kubernetes on AWS (managed EKS), with routine work done through Helm values and Argo CD applications.

Hands-on experience running ClickHouse in production, including replication and Keeper quorum, shard and replica topology, and backup and restore, ideally run through an operator.

Experience building monitoring and alerting from scratch, including service levels that reflect what users actually experience rather than what is easy to measure.

Experience upgrading self-hosted software safely, including schema migrations, rehearsal in a lower environment and a rollback plan for a partly completed migration.

PostgreSQL and Redis operations deep enough to debug metadata-store and queue problems, including backpressure and worker drain.

Strong operational writing: runbooks, SOPs and post-incident reviews that a colleague can follow unaided during an incident.

PREFERRED QUALIFICATIONS

Observability engineering, including OpenTelemetry Collector pipelines, alerting design and Grafana dashboards.

OIDC or enterprise SSO integration with a corporate identity provider.

GitHub Actions for plan and apply pipelines with approval gates.

Experience running LLM observability tools such as Langfuse, LangSmith or Arize Phoenix.

Experience in pharma, life sciences or another regulated industry.

WHO YOU ARE

You don’t trust a backup until you’ve restored from it

You want to find problems before users do

You write runbooks for the person on call at 3am, not for yourself

You stay calm in incidents and focus on fixing the process afterward

You automate anything you have to do twice

You work well inside a client team and build trust quickly

You have prior experience in consulting

Prior experience and connections in the Life Sciences industry is preferred

Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.

At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how we stay true to our company values as we build better lives through better digital.

Accommodations are available upon request throughout the recruitment process.

Original posting on Appnovation's site ↗

Listed on hirly, a job board. hirly is not the employer: Appnovation is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job