hirly

DataSnipper

Senior Software Engineer - LLM Ops & Evals

Amsterdam

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at DataSnipper first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Role family
Engineering
Seniority
Senior
Country
NL
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Every AI call in DataSnipper goes through us. Product teams do not talk to model providers directly, they talk to our gateway. We are responsible for how inference is routed, how it fails over, what it costs, and how anyone can tell whether the output is any good.

The second half of the job is evaluation. We are building the platform teams use to measure AI quality: versioned datasets, experiment tracking, and evaluation runs they can act on. It is a hard problem and largely an open one, so you will have real influence over how we solve it.

This is a small team with a large blast radius. You will own real production systems, set the standards other teams build against, and see your work in front of hundreds of thousands of users in audit and finance.

About DataSnipper

  • Audit and finance are still massively manual and we are changing that. DataSnipper is a $1B, bootstrapped unicorn with 600,000+ users across 180+ countries, already embedded in the daily workflows of top audit and accounting firms.
  • Now, we are taking things further with our Excel Agent, bringing AI directly into where the work actually happens. Unlike generic AI tools, we do not sit on the sidelines. Our AI operates inside Excel, with access to real documents and audit evidence, meaning it does not just generate answers, it does the work, with full traceability.
  • We are not just applying AI, we are redefining how audit gets done. If you want to build something category-defining at scale, this is the place.

What you will do

Technical Delivery

Own the LLM gateway: routing, provider failover, rate limits, retries, and cost attribution across multiple model providers

Deploy, version and deprecate models across clouds, regions and environments, including quota and capacity planning, managed as infrastructure as code

Build the shared evaluation platform: versioned datasets, experiment tracking, run and result schemas, reporting, and trace linkage back to the run

Own the infrastructure for async and long-running AI workloads

Reliability, Security & On-Call

Own observability for AI traffic: latency, retries and fallbacks, token usage, cost and errors, per team and per use case

Take part in the on-call rotation, run incidents, and close the follow-ups

Implement the security and compliance controls the platform is held to: retention, access control, RBAC and SSO

Collaboration & Impact

Define and maintain clean integration contracts between the platform and the teams that consume it

Partner with product and ML engineers to turn their requirements into platform capabilities that are self-service rather than a request queue

What you will bring

Must-Have

5+ years in backend or platform engineering, with strong production Python

Experience building or running LLM inference infrastructure: a gateway or routing layer with multiple providers, failover, rate limiting and cost attribution

Experience with cloud at the infrastructure level and infrastructure as code ( we run across Azure and GCP with Terraform )

Experience running a shared service in production: on-call, incidents, postmortems, SLOs

Hands-on experience with observability tools (OpenTelemetry, Grafana), including instrumenting services and designing dashboards and alerts

Comfort with privacy and compliance work: PII handling, anonymisation, retention, access control

Experience building platform or shared-service capabilities consumed by multiple internal teams

Nice-to-Have

Experience with LLM or agent evaluation

Temporal or another durable workflow engine

Self-hosted inference, capacity planning, load testing

Synthetic data or document anonymisation pipelines

Document AI: VLMs, OCR, structured extraction and the metrics that go with it

Domain experience in audit, accounting or fintech

What we expect

Ownership : You own work end-to-end, anticipate issues, and ensure high-quality delivery without close supervision

Growth Mindset : You encourage open feedback exchange and provide clear, balanced feedback that helps others grow

Collaboration : You build strong cross-functional relationships and influence peers through expertise, data, and empathy

Adaptability : You navigate ambiguity calmly, model positive behavior, and help peers adjust through clear communication

Judgment : You exercise sound judgment in ambiguous situations, balance speed and accuracy, and adjust priorities proactively

Recruitment steps

Recruiter screen

Hiring Manager interview

Peer programming session

System design interview

Final interviews with Engineering leadership

Original posting on DataSnipper's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job