hirly

Hamming AI

Staff Backend Engineer

Austin

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Hamming AI first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Role family
Engineering
Seniority
Lead / management
Country
US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

Location: Remote (North America) or Austin, TX

Employment Type: Full-time (no contractors)

Department: Engineering

Why now

Hamming builds three products for voice and chat AI agents: testing/simulation to validate behavior before launch; red-teaming to probe for prompt injection, jailbreaks, PII leakage, and policy violations; and production monitoring/observability to detect failures in live conversations and turn them into regression tests.

We are one of the fastest engineering teams in the world. We prod deploy 4x / day.

I’m looking for someone who can own reliability and scale across our LLM-enabled platform, shipping precise, outcome-driven improvements to high-availability systems.

— Sumanyu (CEO)

Previously: grew Citizen 4× and scaled an AI sales program to $100Ms/yr at Tesla.

Devin Case Study

Ranked #1 Eng team

OpenAI Dev Day 100billion token list

What you’ll do

Own core services in TypeScript/Node.js and Python that orchestrate LiveKit , Temporal , STT/TTS, and LLM tooling for real-time voice agents.

Scale 1 → N → 100× : take what works today and harden it for 10K parallel calls with 99.99% uptime. Turn human playbooks into productized systems.

Harden pipelines for ingestion, evaluation, and analytics so telephony events, recordings, and outcomes propagate reliably across services.

Level-up observability : deepen OpenTelemetry/SigNoz and trace-first practices to shrink mean-time-to-truth in prod.

Prototype → test → prod : partner with product to ship new LLM-driven behaviors with clear success metrics, guardrails, and regressions blocked in CI.

Infrastructure readiness : CI/CD, environment automation, incident response playbooks—customer conversations stay online.

You might be a fit if you

Have senior/staff experience running distributed backends with real-time/streaming constraints.

Are fluent in TypeScript/Node.js and comfortable jumping into Python for ML/audio jobs.

Know Temporal (or similar workflow engines), queues, Redis, and PostgreSQL .

Have shipped production LLM apps and understand prompt/tool design, evals, and guardrail instrumentation.

Operate cloud-native on AWS with Terraform ; k8s doesn’t scare you.

Are a power user of Cursor/Zed/Devin and were using code-gen before it was cool.

Have intuition for what current-gen LLMs can/can’t do—and what tomorrow’s models will unlock.

Think independently, grind with customers , and do whatever it takes—without dropping the quality bar.

Bonus: built 0→1 real-time systems in Telecom/Networking, Autonomous Vehicles, or HFT; founded something; built AI voice apps.

Interesting problems you’ll touch

Voice simulations that feel real : accents, overlapping speech, crosstalk, background noise, barge-ins.

Massive concurrency : 10,000+ parallel calls with deterministic behavior and graceful degradation.

Temporal-driven orchestration for long-running, interruptible call flows.

Closed-loop reliability : turn prod failures into auto-generated tests and blocked deploys.

Trace-everything culture: make “what happened?” a 30-second question, not a war room.

How we work

Outcomes over output : we adjust roadmaps when new data lands.

Demo early and document decisions so context moves fast.

Own incidents : lead the investigation, write crisp notes, land durable fixes.

Direct, candid, respectful communication keeps remote teammates in lockstep with Austin HQ.

Our stack

App : Next.js, TypeScript, Tailwind

AI : OpenAI, Anthropic, STT/TTS providers

Realtime/Orchestration : LiveKit, Pipecat/Daily, Temporal

Infra/DB : AWS, k8s, PostgreSQL, Redis, Terraform

Observability : OpenTelemetry, SigNoz

Apply

If you want to make AI voice agents reliable at scale , let’s talk.

Send a short note (links to work > resumes) and tell us about something reliability-critical you shipped: what broke, what you fixed, and how you knew it worked.

Original posting on Hamming AI's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job