hirly

Clera

Senior Voice AI Engineer

remote

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Clera first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Senior
Stated salary
$96,000 per year
Work mode
Remote-friendly
First seen by hirly
2 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About the Role

As a founding engineer on a small conversational AI team, you will own the real-time voice layer, from incoming speech through AI reasoning to spoken responses. You will help make natural, responsive voice interactions work reliably in production, with a focus on end-to-end latency.

What You'll Do

Build and own streaming speech-to-text, LLM turn-taking, text-to-speech, and telephony or WebRTC transport.

Measure and reduce latency, targeting first audio under 800 milliseconds on real calls.

Address interruptions, barge-in, silence detection, overlapping speech, poor audio, accents, and mid-sentence changes.

Build an evaluation harness from recorded calls, transcripts, and scored turns to detect regressions and guide product decisions.

Compare voice providers and models through evidence-based testing, and make changes based on results.

Instrument production systems for turn latency, transcription confidence, drop-offs, and cost per minute.

Work directly with founders and make technical decisions in a fast-moving team.

What We're Looking For

At least 5 years building production software, including 2 or more years shipping voice, speech, or real-time audio systems.

Experience building and shipping end-to-end real-time voice pipelines, including streaming speech recognition, LLM turn-taking, speech synthesis, and telephony or WebRTC.

Strong Python or TypeScript skills and comfort working in both.

Hands-on experience with an audio stack such as LiveKit, Pipecat, Vapi, Twilio Media Streams, Daily, or a custom WebSocket implementation.

Experience debugging audio at the frame level, including sample rates, codecs, jitter, and voice activity detection thresholds.

Experience building LLM evaluation harnesses, optimizing latency against real-world targets, and using evaluation results to make product decisions.

Clear written English for asynchronous communication. Experience with speech model serving or fine-tuning, SIP, telephony, or LLM orchestration frameworks is a plus.

Compensation & Benefits

Compensation is $96,000 USD annually, regardless of location. Visa sponsorship is not available.

Location

Fully remote, anywhere in the world. Core team overlap is 13:00 to 17:00 UTC.

Original posting on Clera's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job