Newxel
Senior ML Engineer – Speech & Voice (NXJ-212)
Poland (Hybrid/Entity)
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Data & ML
- Seniority
- Senior
- Country
- PL
- Work mode
- On-site / unstated
- First seen by hirly
- 7 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
The Role
GPU-based microservices handle the core speech pipeline—STT consumption, word-level alignment, diarization, and speaker identification. The primary challenge isn't just serving models, but establishing rigorous data-driven evaluation to separate real accuracy gains from benchmark noise on difficult sports audio. The Senior ML Engineer will own the speech services, lead fair model bake-offs, and hold full authority over which models reach production.
About the Product
The platform delivers real-time AI video processing and automated content generation for professional sports leagues globally. The underlying pipeline operates in high-noise live broadcast environments, demanding tight latency budgets, high throughput, and robust handling of overlapping speech and crowd noise.
Technology Stack: The platform delivers real-time AI video processing and automated content generation for professional sports leagues globally. The underlying pipeline operates in high-noise live broadcast environments, demanding tight latency budgets, high throughput, and robust handling of overlapping speech and crowd noise.
What You’ll Be Doing
Own and scale the core speech pipeline covering alignment, diarization, speaker identification, and enrollment signature-matching workflows.
Optimize GPU inference performance for throughput, memory footprint, and cost through batching, precision tuning, and compilation.
Build and maintain labeled sports test benchmarks stratified by speaker count, audio quality, language, and background noise.
Define and track production evaluation metrics including DER, WER, word-level speaker attribution, speaker-count error, latency, and GPU cost.
Conduct rigorous model bake-offs and write clear decision memos detailing trade-offs, licensing constraints, and statistical uncertainty.
Maintain a shadow or A/B deployment framework to gate every production release with empirical evidence.
What We Expect
Must-have
3–5+ years of experience shipping ML systems into production, with 2+ years dedicated to speech or audio pipelines.
Deep Python and PyTorch expertise, including hands-on GPU profiling, memory optimization, and inference acceleration.
Hands-on experience with at least two key speech domains: diarization (pyannote, NeMo/Sortformer, VBx), speaker embeddings (ECAPA-TDNN), or ASR and forced alignment (Whisper, Parakeet, wav2vec-family).
Strict evaluation rigor: demonstrated experience building test sets, computing DER/WER/EER correctly (handling collars and reference pitfalls), and running statistical hypothesis tests.
Demonstrated ability to critically analyze academic papers or model cards and reproduce published claims on internal datasets.
Hands-on experience deploying containerized GPU services in production environments using Docker and queue/API architectures.
Nice-to-have
Hands-on experience with Azure ecosystem tools (Service Bus, Blob Storage) and event-driven microservices.
Experience with domain adaptation or fine-tuning speech models on noisy, broadcast, or far-field sports audio.
Experience handling multilingual STT pipelines, particularly with Hebrew, Arabic, or Spanish.
Experience designing shadow deployment paths, experiment tracking workflows, and annotation team processes.
Why This Role Is Worth Your Time
Direct technical ownership of production GPU microservices where benchmark evaluations directly determine what reaches live users.
Applied research opportunity on challenging sports audio conditions (crowd noise, overlapping commentators, PA systems) rather than clean laboratory datasets.
Clear, evidence-based engineering culture where architectural and model choices are driven by data-backed decision memos rather than hype.
Listed on hirly, a job board. hirly is not the employer: Newxel is hiring for this role.
Similar jobs
- AI/ML Engineer - Junior/Mid/SeniorAccentureFirst seen today
- Senior ML Engineer - Speech (m/f/d)Voize · BerlinFirst seen todayremote
- Senior ML EngineerSolidgate · WarsawFirst seen todayremote
- Senior ML Engineer / Applied ML Engineer | NDAGt Hq · Warsaw, Poland First seen 12d ago
- Senior ML EngineerWlt · WarszawaFirst seen 31d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job