hirly

Udio

Senior Backend Engineer, Data Modeling and Ingestion Platform

New York

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Udio first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.6M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Engineering
Seniority
Senior
Country
US
Work mode
Remote-friendly
First seen by hirly
2 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About the Role

We are looking for a Senior Backend Engineer to lead the unification of large, highly rich, and heterogeneous datasets sourced from a wide range of external providers. These datasets are used to power our generative audio models.

Your work will create the foundational dataset that powers our research by building robust, scalable systems for linking, deduplicating, reconciling, and enriching data at massive scale. This role centers on high-impact bulk ingestion and advanced data linkage . You will design the logic, algorithms, and strategies that transform many independent datasets into a unified, high-quality canonical asset used throughout the company.

You will collaborate closely with ML researchers and product teams, working with tools such as BigQuery, Dataflow/Beam, TFRecords , and—where beneficial—distributed systems frameworks like Ray . Familiarity with ML workflows using JAX or multihost training is a plus, as the datasets you produce will directly support that ecosystem.

What You'll Do

Build high-throughput bulk ingestion workflows to integrate datasets from multiple external providers.

Design and implement scalable entity-resolution solutions, including record linking, deduplication, clustering, and conflict arbitration.

Create and refine matching logic, decision rules, and similarity functions to align datasets with high accuracy and strong coverage.

Define and track data quality indicators , such as overlap metrics, match precision/recall, duplicate rates, and completeness.

Prepare training-ready datasets in formats such as TFRecords , and structure data to meet ML research requirements.

Develop processing components using Dataflow (Beam) and manage large analytical workloads in BigQuery .

Leverage frameworks like Ray to accelerate large-scale experiments, feature extraction, and research-oriented data preparation.

Collaborate with ML researchers to anticipate downstream requirements and evolve linkage strategies as new sources and use cases emerge.

What We're Looking For

Experience working with large, heterogeneous datasets from multiple providers or domains.

Strong background in entity resolution , deduplication, data unification, or related large-scale data integration techniques.

Proficiency in Python , with an emphasis on efficient, scalable data processing.

Experience with BigQuery, Google Dataflow/Apache Beam , or similar batch-processing frameworks.

Familiarity with data validation, normalization, reconciliation , and building consistent views across diverse data sources.

Ability to craft well-structured matching and decision strategies that balance accuracy, completeness, and computational efficiency.

Comfortable iterating quickly on pragmatic solutions, balancing correctness with time-to-delivery.

Clear communication skills and the ability to collaborate closely with ML and research teams.

Nice to Have

Knowledge of architecting Google Cloud Platform systems at scale

Experience with distributed compute frameworks such as Ray , Spark , or Flink .

Understanding of JAX-based ML pipelines , multihost training setups, or large-scale data preparation for accelerator-backed workflows.

Familiarity with TFRecords or other high-volume training data formats.

Exposure to ranking, clustering, or statistical similarity modeling.

Experience with Go , NextJS , and/or React Native to contribute to full-stack development

Why Join Us

You will design the core dataset that underpins our research, product development, and generative audio models.

You'll work on large-scale data challenges that require creativity, algorithmic thinking, and engineering excellence.

You'll join a small, fast-moving team where your decisions shape the direction of our data and research capabilities.

Benefits

Highly competitive salary and equity

Quarterly productivity budget

Flexible time off

Fantastic office location in Manhattan

Productivity package, including ChatGPT Plus, Claude Code, and Copilot

Top notch private health, dental, and vision insurance for you and your dependents

401(k) plan options with employer matching

Concierge medical/primary care through One Medical and Rightway

Mental health support from Spring Health

Personalized life insurance, travel assistance, and many other perks

Udio’s success hinges on hiring great people and creating an environment where we can be happy, feel challenged, and do our best work.

Udio provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

This role is eligible for a compensation package of base salary, equity, and benefits. The starting base salary range for this role is $180,000 - $220,000. Actual salary may vary based on level, work experience, performance, and other factors evaluated during the hiring process.

Original posting on Udio's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job
Senior Backend Engineer – Udio · New York | hirly.me