hirly

Luma

Software Engineer - Data Infrastructure

Redwood City, CA

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Luma first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.3M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

hirly's read of this role

Role family
Engineering
Seniority
Mid level
Stated salary
$170,000 – $360,000 per year
Country
US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About the Role

As a Data Infrastructure Engineer in Research at Luma, you will play a critical role in building and scaling the data infrastructure that supports our cutting-edge multimodal AI systems. Your work will focus on developing high-throughput, large-scale data processing pipelines tailored for machine learning research and internal ML platform needs. You will collaborate closely with ML researchers and product teams to create reliable, efficient, and easy-to-use data infrastructure that empowers innovation and accelerates development. This role requires a strong foundation in distributed systems and data engineering, with an emphasis on supporting complex machine learning workflows rather than traditional product data infrastructure.

Responsibilities

Build and maintain scalable data infrastructure for high-throughput machine learning workflows

Collaborate with ML researchers and product teams to ensure data systems meet evolving needs

Develop and optimize large-scale data pipelines and batch processing jobs

Contribute to the architecture and implementation of reliable, high-performance data platforms

Integrate open-source tools and continuously improve data infrastructure through monitoring and tuning

Participate in cross-functional projects to improve data reliability, scalability, and operational excellence

Support the evaluation and adoption of new programming languages and frameworks relevant to data infrastructure

Engage in continuous improvement of data infrastructure through monitoring, troubleshooting, and performance tuning

Collaborate with research & engineering teams to help define and refine best practices for data infrastructure development

Qualifications

Proficiency in Python (or similar languages with willingness to learn Python) and experience with large-scale, high-throughput data infrastructure

Familiarity with distributed computing frameworks (e.g., Ray, Spark, Beam)

Ability to design and optimize data pipelines for ML research and internal teams

Strong problem-solving skills and understanding of data engineering at scale

Collaborative, product-focused mindset; comfortable in fast-paced environments

Experience sourcing, integrating, and optimizing data from diverse and large datasets

Comfortable working in a fast-paced, product-focused environment with a strong execution mindset

Open to candidates across seniority levels, from mid-level individual contributors to senior engineers and managers.

Nice to have

Prior experience working with complex data infrastructure or AI/ML platforms highly desirable

Experience with open source data infrastructure projects is a plus

Experience working in the robotics industry preferred

Original posting on Luma's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job