hirly

Avra

Member of Technical Staff | ML Systems

São Paulo

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Avra first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Lead / management
Country
BR
Work mode
Remote-friendly
First seen by hirly
23 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About the role

At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area.

In this role, you'll join our ML Systems team, which owns Avra's ML core and the governance of every model we ship. Research produces candidate models and evidence; you build the reliable path from data and training to a governed, reproducible release that can run in our cloud or in any customer environment. ML Systems is an internal platform: its users are our researchers and platform engineers, and its success is measured by the leverage it creates for them.

What you'll do

Build CUDA kernels and compute primitives for training and serving graph neural networks (GNNs).

Evolve Monad, our sampler and distributed-training library, including neighbor sampling and training performance.

Specify our binary data formats (Lance, Arrow, CSR/CSC), and own materializations and feature backfills for training and evaluation.

Define data contracts and consumption requirements with the teams that build our customer and proprietary datasets.

Build and operate experiment tracking, checkpoints, and evaluation infrastructure, with reproducibility by default.

Own the model registry, lineage, versioning, and compatibility across models, embeddings, and downstream models.

Define and run release gates, so every model running in production, batch, or on-premise maps to a governed release.

Make it possible to audit exactly which data, code, configuration, and evidence produced each release.

How we measure success

Time-to-experiment: how quickly a researcher goes from a hypothesis to materialized data, compute, and tracking.

Time-to-governed-release: how quickly a validated candidate becomes an authorized release.

Training throughput per GPU on our foundation model training runs.

100% of production models with complete release records and lineage — no ad hoc models in any environment.

Every release reproducible from its registered data, code, and configuration.

What we're looking for

Strong systems engineering skills and production-quality Python.

Experience with distributed training (e.g., Ray, PyTorch distributed) and multi-node GPU workloads.

Experience with columnar data formats and large-scale data materialization.

Familiarity with ML lifecycle tooling: experiment tracking, model registries, evaluation, and reproducibility.

A product mindset: you treat an internal platform as a product with real users. You don't need to be a data scientist.

Nice to have

CUDA kernel development or GPU performance optimization.

Graph neural networks or graph sampling at scale.

Lance, Arrow, or other columnar/indexed storage formats.

Multi-cloud GPU compute (e.g., SkyPilot).

Model governance or audit requirements in financial services or other regulated environments.

Original posting on Avra's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job