Avra
Member of Technical Staff | Inference Platform
São Paulo
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Lead / management
- Country
- BR
- Work mode
- Remote-friendly
- First seen by hirly
- 23 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About the role
At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area.
In this role, you'll join the Platform team to own where our models execute. Customers consume our models through large batches of millions of records and through real-time APIs, and they make business decisions on every response. You'll run governed model releases reliably and efficiently — in our cloud and on customer-hosted Kubernetes — and make inference fast, predictable, and cheap enough to serve both enterprise and mid-market customers.
What you'll do
Evolve Sophos, our online and batch inference runtime, built on Kubernetes.
Run large batch inference on ephemeral jobs, with multi-dimensional admission control (CPU, memory, GPU).
Build and extend the controller and its Kubernetes custom resources.
Optimize each model's inference engine and feature processing.
Serve graphs and data efficiently.
Own execution of training, post-training, and fine-tuning jobs, in our cloud and in customer dataplanes / on-premisse cloud.
Drive autoscaling, GPU serving, performance, and cost optimization, with telemetry for every model we run.
How we measure success
99.9% serving availability.
p95/p99 latency for online inference and throughput for batch.
Cost per prediction and per training job.
GPU utilization: paid capacity versus capacity actually used.
Training and batch jobs that finish on time and succeed without manual retries.
What we're looking for
Experience running model serving or large-scale batch compute on Kubernetes.
Experience building Kubernetes controllers or operators.
Skill at profiling and optimizing data-heavy Python pipelines.
A clear sense of cost: you treat compute efficiency as a product feature.
Production-quality code and reviews, and a willingness to operate what you build.
Nice to have
Ray, Ray Serve, or KubeRay in production.
Admission-control systems.
GPU serving and performance optimization.
Arrow, Parquet, Lance, or other columnar formats.
Shipping software to customer-hosted Kubernetes.
GCP/AWS and GKE/EKS, and financial services or regulated environments.
Similar jobs
- Sr. Member Technical Staff - ESD and Latch-Up - HBMMicron · Folsom, CAFirst seen 6d ago
- Member Technical Staff - Applied AI Engineer (US Timing) Composio · BangaloreFirst seen 25d ago
- Member Technical StaffPirros · Los Angeles OfficeFirst seen 31d ago
- Member TechnicalBroadridge · Bengaluru-EPIP Industrial AreaFirst seen 8d ago
- Senior Member TechnicalBroadridge · Hyderabad-Hi-Tec CityFirst seen 10d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job