hirly

Avra

Member of Technical Staff | Observability & Reliability

São Paulo

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Avra first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Lead / management
Country
BR
Work mode
Remote-friendly
First seen by hirly
23 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About the role

At Avra, every technical IC is a Member of Technical Staff (MTS). The title doesn't put anyone in a silo: you own systems and outcomes, not steps in a function, and you keep building depth in your area.

In this role, you'll join the Platform team as our go-to expert on observability and reliability. Our customers make real-time decisions based on our responses, so when we're down, their operations stop. Avra's cloud is just one more dataplane, alongside the dataplanes we operate inside customer environments — so observability and reliability have to work the same way everywhere.

What you'll do

Evolve our observability stack for logs, metrics, traces, and alerting.

Make sure every dataplane, in our cloud and on-premise, reports its active release, health, heartbeat, logs, metrics, and usage to the control plane.

Bring telemetry into customer clusters within a model where agents only make outbound connections.

Detect drift between the desired state and what's actually running in each environment.

Monitor the health of our deployment and runtime agents.

Provide visibility into ephemeral workloads, such as the Ray clusters that run our batch inference.

Define SLOs, lead incident response and postmortems, and reduce MTTR — including when a fix requires coordinating with the customer.

Reduce telemetry cost: less redundant data, more useful signal.

How we measure success

99.9% serving availability, with incidents trending down.

MTTR, including on-premise incidents.

Near-zero drift between desired and actual state.

All agents active and reporting, across every dataplane.

What we're looking for

Deep experience with OpenTelemetry and observability backends.

Hands-on practice with SLOs, error budgets, actionable alerting, and incident management.

Strong experience with Kubernetes and infrastructure as code (Terraform / Helm ).

Experience operating software in environments you don't fully control.

Production-quality code and reviews, and a willingness to operate what you build.

Nice to have

Shipping software to customer-hosted Kubernetes (e.g., Helm, outbound-only connectivity).

GCP or GKE, AWS or EKS.

ML multi-node/multi-cluster workloads in production.

Financial services or regulated environments.

Original posting on Avra's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job