Matter Intelligence
Data Infra - Telemetry and Observability (IC)
San Francisco
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.6M live jobs from 190,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Seniority
- Mid level
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 18 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About Matter Intelligence
Welcome to Matter, where we are building the future of vision AI: pairing a world-first sensor that sees molecular chemistry, temperature, and 3D shape with a Large World Model that will be the most powerful intelligence engine for the physical world. This system doesn't just see what something looks like; it understands everything from a single pixel. We call this Superintelligent Vision.
Our team has delivered technologies to Mars for NASA/JPL, designed advanced sensors for U.S. Defense, and frontier artificial intelligence systems. We are now building the next generation of space- and airborne-based sensing systems.
About the Role
Matter is hiring a Telemetry and Observability Engineer to make datasets, training runs, models, environments, agents, and missions observable as one AI system. Reporting to Ignacio Cases Martin, this individual contributor will connect flight and mission reality with infrastructure, model, and product behavior through shared telemetry, traces, alerts, reliability mechanisms, and operational context.
Key Responsibilities
Build a common telemetry model linking mission state, datasets, code, checkpoints, accelerators, model and prompt versions, environment episodes, agent steps, tools, decisions, cost, latency, and feedback.
Instrument distributed training, data loading, GPU utilization, checkpointing, experiment health, batch and online inference, and model-serving behavior.
Instrument agent workflows across prompts, context, retrieval, memory, planning, tool calls, graph state, human intervention, evidence, outcomes, and safety controls.
Connect planned-versus-actual mission state, hardware context, time, geometry, and geolocation to relevant datasets and model results.
Establish actionable service objectives, alerts, incident interfaces, rollout and rollback signals, and escalation paths across platform, model, environment, and agent failures.
Build capacity and cost visibility across storage, processing, accelerators, training, evaluation, inference, retrieval, and agent execution.
Qualifications
Required
Experience with ML observability, telemetry, time-series systems, distributed training or inference, cloud infrastructure, or production AI platforms.
Strong hands-on engineering skills in observability, automation, APIs, infrastructure, incident tooling, and systems debugging.
Understanding of metrics, logs, traces, events, model and data monitoring, service objectives, capacity, release safety, and incident diagnosis.
Ability to reason about AI failure modes including bad datasets, training instability, stale checkpoints, drift, version mismatch, environment bugs, and agent-tool failure.
Experience designing telemetry schemas and correlation workflows across multiple systems, time domains, and ownership boundaries.
Preferred
Experience with MLOps, GPU workloads, model serving, agent observability, reinforcement-learning environments, mission telemetry, or scientific instrumentation.
Experience operating high-throughput inference, streaming systems, geospatial pipelines, or mixed cloud and edge deployments.
Experience with reliability platforms or observability products used by multiple engineering teams.
Familiarity with aerospace, defense, regulated operations, or other settings requiring formal change control and incident evidence.
What Success Looks Like
Teams can correlate mission, data, infrastructure, model, environment, and agent behavior during normal operation and incidents.
Alerts and service objectives identify actionable failures without obscuring scientific or operational context.
Capacity, cost, version, and reliability signals support safe deployments and faster diagnosis across the AI stack.
Location
This role is based in San Francisco, CA, and requires onsite work.
ITAR Requirements
To comply with U.S. export regulations, applicants must be one of the following:
A U.S. citizen or national
A lawful permanent resident (green card holder)
Eligible to obtain required authorizations from the U.S. Department of State
Employee Offerings and Benefits
At Matter, we believe in rewarding high performance and providing the support you need to thrive. Our compensation and benefits package includes:
Competitive compensation based on experience
Early-stage equity package
100% employer-paid health, dental, and vision coverage
Opportunity to work on novel sensing, data, and AI systems with real-world deployment paths to the largest industries in the world
Matter Intelligence is an equal opportunity employer. We welcome candidates from all backgrounds who can raise the ambition and performance of the team.
Listed on hirly, a job board. hirly is not the employer: Matter Intelligence is hiring for this role.
Similar jobs
- Software Engineer, Data InfrastructureThinking Machines Lab · San FranciscoFirst seen 2d ago
- Robotics Data Infrastructure EngineerVerne Robotics · San FranciscoFirst seen 6d ago
- AI Context & Data Infrastructure EngineerTown.com, Inc. · San FranciscoFirst seen 6d ago
- Data Infrastructure EngineerDroyd · San Francisco, CAFirst seen 6d ago
- C2C Forward Fellow, Data InfrastructureFoundationccc · California RemoteFirst seen 2d agoremote
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job