Together AI
Senior Software Engineer — Infra Agent Systems
Amsterdam
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.4M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Engineering
- Seniority
- Senior
- Country
- NL
- Work mode
- Remote-friendly
- First seen by hirly
- 6 Sept 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
About the Role
Together AI runs one of the largest GPU fleets in the world. The Infra Agent Systems team builds the software systems that power and automate that infrastructure.
We develop production AI agents that diagnose hardware failures, investigate incidents, correlate signals across the fleet, and automate operational workflows. Alongside these agents, we build the platform they run on, including knowledge graphs, retrieval systems, orchestration frameworks, and developer tooling.
You’ll work across two areas:
Infrastructure Agent Systems — Build production AI agents that help operate our GPU fleet by diagnosing failures, investigating incidents, gathering evidence from live systems, and assisting with remediation. These agents are used every day by our infrastructure and datacenter teams through APIs, CLI, dashboards, and Slack.
Core Agent Platform — Build the platform that powers these agents, including knowledge graphs, search and retrieval, orchestration, evaluation, and the tooling that enables agents to reason, act, and continuously improve.
We’re working on something that hasn’t really been done before: building knowledge graphs and self-improving AI agents that understand, operate, and continuously improve large-scale AI infrastructure.
This is an opportunity to work at the intersection of AI agents, distributed systems, infrastructure, and automation , solving challenging engineering problems with real production impact. There’s an enormous amount to build, learn, and shape as we define the future of autonomous infrastructure.
responsible for delivering the software but also for operating and supporting it in production.
Why this Role
You’ll work on two hard problems at the same time: making AI agents trustworthy enough to operate production infrastructure, and building the knowledge, retrieval, and distributed systems that make those agents effective.
You’ll have the opportunity to build foundational systems from the ground up, work on infrastructure at massive scale, and help define how self-improving AI agents operate real-world AI infrastructure.
Hybrid in Amsterdam
Responsibilities
Design and build production AI agent systems that diagnose, investigate, and remediate infrastructure issues across one of the world’s largest GPU fleets.
Build the distributed services, orchestration framework, knowledge graph, and retrieval systems that power infrastructure agents.
Develop fleet intelligence systems that combine telemetry, infrastructure state, operational knowledge, and historical incidents to help agents make better decisions.
Integrate with observability, incident management, ticketing, fleet inventory, source control, chat, and internal infrastructure systems through well-designed APIs.
Own services end to end, including architecture, implementation, testing, deployment, observability, and production operations.
Improve agent performance through evaluations, retrieval improvements, better tools, and production feedback loops.
Turn what agents learn in production into reliable, reviewed software and automation.
Requirements
5+ years of experience building production backend systems, distributed systems, or infrastructure platforms.
Strong systems design skills and experience owning significant systems from design through production.
Depth in at least one of the following:
AI agent systems, orchestration, tool use, evaluation, or grounding
Knowledge graphs or graph data modeling
Search, retrieval, ranking, RAG, or semantic search systems
Strong backend engineering experience, including API design, service boundaries, data modeling, and integrations across complex systems.
Experience with Kubernetes, GitOps such as ArgoCD, infrastructure-as-code, and cloud platforms.
Comfortable working across languages such as Go, TypeScript, Python, or Rust.
Experience in the following is a plus:
GPU infrastructure, datacenters, bare-metal systems, hardware failure modes, BMC/IPMI, or cluster schedulers
Graph databases
Event-driven systems and messaging platforms such as NATS or Kafka
Observability platforms such as Prometheus and Grafana
Building evaluation frameworks or improving the quality and reliability of LLM-powered systems
About Together AI
Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinforcement learning for creating frontier-level specialized models, and pre-training at massive scale for fully custom intelligence, all around a marketplace of leading open models that teams can run, adapt, and own. Trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, Together serves 400+ trillion tokens a month.
Equal Opportunity
Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.
Please see our privacy policy at https://www.together.ai/privacy .
Similar jobs
- Software Engineer III (Full Stack)Relx · AmsterdamFirst seen today
- Senior Software Engineer - SDK and Visualization EnginesTomTom · Amsterdam, The NetherlandsFirst seen today
- Senior Software Engineer: Platform SREFlexport · Amsterdam, NetherlandsFirst seen 3d ago
- Senior Software Engineer | Data Infrastructure | RustBlockTech · AmsterdamFirst seen 4d ago
- Senior Software Engineer | Core Tech | Python/RustBlockTech · AmsterdamFirst seen 4d ago
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job