hirly

Code Metal

Applied AI Research Engineer

Boston Hub · Remote · San Franscisco Hub

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Code Metal first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Seniority
Mid level
Country
US
Work mode
On-site / unstated
First seen by hirly
28 Sept 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Code Metal

  • Code Metal is the leader in automated software engineering you can trust. As AI writes more of the world's code, the bottleneck in software has shifted from writing code to verifying it works, and AI cannot verify its own work with certainty. Code Metal takes a fundamentally different approach: constrain AI to what it does reliably, verify every step independently of the model using formal methods, and keep engineers in the loop on the decisions that matter. The result isn't code that probably works — it's code that is provably correct, with auditable proof. Customers including the U.S. Air Force, L3Harris, RTX, and Toshiba use Code Metal to modernize legacy code, optimize performance on real hardware, and move prototypes to production, fast. Founded in 2023 with offices in Boston and San Francisco, Code Metal is funded by Accel, Salesforce Ventures, B Capital, Smith Point Capital, J2 Ventures, Shield Capital, Overmatch, RTX, and others.
  • Learn more at codemetal.ai .

About the Role

We're building next-generation AI systems that help military planners explore, compare, and evaluate operational courses of action. Our work combines frontier language models, simulation, planning, and verification into human-in-the-loop decision-support systems for defense applications. As an Applied AI Research Engineer, you’ll focus on human machine teaming and agentic AI to build systems that allow warfighters, planners, analysts, and decision-makers to explore operational choices with speed, confidence, and control.

This role focuses on designing and building agentic AI systems – not chatbots. You'll develop multi-agent workflows, fine-tune and evaluate models, build retrieval pipelines, experiment with post-training techniques, and integrate AI with simulation and planning software. You'll work closely with AI researchers, software engineers, and defense experts to turn research ideas into production-ready capabilities. The goal is to make complex planning, wargaming, adjudication, and analysis workflows faster, more explainable, and more trustworthy.

Research Areas of Interest

An incomplete list of ongoing and near-term directions:

Human-machine teaming for AI-assisted course-of-action development, comparison, critique, refinement, and operational decision support

Agentic planning systems that integrate language models with simulation, doctrine retrieval, external tools, structured outputs, and deterministic verification

Adapting and optimizing foundation models through fine-tuning, post-training, distillation, reinforcement learning, and rigorous evaluation for planning and decision-support tasks

Multi-agent AI systems for Red/Blue planning, control-cell support, adjudication, branch-and-sequel analysis, and collaborative planning workflows

Building reliable AI systems using self-correction, structured reasoning, constraint-aware generation, verification, and robust tool use

Learning from human expertise through planner feedback, preferences, approvals, synthetic data generation, and human-in-the-loop improvement

Trustworthy AI for high-consequence applications, with an emphasis on explainability, provenance, traceability, auditability, uncertainty estimation, and model behavior analysis

What You’ll Do

Design and build agentic AI systems for planning, decision support, and human-machine teaming

Develop AI pipelines that integrate foundation models, retrieval, simulation, external tools, and deterministic software

Design, run, and analyze experiments to evaluate model and agent performance, reliability, traceability, latency, cost, and user trust

Fine-tune, distill, and evaluate foundation models for domain-specific planning, reasoning, and decision-support tasks

Build datasets, retrieval pipelines, automated benchmarks, and experiment infrastructure to support continuous model improvement and reproducible research

Partner with software engineers to transition research prototypes into scalable AI services

Collaborate with domain experts to translate operational workflows into AI-enabled capabilities while ensuring AI outputs remain explainable, reviewable, and under human control

Why Code Metal?

Mission with impact: Build AI systems that help users reason through high-consequence operational decisions.

AI beyond demos: Work on systems where models are paired with software, verification, simulation, guardrails, and human oversight.

Greenfield research: Explore ambitious ideas in GenAI, RL, agentic workflows, evaluation, and human-machine teaming.

Small-team velocity: Move quickly from research question to prototype to user-facing capability.

Real users: See your work tested by planners, analysts, engineers, and operational stakeholders.

Must-Have Credentials

Bachelor's or Master's degree in Computer Science, Machine Learning, Engineering, Mathematics, Physics, or a related technical field, or equivalent practical experience.

3+ years building AI, machine learning, or applied research systems.

Strong Python engineering skills.

Experience with PyTorch and modern LLM tooling (Transformers, vLLM, Hugging Face, etc.).

Experience building or deploying agentic AI systems, tool-calling workflows, or multi-step reasoning pipelines.

Experience fine-tuning, evaluating, or serving language models.

Experience with retrieval-augmented generation, embeddings, vector search, or knowledge retrieval systems.

Strong understanding of experiment design, benchmarking, and model evaluation.

Ability to move quickly from research prototype to production-quality implementation.

Eligible to obtain a U.S. security clearance.

Benefits

Pay depends on experience, but we strive to be at the upper end of the salary range

Health care plan with 100% premium coverage, including medical, dental, and vision

401k with 5% matching

Paid Time Off (uncapped vacation, plus sick and public holidays)

Flexible hybrid or remote work arrangement

Relocation assistance for qualifying employees

Wage Transparency - The salary range for this role is not a guarantee of compensation or salary, as the final offer amount may vary based on factors including, but not limited to, individual proficiency, skills, experience, and location.

We are an equal opportunity employer. US Citizenship may be required for certain project assignments involving security clearance.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

Original posting on Code Metal's site ↗

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job