hirly

Jda

Director, Model Behavior & Evaluation Systems

Paris

See how you match this job — and similar ones. Free.

Upload your resume and hirly scores it against this role at Jda first, then against similar open jobs, and shows where you fit and why.

PDF or DOCX, up to 12MB. No sign-up to see your matches.

Get past the screening software and onto a recruiter's desk

hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.

  • Keywords matched to this posting
  • Fit score before you apply
  • Cover letter included

Matched against 2.7M live jobs from 200,000+ employers in 200+ countries.

Tailor my resume for this job →

Apply from your AI assistant

Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.

Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.

hirly's read of this role

Role family
Supply chain
Seniority
Director
Country
FR
Work mode
On-site / unstated
First seen by hirly
3 Oct 2026

Derived automatically from the posting. Upload your resume above to see how the role scores against it.

the posting

About Blue Yonder

Blue Yonder is the AI company for supply chain. Our platform helps the world's leading companies plan, fulfill, deliver, and operate more resilient supply chains across complex global networks.

We are building toward the autonomous supply chain: intelligent systems that can understand operational context, reason through tradeoffs, use tools, collaborate with people, and take action across real supply chain workflows.

About Autonomy Labs

Autonomy Labs' mission is to find the fastest possible path to an autonomous supply chain.

We build LLM agents, learning systems, model training pipelines, evaluations, simulations, and decision-making systems for some of the hardest problems in global supply chain. The work spans LLMs, agentic workflows, tool use, software automation, evaluation, post-training, optimization, and production engineering.

In short, we are having a lot of fun.

Your Mission

We are looking for a deeply technical Director of Model Behavior & Evaluation Systems to own the behavioral quality system for Blue Yonder's LLM agents.

Our agents are not generic chatbots. They are being trained to operate supply chain software: querying state, calling APIs, interpreting operational context, proposing actions, handling exceptions, asking for missing information, and helping users make decisions in complex enterprise environments.

Your mission is to make model behavior a product-quality system, not a collection of dashboards. You will define what "good" means for agents operating supply chain workflows, establish the release gates that determine when behavior is ready to ship, and build the feedback loops that turn traces, customer feedback, SME review, telemetry, red-teaming, and eval failures into model improvements.

This is a director-level technical leadership role. You will lead through systems, standards, people, and decisions. You should be close enough to model traces, evals, post-training, tool use, and customer workflows to make strong technical calls, while operating at the level of ownership boundaries, launch authority, roadmap sequencing, and team building.

The stack is real and close to the work. You should expect to operate around Python, PyTorch, Hugging Face Transformers and Datasets, NVIDIA NeMo RL, OpenAI Agents SDK, Langfuse, LLM evaluation harnesses, tool-calling traces, model checkpoints, reward and preference data, synthetic scenarios, experiment reports, and production observability. You do not need to be the person implementing every pipeline, but you do need the technical depth to challenge designs, read artifacts, understand failure modes, and guide senior engineers toward better systems.

What You'll Do:

Own the behavioral quality bar for Blue Yonder's LLM agents across customer-facing supply chain workflows.

Build and lead the model behavior and evaluation systems function across behavior specs, eval governance, SME review, release gates, regression coverage, and launch readiness.

Define launch criteria across operational correctness, tool-use accuracy, workflow completion, escalation quality, safe fallback behavior, refusal quality, consistency, and customer trust.

Establish evaluation authority so evals become release decision infrastructure, not just model-quality reporting.

Set the technical direction for behavior and eval infrastructure across Python eval harnesses, OpenAI Agents SDK workflows, Langfuse traces, LLM-as-judge workflows, deterministic checks, trace analysis, reward/report versioning, and model-candidate comparison.

Convert model traces, tool-call failures, SME feedback, red-team findings, telemetry, and customer-facing failures into behavior specs, eval requirements, training data needs, and model improvement priorities.

Partner with the reinforcement learning and post-training organization to turn behavior gaps into SFT data, preference data, reward criteria, curriculum, NeMo RL experiments, model-candidate decisions, and regression tests.

Partner with workflow, data, product, and domain experts to turn supply chain workflow truth into durable scenario coverage, rubrics, synthetic scenarios, eval datasets, and training data requirements.

Partner with agent architecture and product engineering teams to ensure prompts, tools, APIs, skills, system instructions, and product workflows express the intended model behavior consistently.

Review model traces, eval outputs, experiment summaries, dataset slices, reward reports, and post-training results closely enough to make informed launch and roadmap decisions.

Own customer and user behavior discovery for agent workflows: what users expect agents to do, explain, ask, verify, escalate, refuse, and act on.

Lead red-teaming and behavioral risk programs for hallucinated operational claims, incorrect tool use, overconfidence, poor escalation, prompt injection, unsafe recommendations, sycophancy, and unhelpful refusals.

Define the operating cadence for behavioral quality reviews, model release readiness, model cards, issue triage, regression management, and post-launch behavior monitoring.

Build a high-performing team or function around model behavior, evaluation governance, SME review systems, and behavioral launch quality.

Communicate model behavior strategy, quality tradeoffs, launch readiness, and residual risks clearly to executives, customers, product teams, engineering teams, and domain experts.

What We're Looking For - We want to talk if you:

Have led technical work on LLM products, AI agents, model behavior, model evaluation, post-training, AI alignment, or AI product quality.

Have built or led evaluation systems for open-ended LLM behavior, including eval datasets, graders, rubrics, scorecards, SME review loops, regression tests, or launch gates.

Have hands-on depth with LLM tool calling, function calling, API agents, workflow agents, software-operating agents, or enterprise automation.

Are technically fluent in the modern LLM stack, including Python, PyTorch, Hugging Face Transformers, Hugging Face Datasets, OpenAI Agents SDK, Langfuse, model checkpoints, tokenization, inference behavior, and experiment analysis.

Understand post-training workflows well enough to partner deeply with ML teams, including supervised fine-tuning, preference data, RLHF/RLAIF, reward modeling, NeMo RL, system prompting, synthetic data, and model launch evaluation.

Can reason about practical LLM behavior failures: hallucination, sycophancy, overconfidence, verbosity, refusal behavior, instruction hierarchy, prompt injection, tool-use failures, eval drift, and behavioral regressions.

Have enough technical fluency to inspect model traces, tool-call transcripts, eval failures, telemetry, datasets, experiment configs, and model-quality reports, then guide teams toward concrete fixes.

Have strong product judgment and experience translating user or customer needs into shipped product behavior.

Can lead cross-functional work across product, research, engineering, design, security, legal, customer success, domain experts, and executives.

Are comfortable making launch-readiness decisions under ambiguity, with explicit tradeoffs and evidence.

Are an exceptional written communicator who can turn ambiguous behavioral principles into clear guidelines, examples, rubrics, specs, and decisions.

Have experience defining model specifications, behavior guidelines, system instruction hierarchies, annotation rubrics, policy taxonomies, or model behavior standards.

Have managed, built, or strongly influenced senior technical teams.

Experience That Stands Out - Strong candidates may also bring experience in one or more of the following areas:

Owning model behavior, assistant quality, agent behavior, LLM product quality, or AI alignment work at a frontier lab, AI product company, or enterprise AI platform.

Scaling an evaluation, model-quality, or launch-readiness function fro

Original posting on Jda's site ↗

Listed on hirly, a job board. hirly is not the employer: Jda is hiring for this role.

Browse similar roles

Want this one?

Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.

Tailor my resume for this job