Kcura
Product Manager - AI Evaluations Tooling
Illinois
Get past the screening software and onto a recruiter's desk
hirly rewrites your resume for this job — matching the keywords and skills in the posting, moving your most relevant experience to the top, and writing a cover letter to fit. About 30 seconds.
- Keywords matched to this posting
- Fit score before you apply
- Cover letter included
Matched against 2.5M live jobs from 200,000+ employers in 200+ countries.
Tailor my resume for this job →Apply from your AI assistant
Connect hirly to Claude and ask it to apply to this job. hirly tailors your resume, fills the employer’s form and asks before sending. ChatGPT: manual setup today.
Some employer sites stop an application at a CAPTCHA or sign-in and hand it back with a link. Applying needs a paid plan. Works with any assistant that supports MCP.
hirly's read of this role
- Role family
- Product management
- Seniority
- Lead / management
- Stated salary
- $115,000 per year
- Country
- US
- Work mode
- On-site / unstated
- First seen by hirly
- 1 Oct 2026
Derived automatically from the posting. Upload your resume above to see how the role scores against it.
the posting
Posting Type
Remote/Hybrid
Job Overview
Relativity's AI portfolio spans 20+ products and we are moving quickly into agentic capabilities. As that portfolio grows, so does the opportunity to give every team building AI at Relativity a shared, standardized way to understand and improve their own AI performance. The Evaluations Tooling team exists to build that foundation: a platform that standardizes how we evaluate AI, accelerates how quickly teams can ship it with confidence, and turns AI quality into shared evidence that applied scientists, application teams, product managers, and legal experts can all act on.
We are looking for a Product Manager to own the product strategy and roadmap for Relativity's evaluation platform. Your mission is to enable every team at Relativity building AI products to understand, measure, and improve their own AI performance. That means productizing what applied scientists have already validated into a platform that application teams, PMs, and legal subject-matter experts can use directly; building a world-class agentic testing system that lets us ship new agentic capabilities quickly and confidently; and using production monitoring and team-facing dashboards to move the organization from reacting to quality problems to proactively getting ahead of them.
This is a platform PM role with unusually direct leverage. A successful candidate is energized by deeply technical customers, comfortable making build-vs-buy calls in a fast-moving vendor landscape, and able to hold a categorical bar ("every AI product ships behind elite evals") while sequencing pragmatically to get there.
Job Description and Requirements
Role Responsibilities
Own the product vision, strategy, and roadmap for the Evaluations Platform across its pillars: offline evaluation pipeline, online evaluation and production monitoring, eval discovery and reuse, SME authoring and approval, and agentic evaluation.
Serve every team building AI at Relativity. Run continuous discovery with applied scientists, application engineering teams, and product PMs to understand how each of them experiences AI quality today and design the platform so teams can measure and improve their own performance without routing every question through an applied scientist.
Lead the orchestration of a world-class agentic testing system: trajectory evaluation, tool-call correctness, intermediate-state rubrics, long-running judge orchestration, and simulation of multi-step flows. Make it fast and safe to ship new agentic capabilities and make the platform a competitive advantage in how quickly Relativity can iterate on agents.
Work closely with legal subject-matter experts to define the tasks our AI is evaluated against, what a correct outcome looks like, where the hard cases are, and how judgment criteria should be expressed so they can be encoded, scored, and reused across products. Build the SME authoring and approval workflow around how experts work.
Own the team-facing quality dashboards. Every AI product team should be able to open a view and understand their own performance, where they stand against their quality bar, and what changed since last release, with no engineering help required.
Lead build-vs-buy evaluations for tracing infrastructure and commercial eval tooling and run a recurring review as the vendor landscape shifts.
Partner with engineering and applied science on judge calibration so LLM-as-judge is trustworthy enough to gate deployments.
Clearly articulate technical tradeoffs to stakeholders ranging from applied scientists to attorneys to executive leadership.
Preferred qualifications:
Experience with LLM-based products and a working understanding of how they are evaluated (rubrics, LLM-as-judge, offline vs. online evaluation). Direct experience with agentic systems is a plus, not a requirement.
Experience as a platform, developer-tools, or data/ML PM serving technical customers, or a demonstrated ability to learn a technical domain quickly and earn credibility with engineers and scientists.
Comfort bringing non-engineering domain experts into technical workflows and translating their judgment into something a system can act on.
Strong analytical instincts: sets metrics, tests assumptions, and changes course when the evidence says to.
Mi nimum qualifications:
4+ years of product management experience at a technology company, with at least 2 years on technical, platform, infrastructure, or data/ML products.
Solid understanding of the software development lifecycle and agile practices.
Experience conducting user discovery with technical audiences and translating findings into roadmap decisions.
Excellent written and verbal communication skills, with the ability to explain complex technical concepts to both technical and non-technical audiences.
.
Relativity is committed to competitive, fair, and equitable compensation practices.
This position is eligible for total compensation which includes a competitive base salary, an annual performance bonus, and long-term incentives.
The expected salary range for this role is between following values:
$115,000 and $173,000
The final offered salary will be based on several factors, including but not limited to the candidate's depth of experience, skill set, qualifications, and internal pay equity. Hiring at the top end of the range would not be typical, to allow for future meaningful salary growth in this position.
Required Skills:
Agile Methodology, Cross-Functional Teamwork, Market Research, Market Strategy, Product Lifecycle, Product Lifecycle Management (PLM), Product Management, Product Strategies, Stakeholder Management, Team Leadership
Similar jobs
- Senior Product ManagerEpicgames · BLANK,BLANK,Multiple LocationsFirst seen today
- Senior Product ManagerEpicgames · Cary,North Carolina,United StatesFirst seen today
- Sr Product Manager - Enterprise SolutionsWorkday · 2 LocationsFirst seen today
- Sr Principal Product Manager, Customer ZeroWorkday · USA, CA, PleasantonFirst seen today
- Product ManagerWestlake · 7 LocationsFirst seen today
Browse similar roles
Want this one?
Upload your resume and hirly rewrites it for this job and writes the cover letter — in about thirty seconds, before you sign up.
Tailor my resume for this job