Skip to main content

AI evaluation, grounded in real human judgment

Forum AI offers a suite of judges calibrated to leading experts, with a platform to adopt, customize, and deploy them.

Featured on:
How it works

Expert-built judges, tailored for your requirements.

01Judges

Expert-calibrated judges for different use cases

Start with judges for the decisions your AI needs to get right. Relevant practitioners define the standards and reference cases that guide how those judges are built and tested.

02Customize

Guided customization workflows to tailor the judges

Tailor judges with your reference examples and policies while preserving expert integrity. Validate them against your use case requirements.

03Evaluate

Offline evaluation and online monitoring

Use judges to monitor live systems or run offline evaluations with Forum AI’s expert-designed benchmarks. Review results, explanations, and supporting evidence before launch and for selected interactions after deployment.

Who Forum AI is for

A customer support agent can close a ticket but frustrate a user.

A workplace assistant can send an email without the critical information.

A coding agent can ship a change but add unnecessary complexity.

Forum AI partnered with experienced operators to develop template judges for a wide range of relevant factors, designed to be customized for your use cases.

Features

  • Online and offline evaluation

    Ongoing agent trace monitoring, plus expert-designed prompt sets for offline evaluations.

  • One click customization

    Upload sample agent traces and specify your priorities — we'll automatically tailor the relevant judges to your use case.

  • Internal calibration tooling

    Align judges with your team through built-in calibration flows that adapt to your needs over time.

Popular evaluations

  • Customer Communication

    • Policy Adherence
    • Tone & Language
    • Escalation to Human
    • User Frustration

    26 contributing experts

  • Productivity

    • Goal Alignment
    • Completeness
    • Efficiency
    • Appropriate Initiative

    21 contributing experts

  • Safety

    • Data Privacy
    • Access Boundaries
    • Harmful Actions
    • Human Oversight

    34 contributing experts

  • Coding

    • Correctness
    • Simplicity
    • Maintainability
    • Efficiency

    23 contributing experts

  • Finance

    • Numerical Accuracy
    • Policy Adherence
    • Auditability
    • Risk Identification

    29 contributing experts

Platform

Adopt, customize, and deploy expert-calibrated judges in one place

Everything runs through a single workspace — drafting and calibrating judges, watching live traffic, and replaying history to see whether a change actually helped.

Travel assistant · judges3 drafted
What matters to you
Refunds issued without approval
Drafted from expert rubrics
Authority & escalationReady
Policy applicationReady
Cancellation pressureDraft

Judge Customization

Describe what matters and Forum drafts judges from the matching expert rubrics, then calibrates them against your own labels until they agree with you.

Monitoring · live2 flagged
Policy applicationPass
Resolution qualityPass
Authority & escalationFlagged
Cancellation pressurePass
Policy applicationPass
Resolution qualityFlagged
Authority & escalationPass
Policy applicationPass
Cancellation pressurePass
Resolution qualityPass
Policy applicationPass
Resolution qualityPass
Authority & escalationFlagged
Cancellation pressurePass
Policy applicationPass
Resolution qualityFlagged
Authority & escalationPass
Policy applicationPass
Cancellation pressurePass
Resolution qualityPass
Policy applicationPass
Resolution qualityPass
Authority & escalationFlagged
Cancellation pressurePass
Policy applicationPass
Resolution qualityFlagged
Authority & escalationPass
Policy applicationPass
Cancellation pressurePass
Resolution qualityPass
Sampling live traffic · alerts on

Online Monitoring

Run the accepted judges against live agent traffic. Findings and evidence gaps surface as they happen, with alerts on the ones worth interrupting you for.

Your Customer Service Benchmark7 runs
Score over time+20 pts
50607080
MarSep
Authority & escalationPolicy applicationCancellation pressure
75Benchmark score
17Concerning flags

Offline Evaluation

Replay historical activity against your own benchmark and watch the score move run over run, with the judgments that need attention surfaced alongside it.

Insights

Latest research and insights

Preview of NewsBench
Benchmark

NewsBench

As AI informs voters, shapes policy, and drives real-world decisions, it must understand what's happening in the world around it. We partnered with world-leading experts to build a benchmark for high-stakes news coverage.

View benchmark
Preview of NewsBench: Expert-Grounded Evaluation of Epistemic Quality in AI News Reporting
Whitepaper

NewsBench: Expert-Grounded Evaluation of Epistemic Quality in AI News Reporting

As AI becomes a primary source of news, what matters is not just whether models avoid bias but whether they are accurate, well-sourced, and fair. NewsBench reframes evaluation around editorial standards set by senior journalists, policy experts, and intelligence analysts — measuring frontier models on source quality, factuality, and neutrality.

Read paper
Preview of Distilling Expert Judgment at Scale
Whitepaper

Distilling Expert Judgment at Scale

Frontier AI is being deployed where the stakes are high and expert judgment is required. We show how to encode that judgment — not just experts' conclusions, but their reasoning — into automated systems that scale, outperforming uncalibrated frontier models on every source-quality metric.

Read paper
Preview of How We Turn Expert Insight Into Action
Blog

How We Turn Expert Insight Into Action

From expert interviews to AI judges — how we transform domain expertise into scalable evaluation systems that improve AI where it matters most.

Read post
Preview of Not All Queries Are Created Equal
Blog

Not All Queries Are Created Equal

Engineering a classification system for LLM evaluation — why the type of question matters as much as the answer when measuring AI performance.

Read post
Preview of How We Pick the Right Experts to Evaluate AI
Blog

How We Pick the Right Experts to Evaluate AI

Four principles that guide our work — building the expert network that holds AI systems to the highest standards of accuracy and nuance.

Read post
Preview of Speed-Running Content Moderation
Blog

Speed-Running Content Moderation

What fifteen years of social media safety teaches about evaluating AI — lessons from the front lines applied to a new generation of challenges.

Read post
Preview of Ex-Meta Executive, CNN Anchor Campbell Brown Launches Forum AI
Press

Ex-Meta Executive, CNN Anchor Campbell Brown Launches Forum AI

The Wrap covers Forum AI's launch with $3 million in seed funding to bring expert judgment to AI evaluation.

Read article
Get started

Put real human judgment behind your AI

Careers

Join our team

Help us build the judgment layer for AI — scaling expert evaluation across the domains that matter most.

Apply