AI evaluation, grounded in real human judgment
Expert-calibrated judges for evaluating, monitoring, and improving AI at scale
Expert-calibrated judges. Use ours or build your own.
Judges calibrated to our network of industry experts
Adopt judges off the shelf, built with practitioners who define the standards and reference cases behind every verdict. Use them as they are, or tailor them with your own policies and examples when your use case calls for it.
Browse judge catalogJudges calibrated to your internal team
Start from scratch and calibrate judges to the people who know your product best. Your reviewers set the standards and label reference cases, and our calibration tooling turns their judgment into judges you can run at scale.
Judges calibrated to our network of industry experts
Adopt judges off the shelf, built with practitioners who define the standards and reference cases behind every verdict. Use them as they are, or tailor them with your own policies and examples when your use case calls for it.
Browse judge catalogJudges calibrated to your internal team
Start from scratch and calibrate judges to the people who know your product best. Your reviewers set the standards and label reference cases, and our calibration tooling turns their judgment into judges you can run at scale.
Evaluate, monitor, and improve AI systems at scale
Judges can be used for online and offline evaluation of any AI system, directly on the platform or through our API. Issues are routed to human review queues and turned into fine-tuning data or system prompt suggestions.
Who Forum AI is for
A customer support agent can close a ticket but frustrate a user.
A workplace assistant can send an email without the critical information.
A coding agent can ship a change but add unnecessary complexity.
Forum AI partnered with experienced operators to develop template judges for a wide range of relevant factors, designed to be customized for your use cases.
Features
Online and offline evaluation
Ongoing agent trace monitoring, plus expert-designed prompt sets for offline evaluations.
One click customization
Upload sample agent traces and specify your priorities — we'll automatically tailor the relevant judges to your use case.
Internal calibration tooling
Build judges with your team through built-in calibration flows that adapt to your needs over time.
Popular use cases
- Customer Communication6 judges
- Productivity4 judges
- Safety7 judges
- Coding3 judges
- Finance5 judges
AI models are powering increasingly consequential tasks.
They offer mental health support. They guide voters during elections. They support new parents.
Forum AI partners with leading domain experts to develop industry-standard evaluations that both capture the necessary nuance and offer third-party defensibility.
Offerings
Standardized evaluation reports
Recurring industry-standard evaluations to track progress, compare competitors, and surface risks.
Commissioned evaluations
Bespoke prompt sets, fact sets, and judges tailored to your priorities for internal hillclimbing.
Popular use cases
- Mental Health6 judges
- News & Politics7 judges
- Finance5 judges
- Education4 judges
- Primary Health5 judges
Forum AI’s expert-calibrated judges
Every judge here works off the shelf. Send us your policy and a sample of your traces and we rebuild any of them around how your team actually works, then calibrate it until it agrees with you.
Adopt, customize, and deploy expert-calibrated judges in one place
Everything runs through a single workspace — drafting and calibrating judges, watching live traffic, and replaying history to see whether a change actually helped.
Judge Customization
Describe what matters and Forum drafts judges from the matching expert rubrics, then calibrates them against your own labels until they agree with you.
Online Monitoring
Run the accepted judges against live agent traffic. Findings and evidence gaps surface as they happen, with alerts on the ones worth interrupting you for.
Offline Evaluation
Replay historical activity against your own benchmark and watch the score move run over run, with the judgments that need attention surfaced alongside it.
Forum AI combines three data layers to produce judges that capture what matters
Latest research and insights
In the news
Forum AI's Campbell Brown: AI Needs Public Quality Testing
Forum AI study finds chatbots fail on election accuracy and sourcing 90% of the time
Forum AI's Campbell Brown on AI's accuracy gap in news and information
Forum AI and Allstate CEOs discuss AI's next chapter
Forum AI hosts release of Stanford HAI's 2026 AI Index Report
Forum AI cofounder discusses AI's judgment at Eye on AI
Campbell Brown co-launches Forum AI
Ex-Meta Executive, CNN Anchor Campbell Brown Launches Forum AI With $3 Million in Funding
Join our team
Help us build the judgment layer for AI — scaling expert evaluation across the domains that matter most.
Apply















































