AI evaluation, grounded in real human judgment
Forum AI offers a suite of judges calibrated to leading experts, with a platform to adopt, customize, and deploy them.
Expert-built judges, tailored for your requirements.
Expert-calibrated judges for different use cases
Start with judges for the decisions your AI needs to get right. Relevant practitioners define the standards and reference cases that guide how those judges are built and tested.
Guided customization workflows to tailor the judges
Tailor judges with your reference examples and policies while preserving expert integrity. Validate them against your use case requirements.
Offline evaluation and online monitoring
Use judges to monitor live systems or run offline evaluations with Forum AI’s expert-designed benchmarks. Review results, explanations, and supporting evidence before launch and for selected interactions after deployment.
Who Forum AI is for
A customer support agent can close a ticket but frustrate a user.
A workplace assistant can send an email without the critical information.
A coding agent can ship a change but add unnecessary complexity.
Forum AI partnered with experienced operators to develop template judges for a wide range of relevant factors, designed to be customized for your use cases.
Features
Online and offline evaluation
Ongoing agent trace monitoring, plus expert-designed prompt sets for offline evaluations.
One click customization
Upload sample agent traces and specify your priorities — we'll automatically tailor the relevant judges to your use case.
Internal calibration tooling
Align judges with your team through built-in calibration flows that adapt to your needs over time.
Popular evaluations
Customer Communication
- Policy Adherence
- Tone & Language
- Escalation to Human
- User Frustration
26 contributing experts
Productivity
- Goal Alignment
- Completeness
- Efficiency
- Appropriate Initiative
21 contributing experts
Safety
- Data Privacy
- Access Boundaries
- Harmful Actions
- Human Oversight
34 contributing experts
Coding
- Correctness
- Simplicity
- Maintainability
- Efficiency
23 contributing experts
Finance
- Numerical Accuracy
- Policy Adherence
- Auditability
- Risk Identification
29 contributing experts
AI models are powering increasingly consequential tasks.
They offer mental health support. They guide voters during elections. They support new parents.
Forum AI partners with leading domain experts to develop industry-standard evaluations that both capture the necessary nuance and offer third-party defensibility.
Offerings
Standardized evaluation reports
Recurring industry-standard evaluations to track progress, compare competitors, and surface risks.
Commissioned evaluations
Bespoke prompt sets, fact sets, and judges tailored to your priorities for internal hillclimbing.
Popular evaluations
Mental Health
- Clinical Safety
- Empathy
- Crisis Recognition
- Therapeutic Boundaries
38 contributing experts
News & Politics
- Factual Accuracy
- Neutrality
- Source Quality
- Loaded Premise Handling
31 contributing experts
Finance
- Factual Accuracy
- Risk Disclosure
- Suitability
- Uncertainty Communication
27 contributing experts
Education
- Instructional Accuracy
- Learner Adaptation
- Guided Reasoning
- Constructive Feedback
24 contributing experts
Primary Health
- Clinical Accuracy
- Triage & Escalation
- Uncertainty Communication
- Patient Comprehension
41 contributing experts
Adopt, customize, and deploy expert-calibrated judges in one place
Everything runs through a single workspace — drafting and calibrating judges, watching live traffic, and replaying history to see whether a change actually helped.
Judge Customization
Describe what matters and Forum drafts judges from the matching expert rubrics, then calibrates them against your own labels until they agree with you.
Online Monitoring
Run the accepted judges against live agent traffic. Findings and evidence gaps surface as they happen, with alerts on the ones worth interrupting you for.
Offline Evaluation
Replay historical activity against your own benchmark and watch the score move run over run, with the judgments that need attention surfaced alongside it.
Latest research and insights
In the news
Forum AI's Campbell Brown: AI Needs Public Quality Testing
Forum AI study finds chatbots fail on election accuracy and sourcing 90% of the time
Forum AI's Campbell Brown on AI's accuracy gap in news and information
Forum AI and Allstate CEOs discuss AI's next chapter
Forum AI hosts release of Stanford HAI's 2026 AI Index Report
Forum AI cofounder discusses AI's judgment at Eye on AI
Campbell Brown co-launches Forum AI
Ex-Meta Executive, CNN Anchor Campbell Brown Launches Forum AI With $3 Million in Funding
Join our team
Help us build the judgment layer for AI — scaling expert evaluation across the domains that matter most.
Apply







