Eval (LLM-as-judge)
An automatic scan every 3 hours scores LLM responses and detects quality drift.
Sample eval summary from the automatic scan (average score and 24-hour trend)
Mean score
4.40
vs previous scan
Average score trend (last 24 hours)
The judge AI is verified weekly
Every week the judge takes an exam on mutated copies of real calls, and its trustworthiness is shown as discrimination, stability, and human agreement.
Example judge trust stats (weekly automatic verification)
Computed from 24 probe tests
5 default criteria
Five criteria ship built in: helpfulness, accuracy, relevance, safety, and conciseness. Each is a 1-5 rubric you can open on screen to read exactly how scores are assigned. The defaults go through the same judge verification and calibration as everything else, so they stay operational rather than decorative.
AccuracyConcisenessHelpfulnessRelevanceSafetyCustom criteria
Add "matches our brand voice", etc. (Pro+). Each criterion's scoring rubric is editable.