Eval (LLM-as-judge)

An automatic scan every 3 hours scores LLM responses and detects quality drift.

Sample eval summary from the automatic scan (average score and 24-hour trend)

Mean score

4.40

↑ +0.18

vs previous scan

Average score trend (last 24 hours)

5.04.54.03.53.000:0006:0012:0018:0024:00
Last scan: 2 hours ago

The judge AI is verified weekly

Every week the judge takes an exam on mutated copies of real calls, and its trustworthiness is shown as discrimination, stability, and human agreement.

Example judge trust stats (weekly automatic verification)

Judge reliability96%
+2% vs last week

Computed from 24 probe tests

Detection96%24 degraded copies
Stability100%same meaning, same score
Human agreement (±1)92%12 pairs

5 default criteria

Five criteria ship built in: helpfulness, accuracy, relevance, safety, and conciseness. Each is a 1-5 rubric you can open on screen to read exactly how scores are assigned. The defaults go through the same judge verification and calibration as everything else, so they stay operational rather than decorative.

Helpfulness / Accuracy / Relevance / Safety / ConcisenessAccuracyConcisenessHelpfulnessRelevanceSafety

Custom criteria

Add "matches our brand voice", etc. (Pro+). Each criterion's scoring rubric is editable.

“Matches our in-house terminology”Checks consistency of internal wording Pro+
Add a criterion