Threshold Detector
Identifies the exact concurrency level where p99 latency, error rate, or throughput crosses defined thresholds. Outputs a verdict, not raw numbers.
metric: latency_p99threshold: 800msbend_point: 240 usersconfidence: ±15 userssaturation_cause: db_connection_pool (fromObservability Hookup)saturation_cause field comes from cross-referencing the APM correlation. If LEO can name the bottleneck, it does.Consumes outputs from TARA, EVAN, LEO, and SARA in a single pass. Reconciles competing signals and surfaces the few that leadership should act on.
Compares current sprint signals to the prior 4–6 sprints. Flags trajectory not just snapshot — three consecutive declines is treated as a leading indicator.
Checks factual claims in model output against a grounded source — retrieved documents, API responses, supplied context. Flags fabricated facts, unsupported numbers, and invented citations. Scores grounding quality on a per-claim basis.
Verifies the model honors what the system prompt actually said — refused topics stay refused, required output format is produced, tool calls stay in scope, citation policy is enforced. Catches silent spec drift between prompt versions.
Runs counterfactual probes that vary protected attributes (race, gender, age, geography, disability, accent) while holding the underlying scenario constant. Surfaces demographic disparity in outcomes, recommendations, refusal rates, and tone.