Evaluation Runs
live, this tenant
Cases Evaluated
live, summed across runs
Regression Alerts
live, fired for this tenant
Service
live, from /healthz
Measured benchmarks

These are recorded results, not live telemetry, and are labelled as such. Each was measured on this deployment — a Hetzner CX33, 4 vCPU / 8 GB, running k3s — with row-level security enforced and a deterministic mock judge. Reproduction steps and the full conditions are in the repository README.

WhatMeasuredTarget
Read throughput under RLS (AC5) 1,000 RPS held 2 min · p95 2.2 ms · p99 3.55 ms · 0.00% errors 1,000 RPS, p99 < 100 ms
Rollup read over 101,200 rows (AC3) p95 0.07 ms, against a raw-scan control at 104.79 ms p95 < 50 ms
False alarms at zero true effect (AC4) 0.000 on normal and bimodal scores; 0.075 / 0.065 / 0.070 on binary at n = 20 / 50 / 100, against 0.405 / 0.250 / 0.170 for a naive fixed-threshold rule bounded, and stated

The engine is conservative rather than calibrated: firing needs both a minimum effect size and statistical significance, so the false-alarm rate sits below the configured α rather than at it, and detection does not improve with sample size near the effect gate. Both are measured and published rather than smoothed over.

Recent Evaluation Runs
Run ID Dataset Status Mean Score Started
Loading evaluation runs...
Regression Alerts Feed
Loading alerts…