Recent Evaluation Runs
| Run ID | Dataset | Status | Mean Score | Started |
|---|---|---|---|---|
| Loading evaluation runs... | ||||
Regression Alerts Feed
Loading alerts…
These are recorded results, not live telemetry, and are labelled as such. Each was measured on this deployment — a Hetzner CX33, 4 vCPU / 8 GB, running k3s — with row-level security enforced and a deterministic mock judge. Reproduction steps and the full conditions are in the repository README.
| What | Measured | Target |
|---|---|---|
| Read throughput under RLS (AC5) | 1,000 RPS held 2 min · p95 2.2 ms · p99 3.55 ms · 0.00% errors | 1,000 RPS, p99 < 100 ms |
| Rollup read over 101,200 rows (AC3) | p95 0.07 ms, against a raw-scan control at 104.79 ms | p95 < 50 ms |
| False alarms at zero true effect (AC4) | 0.000 on normal and bimodal scores; 0.075 / 0.065 / 0.070 on binary at n = 20 / 50 / 100, against 0.405 / 0.250 / 0.170 for a naive fixed-threshold rule | bounded, and stated |
The engine is conservative rather than calibrated: firing needs both a minimum effect size and statistical significance, so the false-alarm rate sits below the configured α rather than at it, and detection does not improve with sample size near the effect gate. Both are measured and published rather than smoothed over.
| Run ID | Dataset | Status | Mean Score | Started |
|---|---|---|---|---|
| Loading evaluation runs... | ||||