Benchmarks
Repeatability before claims.
Fixed suites expose regressions across generation, Re-Voice, and evaluation. Scores retain their limitations instead of collapsing uncertainty into a single claim.
Execution
Run each case from Generate; the API returns its sealed evaluation report.
No fabricated scores
The UI does not display a score until the backend evaluates a workflow.
Publishable study
Requires held-out corpora, human ratings, agreement, and confidence intervals.