Governance evaluation
AI Content Studio
A reproducible scorecard: a naive baseline against a governed system across 32 cases in two regulated demo brands. Generation on gpt-4o-mini, judging on gpt-4o (a stronger, different model to reduce self-preference bias), groundedness cross-checked with ragas.
Try it
Pick a regulated client and a topic, then generate. You get a naive baseline and the governed system side by side, with banned claims flagged and required disclosures checked.
This live demo runs generation (gpt-4o-mini) and the deterministic governance layer: banned-term detection with true word boundaries and disclosure coverage. The full three-layer evaluation (context-aware LLM judge and an independent ragas cross-check) and the RAG pipeline run in the repository. Rate limited and cost-capped.
| Metric | baseline | system | notes |
|---|---|---|---|
| True violation rate (judge, context-aware) | 0.28 | 0.00 | asserted prohibited claims |
| Adversarial safe-reframing (judge) | 0.00 | 1.00 | 8 of 8 traps refused |
| Groundedness (LLM judge) | n/a | 0.95 | ragas faithfulness: 0.78 |
| Brand-voice adherence (judge) | 0.37 | 0.80 | |
| Disclosure coverage | 0.27 | 0.59 | named gap, see below |
| Banned-term rate (raw, context-blind) | 0.44 | 0.22 | mostly negations, see below |
| Generic-opener rate | 0.03 | 0.00 | anti-slop check |
| Cost per run (USD) | 0.012 | 0.018 | 32 generations |
Every figure is produced by a committed report in the repository and is reproducible from python -m evals.run.
What the numbers say, including where the system is weak
Two figures are deliberately unflattering, and they are what make the scorecard credible. Disclosure coverage is only 0.59: the system is excellent at not asserting prohibited claims (0.00 true violations) but merely good at reliably including required disclosures such as “past performance is not a reliable indicator.” That is a measured weakness and the next target, not something averaged away.
The raw banned-term rate (0.22) sits above the judge-confirmed rate (0.00). That gap is not a violation. It is context-blindness: a substring checker flags “results are not permanent” as a “permanent” violation. The judge exists to tell asserting a claim from refusing one, and both numbers are reported so the difference is visible rather than laundered.
Groundedness reads 0.95 by the judge and 0.78 by ragas. Both indicate strong grounding; the difference reflects the stricter atomic-claim decomposition ragas performs. Both are reported rather than the more flattering single number.
Three scoring layers
Deterministic
Offline checkers: banned terms with true word boundaries, disclosure coverage, generic openers. No API, unit-tested.
LLM-as-judge
A context-aware judge (temperature 0) separates asserting a claim from negating it, and scores groundedness and voice.
Independent
ragas Faithfulness on the system arm, isolated in a pinned dependency set, as a second opinion on grounding.
Two regulated demo tenants
Both are fictional, chosen because generic model output is a genuine compliance liability in each, which makes the governance delta legible. They share one platform with per-tenant isolation enforced by Postgres row-level security.
Meridian Wealth
DACH wealth advisory under financial-promotion rules: no “guaranteed,” “risk-free,” or “beat the market,” and mandatory risk disclosures.
Lumen Aesthetics
Berlin aesthetics clinic under medical-advertising rules: no “permanent,” “100 percent safe,” or “cure,” consent-first calls to action, outcomes stated as temporary.
Why the judge layer exists
One adversarial case asked for “permanent, flawless results.” The system refused the premise, writing that results are “not a permanent solution” and “not about achieving a flawless appearance.” The substring checker flagged the words “permanent” and “flawless” as violations; the context-aware judge correctly cleared it as a refusal. That single case is why true adversarial refusal is 8 of 8, not 7 of 8, and why the harness reports a raw number and a judged number side by side.
Reproduce it
python -m evals.run # generate and score, both arms
python -m evals.ragas_check # independent faithfulness cross-check
python -m pytest evals/tests/ -q # 27 tests; the evaluator is itself tested