Skip to content
Lucas BarriosApplied AI & Operational Transformation

Governance evaluation

AI Content Studio

A reproducible scorecard: a naive baseline against a governed system across 32 cases in two regulated demo brands. Generation on gpt-4o-mini, judging on gpt-4o (a stronger, different model to reduce self-preference bias), groundedness cross-checked with ragas.

Try it

Pick a regulated client and a topic, then generate. You get a naive baseline and the governed system side by side, with banned claims flagged and required disclosures checked.

Two live calls: a naive baseline and the governed system.

This live demo runs generation (gpt-4o-mini) and the deterministic governance layer: banned-term detection with true word boundaries and disclosure coverage. The full three-layer evaluation (context-aware LLM judge and an independent ragas cross-check) and the RAG pipeline run in the repository. Rate limited and cost-capped.

The full evaluation, 32 cases scored offline
Metricbaselinesystemnotes
True violation rate (judge, context-aware)0.280.00asserted prohibited claims
Adversarial safe-reframing (judge)0.001.008 of 8 traps refused
Groundedness (LLM judge)n/a0.95ragas faithfulness: 0.78
Brand-voice adherence (judge)0.370.80
Disclosure coverage0.270.59named gap, see below
Banned-term rate (raw, context-blind)0.440.22mostly negations, see below
Generic-opener rate0.030.00anti-slop check
Cost per run (USD)0.0120.01832 generations

Every figure is produced by a committed report in the repository and is reproducible from python -m evals.run.

What the numbers say, including where the system is weak

Two figures are deliberately unflattering, and they are what make the scorecard credible. Disclosure coverage is only 0.59: the system is excellent at not asserting prohibited claims (0.00 true violations) but merely good at reliably including required disclosures such as “past performance is not a reliable indicator.” That is a measured weakness and the next target, not something averaged away.

The raw banned-term rate (0.22) sits above the judge-confirmed rate (0.00). That gap is not a violation. It is context-blindness: a substring checker flags “results are not permanent” as a “permanent” violation. The judge exists to tell asserting a claim from refusing one, and both numbers are reported so the difference is visible rather than laundered.

Groundedness reads 0.95 by the judge and 0.78 by ragas. Both indicate strong grounding; the difference reflects the stricter atomic-claim decomposition ragas performs. Both are reported rather than the more flattering single number.

Three scoring layers

Deterministic

Offline checkers: banned terms with true word boundaries, disclosure coverage, generic openers. No API, unit-tested.

LLM-as-judge

A context-aware judge (temperature 0) separates asserting a claim from negating it, and scores groundedness and voice.

Independent

ragas Faithfulness on the system arm, isolated in a pinned dependency set, as a second opinion on grounding.

Two regulated demo tenants

Both are fictional, chosen because generic model output is a genuine compliance liability in each, which makes the governance delta legible. They share one platform with per-tenant isolation enforced by Postgres row-level security.

Meridian Wealth

DACH wealth advisory under financial-promotion rules: no “guaranteed,” “risk-free,” or “beat the market,” and mandatory risk disclosures.

Lumen Aesthetics

Berlin aesthetics clinic under medical-advertising rules: no “permanent,” “100 percent safe,” or “cure,” consent-first calls to action, outcomes stated as temporary.

Why the judge layer exists

One adversarial case asked for “permanent, flawless results.” The system refused the premise, writing that results are “not a permanent solution” and “not about achieving a flawless appearance.” The substring checker flagged the words “permanent” and “flawless” as violations; the context-aware judge correctly cleared it as a refusal. That single case is why true adversarial refusal is 8 of 8, not 7 of 8, and why the harness reports a raw number and a judged number side by side.

Reproduce it

python -m evals.run                 # generate and score, both arms
python -m evals.ragas_check         # independent faithfulness cross-check
python -m pytest evals/tests/ -q    # 27 tests; the evaluator is itself tested

Authored by Kairos Consulting as a reference implementation. Tenants are fictional. No fabricated metrics: every number is reproducible from a committed report.