Context
AI Content Studio is a reference implementation of a governed content pipeline for regulated industries. It answers the question most AI content tools skip: not whether you can generate on-brand content, but whether you can prove it stays inside a client's compliance boundary before it ships.
The work is authored by Kairos Consulting as a capability demonstration. Both client tenants are fictional and deliberately regulated, a DACH wealth advisory and a Berlin aesthetics clinic, because generic model output is a genuine compliance liability in each. That makes the value of governance measurable rather than rhetorical.
The headline deliverable is not the application. It is the evaluation harness and its scorecard, because a system that claims to be brand-safe but cannot demonstrate it is worth very little.
Problem
Generative models make brand content cheap to produce and easy to get wrong. A wealth advisory cannot promise guaranteed returns. An aesthetics clinic cannot promise permanent or 100 percent safe results. A general-purpose model will write both without hesitation, and a naive prompt-and-hope tool will happily ship them.
The hard part is not generation. It is verification: showing, before anything reaches a client's audience, that the output honors brand voice, includes mandatory disclosures, and refuses prohibited claims. That evidence is exactly what a regulated buyer needs and what most AI content products cannot provide.
- Prohibited claims that create regulatory exposure: guaranteed returns, permanent results, 100 percent safe.
- Missing mandatory disclosures: past performance, results vary, possible side effects.
- Off-brand voice that erodes trust in a premium or clinical setting.
Workflow
The evaluation runs two arms against the same 32-case golden dataset and scores them identically. The baseline is a naive generic prompt with no brand, no retrieval, and no compliance contract. The system is the intended production pipeline: the client's brand profile injected at generation time, retrieval grounding over that client's knowledge base, and a compliance-aware template.
Each generation is scored in three layers, from most objective to least: deterministic checkers, a context-aware LLM judge, and an independent ragas cross-check. Generation runs on gpt-4o-mini; judging runs on gpt-4o, a stronger and different model chosen to reduce self-preference bias.
The whole loop is reproducible from a single command, and the evaluator itself is unit-tested with 27 tests, so the thing measuring quality is held to the same standard as the thing being measured.
01
Controlled comparison, same model on both arms
Baseline and system arms
Two generation paths scored on the same 32-case dataset: a naive prompt versus the governed pipeline of brand profile, retrieval, and compliance template.
02
Objective, reproducible, unit-tested
Deterministic checks
Offline, API-free checkers for banned terms with true word boundaries, disclosure coverage, and generic openers.
03
Judge model stronger than the generator
LLM-as-judge
A context-aware judge separates asserting a prohibited claim from negating it, and scores groundedness and brand voice at temperature 0.
04
Cross-validation against own interest
Independent cross-check
ragas Faithfulness on the system arm, isolated in a pinned dependency set, as a second opinion on groundedness.
05
Honesty over optics
Dual-number scorecard
Raw context-blind and judge-adjudicated rates reported side by side, with the gap explained as false positives.
Architecture
A Next.js frontend and TypeScript prompt framework sit in front of a FastAPI backend that handles brand governance, retrieval, and generation. State lives in Supabase Postgres with pgvector for embeddings, and row-level security isolates every tenant on every table.
The multi-tenant design is the point an agency or consultancy cares about: one platform serving two opposite regulatory regimes, with each client's brand rules and knowledge base scoped to that client and no other.
Generation and governance
FastAPI services for brand governance, retrieval, and generation.
- Brand profiles: approved and banned terms, voice, compliance notes
- Compliance-aware prompt templates
- OpenAI generation (gpt-4o-mini)
Retrieval
Per-tenant RAG over Supabase pgvector.
- OpenAI embeddings (text-embedding-3-small)
- HNSW cosine index
- Content-hash dedup on ingestion
Data and isolation
Supabase Postgres with strict tenant separation.
- Row-level security on every table
- Two seeded regulated tenants under one demo org
- Version-controlled migrations
Evaluation
The evals package: dataset, arms, checkers, judges, cross-check, scorecard.
- 32-case golden dataset
- Three-layer scoring
- 27 harness unit tests
Governance
Governance is enforced at generation time and verified after. At generation, the brand profile supplies approved and banned terms, voice, and compliance notes; retrieval grounds claims in the client's own knowledge base rather than the model's parametric memory.
Verification is where the project earns its name. A context-aware judge distinguishes asserting a prohibited claim from refusing one, which a naive substring checker cannot do. The system reports both the raw context-blind number and the judge-adjudicated number, so the gap between them is visible rather than hidden.
- Banned-claim enforcement with true word boundaries, so cure does not match manicure or procedure.
- Required-disclosure coverage checked per case against the disclosures that case demands.
- Adversarial traps: 8 of the 32 cases are engineered to elicit a violation, and the system refused all 8.
Metrics
Against the baseline, the governed system eliminated asserted compliance violations and refused every adversarial trap, while staying transparent about where it is still weak. Under context-aware review the true violation rate fell from 0.28 to 0.00, and all 8 adversarial cases were safely reframed rather than taken.
Groundedness reached 0.95 by the system's own judge and 0.78 by the independent ragas cross-check. Both numbers are reported, not the more flattering one; the gap reflects the stricter atomic-claim decomposition ragas performs. Brand-voice adherence rose from 0.37 to 0.80.
Disclosure coverage, the honest weak point, improved from 0.27 to 0.59 and is named as the next target rather than averaged away. The system is far better at not asserting prohibited claims than at reliably including required ones, and the scorecard says so plainly.
- True violation rate
- 0.00
- Adversarial traps refused
- 8 / 8
- Groundedness
- 0.95 / 0.78
- Golden test cases
- 32
Asserted prohibited claims under context-aware judge review, down from 0.28 at baseline.
Cases engineered to elicit a compliance violation, all safely reframed rather than taken.
LLM judge and independent ragas faithfulness, both reported rather than the more flattering one.
Version-controlled cases across two regulated tenants, including 8 adversarial, with 27 tests on the evaluator itself.
Roadmap
The evaluation is the foundation that makes the rest measurable. Each subsequent phase is judged by whether it moves a number the harness already tracks.
- Production RAG: hybrid retrieval and reranking, measured against the current 0.95 judge and 0.78 ragas groundedness.
- Guardrails: prompt-injection defense, PII checks, and explicit OWASP-LLM and EU AI Act mapping, with disclosure coverage at 0.59 as the named target.
- Cost and observability: per-generation tracing and model routing.
- Path consolidation: wire vector retrieval into the primary generation endpoint and unify the two generation paths.
Phase 1
Production RAG
Hybrid retrieval with BM25 and vector, plus reranking, measured against current groundedness.
Phase 2
Guardrails
Prompt-injection defense, PII checks, and OWASP-LLM and EU AI Act mapping; disclosure coverage is the named target.
Phase 3
Cost and observability
Per-generation tracing and model routing.
Phase 4
Path consolidation
Wire vector retrieval into the primary generation endpoint and unify the two generation paths.
Reflection
The most valuable output was not the passing scores. It was the harness surfacing four real defects a finished-looking system had hidden: schema drift that meant retrieval had never run end to end, a dedup check that silently blocked reingestion of failed documents, a half-populated knowledge base, and a retrieval threshold miscalibrated for the embedding model in use.
The second lesson was methodological. A compliance checker that cannot tell an assertion from a refusal will punish careful, correct writing. The context-aware judge exists to fix exactly that, and one adversarial case that refused the trap while quoting the trap's own words proved why the layer is necessary.
Technical depth
System assumptions and operating controls.
Architecture diagram
A multi-tenant content pipeline (Next.js, FastAPI, Supabase pgvector) with brand governance and RAG grounding, evaluated by a three-layer harness: deterministic checkers, a context-aware LLM judge on a stronger model than the generator, and an independent ragas faithfulness cross-check. Generation runs on gpt-4o-mini, judging on gpt-4o. Every reported number is reproducible from a committed report.
01
Frontend
Next.js and a TypeScript prompt framework.
02
Backend
FastAPI: brand governance, retrieval, generation.
03
Retrieval
OpenAI embeddings into Supabase pgvector (HNSW cosine), scoped per tenant.
04
Data
Supabase Postgres, row-level security on every table, two seeded regulated tenants.
05
Evaluation
evals package: dataset, arms, checkers, judges, ragas cross-check, scorecard.
Knowledge source assumptions
Per-tenant knowledge base (brand guidelines, product and treatment docs)
Kairos
Each client's own approved documents are the ground truth for factual claims.
text-embedding-3-small
Kairos
On-topic cosine similarity runs roughly 0.3 to 0.5, so the retrieval threshold is set to 0.15 for this corpus size; Phase 1 reranking restores precision.
Evaluation metrics
True violation rate
0.00 (system), down from 0.28
Context-aware LLM judge (gpt-4o), assert versus negate
Adversarial safe-reframing
8 / 8 refused
Judge adjudication on 8 trap cases
Groundedness
0.95 judge, 0.78 ragas
LLM judge and independent ragas Faithfulness
Brand-voice adherence
0.80 (system), up from 0.37
LLM judge against a per-tenant voice descriptor
Disclosure coverage
0.59 (named gap)
Deterministic per-case disclosure-family check
Risk and failure scenarios
Context-blind false positives
A substring checker flags negations such as results are not permanent as violations.
A context-aware judge adjudicates; raw and judged numbers are both reported.
Silent brand or retrieval degradation
An empty brand block or zero retrieved chunks could score a broken pipeline as valid.
The run flags degraded generations loudly instead of scoring them silently.
Self-preference bias
A model judging its own family inflates scores.
Judging runs on gpt-4o while generation runs on gpt-4o-mini; a cross-provider judge is supported via JUDGE_MODEL.
Under-disclosed output
The system asserts fewer prohibited claims than it reliably includes required disclosures (0.59 coverage).
Named as the next guardrail target rather than hidden.
Groundedness metric optimism
A single holistic judge can read higher than atomic-claim checking.
An independent ragas cross-check of 0.78 is reported beside the judge's 0.95.
Human review checkpoints
Scorecard review before any claim is published
Kairos
Ship only numbers reproducible from a committed report
Diff review before merge to main
Founder
No fabricated metrics, clients, or history