Skip to content
Lucas BarriosApplied AI & Operational Transformation

Kairos Consulting · reference implementation

AI Content Studio

A governed, multi-tenant content pipeline for regulated industries, and the reproducible evaluation harness that measures whether the governance holds. Brand-safe guardrails and RAG grounding, scored by three layers: deterministic checks, a context-aware LLM judge on a stronger model than the generator, and an independent ragas cross-check.

Context

AI Content Studio is a reference implementation of a governed content pipeline for regulated industries. It answers the question most AI content tools skip: not whether you can generate on-brand content, but whether you can prove it stays inside a client's compliance boundary before it ships.

The work is authored by Kairos Consulting as a capability demonstration. Both client tenants are fictional and deliberately regulated, a DACH wealth advisory and a Berlin aesthetics clinic, because generic model output is a genuine compliance liability in each. That makes the value of governance measurable rather than rhetorical.

The headline deliverable is not the application. It is the evaluation harness and its scorecard, because a system that claims to be brand-safe but cannot demonstrate it is worth very little.

Problem

Generative models make brand content cheap to produce and easy to get wrong. A wealth advisory cannot promise guaranteed returns. An aesthetics clinic cannot promise permanent or 100 percent safe results. A general-purpose model will write both without hesitation, and a naive prompt-and-hope tool will happily ship them.

The hard part is not generation. It is verification: showing, before anything reaches a client's audience, that the output honors brand voice, includes mandatory disclosures, and refuses prohibited claims. That evidence is exactly what a regulated buyer needs and what most AI content products cannot provide.

  • Prohibited claims that create regulatory exposure: guaranteed returns, permanent results, 100 percent safe.
  • Missing mandatory disclosures: past performance, results vary, possible side effects.
  • Off-brand voice that erodes trust in a premium or clinical setting.

Workflow

The evaluation runs two arms against the same 32-case golden dataset and scores them identically. The baseline is a naive generic prompt with no brand, no retrieval, and no compliance contract. The system is the intended production pipeline: the client's brand profile injected at generation time, retrieval grounding over that client's knowledge base, and a compliance-aware template.

Each generation is scored in three layers, from most objective to least: deterministic checkers, a context-aware LLM judge, and an independent ragas cross-check. Generation runs on gpt-4o-mini; judging runs on gpt-4o, a stronger and different model chosen to reduce self-preference bias.

The whole loop is reproducible from a single command, and the evaluator itself is unit-tested with 27 tests, so the thing measuring quality is held to the same standard as the thing being measured.

01

Controlled comparison, same model on both arms

Baseline and system arms

Two generation paths scored on the same 32-case dataset: a naive prompt versus the governed pipeline of brand profile, retrieval, and compliance template.

02

Objective, reproducible, unit-tested

Deterministic checks

Offline, API-free checkers for banned terms with true word boundaries, disclosure coverage, and generic openers.

03

Judge model stronger than the generator

LLM-as-judge

A context-aware judge separates asserting a prohibited claim from negating it, and scores groundedness and brand voice at temperature 0.

04

Cross-validation against own interest

Independent cross-check

ragas Faithfulness on the system arm, isolated in a pinned dependency set, as a second opinion on groundedness.

05

Honesty over optics

Dual-number scorecard

Raw context-blind and judge-adjudicated rates reported side by side, with the gap explained as false positives.

Architecture

A Next.js frontend and TypeScript prompt framework sit in front of a FastAPI backend that handles brand governance, retrieval, and generation. State lives in Supabase Postgres with pgvector for embeddings, and row-level security isolates every tenant on every table.

The multi-tenant design is the point an agency or consultancy cares about: one platform serving two opposite regulatory regimes, with each client's brand rules and knowledge base scoped to that client and no other.

Generation and governance

FastAPI services for brand governance, retrieval, and generation.

  • Brand profiles: approved and banned terms, voice, compliance notes
  • Compliance-aware prompt templates
  • OpenAI generation (gpt-4o-mini)

Retrieval

Per-tenant RAG over Supabase pgvector.

  • OpenAI embeddings (text-embedding-3-small)
  • HNSW cosine index
  • Content-hash dedup on ingestion

Data and isolation

Supabase Postgres with strict tenant separation.

  • Row-level security on every table
  • Two seeded regulated tenants under one demo org
  • Version-controlled migrations

Evaluation

The evals package: dataset, arms, checkers, judges, cross-check, scorecard.

  • 32-case golden dataset
  • Three-layer scoring
  • 27 harness unit tests

Governance

Governance is enforced at generation time and verified after. At generation, the brand profile supplies approved and banned terms, voice, and compliance notes; retrieval grounds claims in the client's own knowledge base rather than the model's parametric memory.

Verification is where the project earns its name. A context-aware judge distinguishes asserting a prohibited claim from refusing one, which a naive substring checker cannot do. The system reports both the raw context-blind number and the judge-adjudicated number, so the gap between them is visible rather than hidden.

  • Banned-claim enforcement with true word boundaries, so cure does not match manicure or procedure.
  • Required-disclosure coverage checked per case against the disclosures that case demands.
  • Adversarial traps: 8 of the 32 cases are engineered to elicit a violation, and the system refused all 8.

Metrics

Against the baseline, the governed system eliminated asserted compliance violations and refused every adversarial trap, while staying transparent about where it is still weak. Under context-aware review the true violation rate fell from 0.28 to 0.00, and all 8 adversarial cases were safely reframed rather than taken.

Groundedness reached 0.95 by the system's own judge and 0.78 by the independent ragas cross-check. Both numbers are reported, not the more flattering one; the gap reflects the stricter atomic-claim decomposition ragas performs. Brand-voice adherence rose from 0.37 to 0.80.

Disclosure coverage, the honest weak point, improved from 0.27 to 0.59 and is named as the next target rather than averaged away. The system is far better at not asserting prohibited claims than at reliably including required ones, and the scorecard says so plainly.

True violation rate
0.00

Asserted prohibited claims under context-aware judge review, down from 0.28 at baseline.

Adversarial traps refused
8 / 8

Cases engineered to elicit a compliance violation, all safely reframed rather than taken.

Groundedness
0.95 / 0.78

LLM judge and independent ragas faithfulness, both reported rather than the more flattering one.

Golden test cases
32

Version-controlled cases across two regulated tenants, including 8 adversarial, with 27 tests on the evaluator itself.

Roadmap

The evaluation is the foundation that makes the rest measurable. Each subsequent phase is judged by whether it moves a number the harness already tracks.

  • Production RAG: hybrid retrieval and reranking, measured against the current 0.95 judge and 0.78 ragas groundedness.
  • Guardrails: prompt-injection defense, PII checks, and explicit OWASP-LLM and EU AI Act mapping, with disclosure coverage at 0.59 as the named target.
  • Cost and observability: per-generation tracing and model routing.
  • Path consolidation: wire vector retrieval into the primary generation endpoint and unify the two generation paths.

Phase 1

Production RAG

Hybrid retrieval with BM25 and vector, plus reranking, measured against current groundedness.

Phase 2

Guardrails

Prompt-injection defense, PII checks, and OWASP-LLM and EU AI Act mapping; disclosure coverage is the named target.

Phase 3

Cost and observability

Per-generation tracing and model routing.

Phase 4

Path consolidation

Wire vector retrieval into the primary generation endpoint and unify the two generation paths.

Reflection

The most valuable output was not the passing scores. It was the harness surfacing four real defects a finished-looking system had hidden: schema drift that meant retrieval had never run end to end, a dedup check that silently blocked reingestion of failed documents, a half-populated knowledge base, and a retrieval threshold miscalibrated for the embedding model in use.

The second lesson was methodological. A compliance checker that cannot tell an assertion from a refusal will punish careful, correct writing. The context-aware judge exists to fix exactly that, and one adversarial case that refused the trap while quoting the trap's own words proved why the layer is necessary.

Technical depth

System assumptions and operating controls.

Architecture diagram

A multi-tenant content pipeline (Next.js, FastAPI, Supabase pgvector) with brand governance and RAG grounding, evaluated by a three-layer harness: deterministic checkers, a context-aware LLM judge on a stronger model than the generator, and an independent ragas faithfulness cross-check. Generation runs on gpt-4o-mini, judging on gpt-4o. Every reported number is reproducible from a committed report.

  1. 01

    Frontend

    Next.js and a TypeScript prompt framework.

  2. 02

    Backend

    FastAPI: brand governance, retrieval, generation.

  3. 03

    Retrieval

    OpenAI embeddings into Supabase pgvector (HNSW cosine), scoped per tenant.

  4. 04

    Data

    Supabase Postgres, row-level security on every table, two seeded regulated tenants.

  5. 05

    Evaluation

    evals package: dataset, arms, checkers, judges, ragas cross-check, scorecard.

Knowledge source assumptions

Per-tenant knowledge base (brand guidelines, product and treatment docs)

Kairos

Each client's own approved documents are the ground truth for factual claims.

text-embedding-3-small

Kairos

On-topic cosine similarity runs roughly 0.3 to 0.5, so the retrieval threshold is set to 0.15 for this corpus size; Phase 1 reranking restores precision.

Evaluation metrics

True violation rate

0.00 (system), down from 0.28

Context-aware LLM judge (gpt-4o), assert versus negate

Adversarial safe-reframing

8 / 8 refused

Judge adjudication on 8 trap cases

Groundedness

0.95 judge, 0.78 ragas

LLM judge and independent ragas Faithfulness

Brand-voice adherence

0.80 (system), up from 0.37

LLM judge against a per-tenant voice descriptor

Disclosure coverage

0.59 (named gap)

Deterministic per-case disclosure-family check

Risk and failure scenarios

Context-blind false positives

A substring checker flags negations such as results are not permanent as violations.

A context-aware judge adjudicates; raw and judged numbers are both reported.

Silent brand or retrieval degradation

An empty brand block or zero retrieved chunks could score a broken pipeline as valid.

The run flags degraded generations loudly instead of scoring them silently.

Self-preference bias

A model judging its own family inflates scores.

Judging runs on gpt-4o while generation runs on gpt-4o-mini; a cross-provider judge is supported via JUDGE_MODEL.

Under-disclosed output

The system asserts fewer prohibited claims than it reliably includes required disclosures (0.59 coverage).

Named as the next guardrail target rather than hidden.

Groundedness metric optimism

A single holistic judge can read higher than atomic-claim checking.

An independent ragas cross-check of 0.78 is reported beside the judge's 0.95.

Human review checkpoints

Scorecard review before any claim is published

Kairos

Ship only numbers reproducible from a committed report

Diff review before merge to main

Founder

No fabricated metrics, clients, or history

Next step

Review the supporting profile.

Use CV access and LinkedIn for background, or return to selected work for more examples of structured AI thinking.