Context
Product catalogues need titles, descriptions, features and SEO keywords at scale. Manual creation does not scale past a few hundred rows, and a vision model can draft all four from a photograph and a catalogue record in one call.
The interesting problem is not the drafting. It is that blindly storing model output creates operational and quality risks: a listing is a database record that a storefront renders and a customer reads, and a response that merely looks plausible is not the same as a response that is safe to store.
This build treats the generation step as the least trustworthy part of the system and puts everything around it under deterministic control.
- Catalogues need four structured fields per row, at volume.
- A model response is a proposal, not a record.
- The engineering question is what happens between the response and the database.
Problem
A naive vision-model workflow fails in ways that are individually mundane and collectively expensive. It can return malformed JSON, or JSON wrapped in a markdown fence. It can exceed a field limit that a storefront silently truncates. It can fail on a transient provider error, or retry a request that can never succeed and burn budget doing it.
Two failures are worse than the rest because they are invisible. Image payloads leaking into logs turn a debugging aid into a data-retention problem. And a model can produce claims that neither the image nor the catalogue metadata establishes — a fibre content, a certification, a performance figure — which a reviewer would catch and an automated pipeline would not.
- Malformed or fenced JSON that a naive parser rejects or a lenient one mis-reads.
- Field limits exceeded: a title over 60 characters, a description outside 150–200 words.
- Transient provider errors, and retries on requests that can never succeed.
- Image payloads leaking into logs.
- Claims the image and metadata do not establish.
Workflow
The pipeline is a fixed sequence, not an agent loop. Every step is deterministic except the one model call, and each step can refuse to pass work forward.
Validation happens first and last: the input row is validated before a prompt is built, and the output is validated before anything is written. Between them, the model call is isolated behind an injectable transport so the entire pipeline is testable without a network.
- Nothing reaches the model until the catalogue row and image have passed validation.
- Nothing reaches storage until the response has passed the strict schema.
- Only safe output and metadata are written; the image payload is not logged.
01
Fails closed at the boundary
Validate the input
The catalogue row and image are validated before anything else runs: required fields present and non-blank, price positive, image decodable within type, size and dimension limits.
02
Prompt and validator cannot drift
Construct the prompt
A multimodal prompt is built from the validated product fields, with numeric requirements interpolated from the same module the validators enforce.
03
Retries only what can succeed
Send through an injected transport
The request goes through an injectable transport with selective retries, so the pipeline is fully testable offline and no socket opens in the test suite.
04
Tolerant of format, strict about content
Parse defensively
Plain JSON and fenced JSON are both handled, without the parser becoming permissive about shape.
05
Unknown fields forbidden
Reject on schema violation
Output that violates the strict ProductListing schema is rejected, unknown fields included. There is no partial acceptance.
06
No self-assessment
Score deterministically
Title length, description word count, feature count and keyword count are computed in Python. The model is never asked to grade its own output.
07
Candidates, not findings
Extract claim candidates
A keyword heuristic proposes sentences asserting things a photograph and a catalogue row usually cannot establish, for a human to adjudicate.
08
Nothing unsafe reaches storage
Write only what is safe
Validated output and run metadata are persisted. Image payloads are never logged.
Architecture
The canonical implementation is a Python package with clean layer separation: pure validation, prompt construction, parsing, quality, claims, pipeline and transport. The pure layers perform no I/O, import no provider SDK and make no network calls, which is why they can be exhaustively tested and why the constants they enforce cannot drift from the prompt that requests them.
The prompt interpolates its numeric requirements from the same module the validators read, so the text asking for a 60-character title and the check rejecting a 61-character title can never disagree.
The portfolio implementation on this site is a TypeScript presentation adapter. It ports the deterministic layers — schema validation, quality scoring, claim extraction — faithfully enough to run in the browser-facing route, and nothing more. The Python repository remains the canonical implementation; the transport, retry classification, defensive parsing, safe logging, evaluation harness and review workflow all live there.
Input boundary
Validation of the catalogue row and the uploaded image before any generation step.
- Required fields present and non-blank, price positive
- Image type, size and dimension limits enforced
- Pixels re-encoded so EXIF, including GPS, is dropped
Prompt construction
A pure module that builds the multimodal prompt and performs no I/O.
- Numeric requirements interpolated from the validator module
- Explicit instruction not to invent unobservable attributes
- Prompt text asserted directly in tests
Injected transport
The provider call isolated behind an interface the pipeline receives rather than constructs.
- Selective retries on transient errors only
- A deterministic offline fake for tests and the public demo
- Tests fail if anything opens a socket
Defensive parsing
Response handling that tolerates format variation without relaxing shape requirements.
- Plain and fenced JSON both accepted
- Strict schema with unknown fields forbidden
- A malformed response is rejected, never partially accepted
Deterministic quality and claim checks
Scoring and review preparation computed entirely in Python.
- Title, description, feature and keyword checks
- Empty-field accounting
- Claim candidates flagged by category for human review
CLI, evaluation, review and demo surfaces
The interfaces over the same validated core.
- CLI with a deterministic --mock mode
- Versioned fixture dataset and evaluation harness
- Review store for human adjudication
- Streamlit demo, deterministic by default
Governance
The public showcase is deterministic by default, and that default is the control. No API key is needed, no live model call is made from this site, and the uploaded image is previewed in the browser and never transmitted.
The honest counterpart to that safety is a limitation stated plainly in the UI: deterministic output is not evidence of live model quality. A fixture that passes every check demonstrates that the validation layer works, not that a vision model would.
- Deterministic mode is the public default; no API key is needed for the public showcase.
- The portfolio demo makes no live model calls.
- Uploaded images are previewed locally only in deterministic mode, and no image is persisted.
- The UI explains that deterministic output is not evidence of live model quality.
- The Python implementation validates file type, size, dimensions and EXIF handling.
- Live mode remains opt-in and must not be enabled in this public deployment.
- Claim candidates are review prompts, not hallucination findings.
- Unknown output fields are rejected.
- Credentials are never rendered, logged, or returned to the browser.
Metrics
The measured evidence is about the engineering, not the model. The package has 514 tests passing at 98.26% core-package coverage, with CI enforcing a 90% floor on every push so the figure cannot silently become an unenforced claim. The deterministic evaluation passes 3 of 3 quality checks.
What that 3/3 means is narrow and worth stating precisely: it is a property of the deterministic fixture, not a model result. It shows the validation layer accepts conforming output and the pipeline runs end to end without a network. It says nothing about live model quality.
The constraints the validators enforce are fixed: a title of at most 60 characters, a description of 150–200 words, 5–7 features, and 10–15 distinct keyword terms.
Four things are explicitly not measured: live model quality, human review, hallucination rate, and user or business impact. No live benchmark has been run and no reviewer has adjudicated a claim, so no rate is reported — an unmeasured number is left unmeasured rather than estimated.
- Title maximum: 60 characters.
- Description range: 150–200 words.
- Features: 5–7.
- Keywords: 10–15 distinct terms.
- Tests passing
- 514
- Core-package coverage
- 98.26%
- Deterministic evaluation
- 3 / 3
- Live model quality
- Not measured
Across the package, verified 2026-09-03.
Measured by pytest --cov; CI enforces a 90% floor on every push.
Quality checks passed offline. A property of the fixture, not a model result.
No live benchmark has been run. Hallucination rate, human review and business impact are also unmeasured.
Roadmap
The phases are ordered so that nothing is claimed before it can be measured. Phase 1 is what exists and is public; phases 2 and 3 are what would be needed before any claim about live quality could honestly be made.
Phase 1
Deterministic public showcase
What exists today: a native demo running the validation, quality and claim layers on a deterministic fixture, with no key, no live call and no image transmitted.
Phase 2
Optional server-side live generation
Live vision-model generation behind hard budget controls — per-session and per-deployment caps, explicit opt-in, and cost visible before the call is made.
Phase 3
Human review workflow and measurement
A real review study: claims adjudicated under a rubric, a hallucination rate computed over reviewed verifiable claims only, with confidence intervals rather than a point estimate.
Reflection
The project is about engineering around the model rather than treating the model response as a trusted database record. Almost everything that makes it worth showing sits outside the generation call: the validator that rejects a 61-character title, the retry classifier that refuses to re-send a request that can never succeed, the parser that handles a fenced response without becoming permissive, the logger that never sees an image payload.
The second lesson is about evidence discipline. It would be easy to present a 3/3 deterministic pass rate as a quality result, and easier still to leave the ambiguity uncorrected. Naming what is not measured — live model quality, human review, hallucination rate, business impact — costs a more impressive-looking page and buys the only thing that makes the measured numbers worth reading.
Technical depth
System assumptions and operating controls.
Architecture diagram
A Python package with pure, I/O-free layers for validation, prompt construction, parsing, quality and claims, wrapped by a pipeline that calls the provider through an injectable transport with selective retries. The input boundary validates the catalogue row and image; the output boundary enforces a strict schema with unknown fields forbidden; quality scoring and claim extraction are deterministic Python, never model self-assessment. 514 tests at 98.26% coverage with a 90% CI floor. The portfolio demo is a TypeScript presentation adapter over the same deterministic layers, running in deterministic mode only.
01
Input boundary
Catalogue-row and image validation: required fields, positive price, allowed type, size and dimension limits, EXIF dropped.
02
Prompt construction
Pure multimodal prompt building; numeric requirements interpolated from the validator module.
03
Injected transport
Provider call behind an injected interface, with selective retries and a deterministic offline fake.
04
Defensive parsing
Plain and fenced JSON handled, then checked against the strict ProductListing schema.
05
Quality and claims
Deterministic scoring in Python plus keyword-based claim-candidate extraction for human review.
06
Surfaces
CLI, evaluation harness, review store, Streamlit demo, and a TypeScript adapter for the portfolio demo.
Evaluation metrics
Test suite
514 passing
pytest, verified 2026-09-03
Core-package coverage
98.26%, floor at 90%
pytest --cov, floor enforced by CI on every push
Deterministic evaluation
3 / 3 quality checks passed
Offline fixture run; a property of the fixture, not a model result
Schema conformance
Title ≤ 60 chars, description 150–200 words, 5–7 features, 10–15 distinct keywords
Strict validation with unknown fields forbidden
Live model quality
Not measured
No live benchmark has been run
Hallucination rate
Not measured
Requires human review under a rubric; no claim has been adjudicated
Human review
Not measured
Review workflow exists; no review study has been run
User or business impact
Not measured
No deployment with users exists
Risk and failure scenarios
Malformed or fenced JSON
A naive parser rejects valid content, or a lenient one accepts the wrong shape.
Defensive parsing accepts plain and fenced JSON, then applies the strict schema unchanged.
Field limits exceeded
A title over 60 characters or a description outside 150–200 words reaches a storefront and is silently truncated.
Deterministic validators reject the response; the constants are shared with the prompt so the two cannot drift.
Hallucinated response shape
Unexpected or renamed fields are stored as if they were valid.
The schema forbids unknown fields; there is no partial acceptance.
Transient provider errors
A single network blip fails an entire catalogue batch.
Selective retries through the injected transport, on transient errors only.
Retrying an unrecoverable request
Budget spent re-sending a request that can never succeed.
Retry classification distinguishes transient from permanent failures and stops on the latter.
Image payloads in logs
Debug logging turns into an unintended data-retention and privacy problem.
Safe logging keeps image payloads out of logs entirely; uploads are re-encoded from pixels so EXIF does not survive.
Unsupported claims in copy
Copy asserts a fibre content, certification or performance figure the image and metadata do not establish.
Claim candidates are extracted for human adjudication; the system proposes, it does not judge.
Deterministic output mistaken for model quality
A passing fixture reads as evidence a vision model performs well.
The demo and the case study both state the limitation explicitly and list what is not measured.
Human review checkpoints
Claim adjudication before any hallucination rate is reported
Human reviewer
Mark each candidate supported, unsupported, not verifiable or irrelevant; unverifiable is not false
Listing approval before publication
Merchandiser
Accept, edit or reject the generated listing
Enabling live mode on any deployment
Maintainer
Opt-in only, with budget controls in place; not enabled on the public site