Context
An independent benchmark of AI maturity across eight European industrial machinery manufacturers (GEA, Krones, TRUMPF, Siemens, KION, Sandvik, Heidelberger and Dürr), judged strictly on what each company has published.
Scores measure evidence of a company's public disclosure of its own AI adoption, not internal capability and not products sold. The instrument deliberately assesses each firm as an adopter of AI in its own operations, not as a vendor of AI to others.
- Public-disclosure lens: a low score means undisclosed, not incapable.
- Adopter-not-vendor scope, so it reads internal maturity rather than product catalogues.
Problem
Nearly every firm in the sector has announced something about generative AI; far fewer have shown it running in their own operations. That gap between announcement and operation is the hardest thing to read from the outside, and the trade press rarely separates the two.
Without a fixed standard, any ranking of AI maturity collapses into whoever markets loudest. The benchmark exists to make the judgement explicit, contestable and reproducible.
- No comparable way to read own-operations maturity from mixed corporate disclosure.
- Vendor and customer figures crowd out evidence of internal adoption.
- The weighting of what maturity means is usually hidden inside the score.
Workflow
The rubric consists of six weighted dimensions, each scored 0 to 5 against anchored descriptors. It was committed to git before any company was assessed; the order is visible in the history.
Each company was searched across the same seven source classes, from annual reports to job advertisements, with searches that returned nothing recorded as evidence of absence. Every score carries a direct quotation, a source URL, a date and a document type; a score with no evidence entry is recorded as 0, never inferred.
- Tie-break rule: where evidence sits between two levels, the lower is recorded.
- Assessment window fixed to January 2024 onward, with a mandatory current-year recency check.
- Weighted leaderboard computed by a deterministic engine and published to an interactive explorer.
Architecture
The system is a reproducible scoring pipeline, not an agent. Per-company evidence files feed a scoring engine that validates them and computes weighted 0-5 scores against the fixed rubric; an automated test suite covers the arithmetic and the real data; a report generator recomputes the leaderboard from evidence alone.
A Streamlit explorer publishes the result and lets any reader substitute their own dimension weights and watch the ranking move, so a conclusion that survives only under one weighting is visible rather than hidden.
- Evidence files → scoring engine → 34 tests → `python -m src.report` → Streamlit explorer.
Governance
Beyond scoring, the project specifies one process, aftersales spare-parts quoting, as an agentic redesign in which the points where a human can understand, intervene and override are fixed before the automation. Release to the customer is a human-only step; every quote stays a human-accountable act. It is written as a design specification to satisfy EU AI Act Article 14, not as an implementation.
The work is fully independent (no company was contacted or consulted) and published openly: code under MIT, data and analysis under CC BY 4.0.
- Article 14 oversight located at the consequential step, not distributed vaguely.
- Every score auditable to its source; corrections retained, not overwritten.
Metrics
Across all eight companies, and two full search passes, not one publishes a quantified result from its own AI use: the mean disclosed-outcomes score is 0.75 and no company scores above one. The finding survived a targeted re-run that hunted specifically for internal outcomes.
Governance and deployment are decoupled: KION runs a generative-AI tool in production yet discloses no AI governance, while Krones has the strongest governance in the set around a use case that stayed a pilot. The distribution is two-tier: six firms between 23% and 45%, then a cliff to 11%.
- ~40 vendor-evidence items excluded under the adopter-not-vendor rule.
- Two errors caught on re-run and corrected in the open: Siemens 31→39%, GEA 39→45%.
- Every published number recomputable from the evidence with one command.
- Companies benchmarked
- 8
- Weighted dimensions
- 6
- Publish a quantified internal-AI outcome
- 0
European industrial machinery manufacturers, scored from public disclosure.
Data, process, agentic, governance, workforce, and disclosed outcomes.
Across all eight. The sector runs on AI but has not begun to measure what it is worth to itself.
Roadmap
The rubric and evidence standard are sector-agnostic; the natural next steps are breadth, robustness and freshness.
- Extend the peer set and carry the method into adjacent sectors.
- Add a second assessor to introduce an inter-rater reliability check.
- Re-run on a fixed cadence to track how disclosure moves.
Reflection
The most useful lesson was methodological: a benchmark's credibility does not come from being right the first time, it comes from the corrections being visible. The framework was amended in the open during the work, and two of its own scoring errors were caught and retained alongside the corrections.
Fixing the rubric before the data, and exposing the weighting as a judgement a reader can overrule, is what lets the result be argued with precisely rather than dismissed wholesale.
Technical depth
System assumptions and operating controls.
Architecture diagram
A reproducible scoring pipeline rather than an agentic system: validated per-company evidence, deterministic weighted scoring against a fixed rubric, an automated test suite, and a public explorer that lets a reader reweight the dimensions.
01
Evidence files
One structured file per company, each score tied to a sourced, dated quotation.
02
Scoring engine
Validates the evidence and computes weighted 0-5 scores against the fixed rubric.
03
Test suite
34 pytest tests across validation, arithmetic, ranking and the real dataset.
04
Report generator
`python -m src.report` recomputes the leaderboard from the evidence, reproducibly.
05
Streamlit explorer
Public app where a reader reweights the six dimensions and watches the ranking move.
Evaluation metrics
Reproducibility
Leaderboard recomputable from evidence alone
`python -m src.report`, backed by 34 pytest tests
Ranking robustness
Conclusions survive alternative weightings
Interactive weight-sensitivity in the explorer
Evidence traceability
Every score backed by a sourced, dated quotation
One evidence file per company; source hierarchy in METHODOLOGY §4
Self-correction
Errors surfaced and fixed in the open
Two documented re-run corrections (Siemens 31→39%, GEA 39→45%), superseded findings retained
Risk and failure scenarios
Single assessor, no inter-rater check
Scores reflect one person's judgement
Anchored descriptors and published evidence let any reader dispute a specific score
Disclosure bias
Measures communication practice as much as maturity
Scope stated explicitly; the adopter-not-vendor rule keeps it about own-operations disclosure
Language coverage
Only German and English sources reviewed
Stated as a limitation; assessment window recorded per company
No independent verification
A published claim is evidence it was made, not that it is true
Every score is sourced so a reader can verify it directly
Point-in-time snapshot
Disclosure in this area moves quickly
Scores dated; a refresh cadence is planned