Skip to content
Lucas BarriosApplied AI & Operational Transformation

Independent research

AI Reinvention Benchmark: European Industrial Machinery

An AI-maturity benchmark of eight European machinery manufacturers, scored only on public disclosure against a rubric fixed before assessment. The finding: across all eight, not one publishes a quantified result from its own AI use.

Context

An independent benchmark of AI maturity across eight European industrial machinery manufacturers (GEA, Krones, TRUMPF, Siemens, KION, Sandvik, Heidelberger and Dürr), judged strictly on what each company has published.

Scores measure evidence of a company's public disclosure of its own AI adoption, not internal capability and not products sold. The instrument deliberately assesses each firm as an adopter of AI in its own operations, not as a vendor of AI to others.

  • Public-disclosure lens: a low score means undisclosed, not incapable.
  • Adopter-not-vendor scope, so it reads internal maturity rather than product catalogues.

Problem

Nearly every firm in the sector has announced something about generative AI; far fewer have shown it running in their own operations. That gap between announcement and operation is the hardest thing to read from the outside, and the trade press rarely separates the two.

Without a fixed standard, any ranking of AI maturity collapses into whoever markets loudest. The benchmark exists to make the judgement explicit, contestable and reproducible.

  • No comparable way to read own-operations maturity from mixed corporate disclosure.
  • Vendor and customer figures crowd out evidence of internal adoption.
  • The weighting of what maturity means is usually hidden inside the score.

Workflow

The rubric consists of six weighted dimensions, each scored 0 to 5 against anchored descriptors. It was committed to git before any company was assessed; the order is visible in the history.

Each company was searched across the same seven source classes, from annual reports to job advertisements, with searches that returned nothing recorded as evidence of absence. Every score carries a direct quotation, a source URL, a date and a document type; a score with no evidence entry is recorded as 0, never inferred.

  • Tie-break rule: where evidence sits between two levels, the lower is recorded.
  • Assessment window fixed to January 2024 onward, with a mandatory current-year recency check.
  • Weighted leaderboard computed by a deterministic engine and published to an interactive explorer.

Architecture

The system is a reproducible scoring pipeline, not an agent. Per-company evidence files feed a scoring engine that validates them and computes weighted 0-5 scores against the fixed rubric; an automated test suite covers the arithmetic and the real data; a report generator recomputes the leaderboard from evidence alone.

A Streamlit explorer publishes the result and lets any reader substitute their own dimension weights and watch the ranking move, so a conclusion that survives only under one weighting is visible rather than hidden.

  • Evidence files → scoring engine → 34 tests → `python -m src.report` → Streamlit explorer.

Governance

Beyond scoring, the project specifies one process, aftersales spare-parts quoting, as an agentic redesign in which the points where a human can understand, intervene and override are fixed before the automation. Release to the customer is a human-only step; every quote stays a human-accountable act. It is written as a design specification to satisfy EU AI Act Article 14, not as an implementation.

The work is fully independent (no company was contacted or consulted) and published openly: code under MIT, data and analysis under CC BY 4.0.

  • Article 14 oversight located at the consequential step, not distributed vaguely.
  • Every score auditable to its source; corrections retained, not overwritten.

Metrics

Across all eight companies, and two full search passes, not one publishes a quantified result from its own AI use: the mean disclosed-outcomes score is 0.75 and no company scores above one. The finding survived a targeted re-run that hunted specifically for internal outcomes.

Governance and deployment are decoupled: KION runs a generative-AI tool in production yet discloses no AI governance, while Krones has the strongest governance in the set around a use case that stayed a pilot. The distribution is two-tier: six firms between 23% and 45%, then a cliff to 11%.

  • ~40 vendor-evidence items excluded under the adopter-not-vendor rule.
  • Two errors caught on re-run and corrected in the open: Siemens 31→39%, GEA 39→45%.
  • Every published number recomputable from the evidence with one command.
Companies benchmarked
8

European industrial machinery manufacturers, scored from public disclosure.

Weighted dimensions
6

Data, process, agentic, governance, workforce, and disclosed outcomes.

Publish a quantified internal-AI outcome
0

Across all eight. The sector runs on AI but has not begun to measure what it is worth to itself.

Roadmap

The rubric and evidence standard are sector-agnostic; the natural next steps are breadth, robustness and freshness.

  • Extend the peer set and carry the method into adjacent sectors.
  • Add a second assessor to introduce an inter-rater reliability check.
  • Re-run on a fixed cadence to track how disclosure moves.

Reflection

The most useful lesson was methodological: a benchmark's credibility does not come from being right the first time, it comes from the corrections being visible. The framework was amended in the open during the work, and two of its own scoring errors were caught and retained alongside the corrections.

Fixing the rubric before the data, and exposing the weighting as a judgement a reader can overrule, is what lets the result be argued with precisely rather than dismissed wholesale.

Technical depth

System assumptions and operating controls.

Architecture diagram

A reproducible scoring pipeline rather than an agentic system: validated per-company evidence, deterministic weighted scoring against a fixed rubric, an automated test suite, and a public explorer that lets a reader reweight the dimensions.

  1. 01

    Evidence files

    One structured file per company, each score tied to a sourced, dated quotation.

  2. 02

    Scoring engine

    Validates the evidence and computes weighted 0-5 scores against the fixed rubric.

  3. 03

    Test suite

    34 pytest tests across validation, arithmetic, ranking and the real dataset.

  4. 04

    Report generator

    `python -m src.report` recomputes the leaderboard from the evidence, reproducibly.

  5. 05

    Streamlit explorer

    Public app where a reader reweights the six dimensions and watches the ranking move.

Evaluation metrics

Reproducibility

Leaderboard recomputable from evidence alone

`python -m src.report`, backed by 34 pytest tests

Ranking robustness

Conclusions survive alternative weightings

Interactive weight-sensitivity in the explorer

Evidence traceability

Every score backed by a sourced, dated quotation

One evidence file per company; source hierarchy in METHODOLOGY §4

Self-correction

Errors surfaced and fixed in the open

Two documented re-run corrections (Siemens 31→39%, GEA 39→45%), superseded findings retained

Risk and failure scenarios

Single assessor, no inter-rater check

Scores reflect one person's judgement

Anchored descriptors and published evidence let any reader dispute a specific score

Disclosure bias

Measures communication practice as much as maturity

Scope stated explicitly; the adopter-not-vendor rule keeps it about own-operations disclosure

Language coverage

Only German and English sources reviewed

Stated as a limitation; assessment window recorded per company

No independent verification

A published claim is evidence it was made, not that it is true

Every score is sourced so a reader can verify it directly

Point-in-time snapshot

Disclosure in this area moves quickly

Scores dated; a refresh cadence is planned

Next step

Review the supporting profile.

Use CV access and LinkedIn for background, or return to selected work for more examples of structured AI thinking.