---
title: "Does anyone measure whether their AI affordances work?"
description: "Rarely: 7 of 19 records mention evaluation work at all."
url: "https://state-of-ai-in-design-systems.netlify.app/questions/evals.md"
canonical: "https://state-of-ai-in-design-systems.netlify.app/questions/evals.md"
type: "question"
id: "evals"
data_collected: "2026-07-26/27"
generated: "2026-07-28T06:01:02Z"
report: "State of AI in Design Systems — July 2026"
author: "Kaelig Deloumeau-Prigent"
license: "CC-BY-4.0"
citation: "Deloumeau-Prigent, K. (2026). State of AI in Design Systems. https://state-of-ai-in-design-systems.netlify.app/questions/evals.md"
---

> Snapshot of 2026-07-27. Every claim below links to the source URL it was taken from. Check the source before citing.

# Does anyone measure whether their AI affordances work?

Rarely, and that is one of the weaker spots in the field: 7 of the 19 records
mention evaluation work of any kind — [Atlassian Design System](https://state-of-ai-in-design-systems.netlify.app/systems/atlassian-design-system.md), [HeroUI](https://state-of-ai-in-design-systems.netlify.app/systems/heroui.md), [Nuxt UI](https://state-of-ai-in-design-systems.netlify.app/systems/nuxt-ui.md), [PatternFly](https://state-of-ai-in-design-systems.netlify.app/systems/patternfly.md), [Primer](https://state-of-ai-in-design-systems.netlify.app/systems/primer-github.md), [React Spectrum / Spectrum 2 (S2)](https://state-of-ai-in-design-systems.netlify.app/systems/react-spectrum-s2.md), [shadcn/ui](https://state-of-ai-in-design-systems.netlify.app/systems/shadcn-ui.md).

Published head-to-head numbers are rarer still. Most teams ship an MCP server or a skill and reason
about it from the shape of the output, not from a scored suite. It is the main reason this report
rates surface area rather than quality: there is not enough public measurement to rank anyone.
Read the 29 validation-loop techniques at
https://state-of-ai-in-design-systems.netlify.app/techniques/validation-loop.md for the closest thing the field has — checks that fail a
build, rather than evals that score a model.


Other questions this report answers, and the index of every file: https://state-of-ai-in-design-systems.netlify.app/llms.txt

---

Generated 2026-07-28T06:01:02Z from the State of AI in Design Systems — July 2026 dataset. Index of every machine-readable file: https://state-of-ai-in-design-systems.netlify.app/llms.txt. JSON, SQLite and the MCP endpoint: https://state-of-ai-in-design-systems.netlify.app/ai.md. Kaelig Deloumeau-Prigent, CC BY 4.0.
