Skip to main content
196

Search Lumen

Find components, APIs, guides, and recipes.

GitHub
Reproducible evaluation

Does a UI library save AI tokens? Measure the whole task.

Reusable components can reduce the code an agent writes. Skills, documentation, and tool calls also consume context. A useful comparison counts both, holds the product requirements steady, and checks whether the result works.

By Santiago Molina ·

Our first comparison did not show token savings.

On these two small React tasks, both Lumen approaches used more total tokens than building from scratch. We used gpt-6-astra with xhigh reasoning effort and codex-cli 0.159.0-alpha.12.1, with 3 fresh runs per task and approach. 18 of 18 runs passed the host checks within 2 attempts.

October 4, 2026 · Median total tokens include repairs and cached input.
TaskApproachMedian tokensPassedFirst attempt
Profile dialogFrom scratch116,5213/33/3
Profile dialogLumen + documentation514,6653/30/3
Profile dialogLumen + skill/MCP692,1513/30/3
Notification settingsFrom scratch121,7403/32/3
Notification settingsLumen + documentation254,0193/33/3
Notification settingsLumen + skill/MCP463,3573/31/3

All six Lumen dialog runs needed a repair to satisfy this protocol's strict within-dialog Tab and Shift+Tab containment check. The totals include that work; they do not isolate documentation lookup overhead. This sample does not support a claim of fewer AI tokens or lower bills.

These runs used a local v4 candidate based on 972348de. Source hashes identify the measured files, including candidate changes. Elapsed times were recorded on a shared development machine and are not isolated speed measurements. The sample covers one model, two tasks, and three repetitions per approach.

Download all 18 results as JSON for input, cached input, output, tool calls, elapsed times, attempts, prompts, and source hashes. Cached input is a subset of input; the report never adds it twice. The measured source archive preserves the harness, verifier, token parser, stylesheet, and MCP catalog at their recorded hashes.

Compare three ways to build the same interface.

  1. From scratch: React with native HTML controls and custom CSS.
  2. Lumen with documentation: public React components, installed types, and the package README.
  3. Lumen with skill and MCP: the same components and documentation, plus the portable skill and local component catalog.

The initial fixtures cover a profile dialog and a notification settings screen. Every approach gets the same behavior, responsive, keyboard, and accessibility requirements. This is a bounded React comparison; it does not establish results for other frameworks, native platforms, competitors, or every kind of application.

Count everything needed to complete the task.

  • Input tokens: include instructions, read documentation, tool context, and subsequent turns reported by the client.
  • Output tokens: use the client's completed-turn usage, including the work in repair attempts.
  • Cached input: report it separately as a subset of input; never add it twice.
  • Failures and repairs: retain unsuccessful runs. Each approach gets at most one independently checked repair.
  • Time and quality: report elapsed time, success rate, and the same verification checks alongside token totals.

Total tokens are input plus output, not a monetary price. Caching and billing vary by provider. Missing usage stays unavailable; it must never become a zero-token result. The host's local compilation and browser checks are included in elapsed time but consume no model tokens. This protocol does not ask agents to run their own test suite.

Keep the comparison controlled.

Pin the model, reasoning effort, client version, package snapshot, skill, and prompts. Start each repetition in a fresh fixture and rotate the approach order. Run at least three repetitions per task and approach, then report per-task medians and all run outcomes. Retain the provider's cache measurements: a fresh directory does not guarantee a cold model cache.

The verifier checks strict TypeScript, the assigned component approach, responsive overflow, browser errors, labels, keyboard behavior, preserved input, and automated WCAG checks at 390px and 1440px. Automated checks do not replace visual judgment or assistive-technology testing.

Run the comparison locally.

Clone the repository, follow its contributor setup, and use an authenticated Codex CLI. This opt-in evaluation consumes account usage. The default matrix has 18 runs and permits one repair per run.

bashRun from the Lumen repository
pnpm install --frozen-lockfile
pnpm --filter @santi020k/lumen-react... run build
pnpm --filter @santi020k/lumen-mcp run build
pnpm run test:ai-efficiency
pnpm run eval:ai-efficiency --model <model-id> --effort <reasoning-effort> \
  --repetitions 3 --output /absolute/path/outside-the-repository

Each run saves prompts, transcripts, synthetic output, screenshots, and token measurements outside the repository. results.json records the comparison and source hashes. Read the full protocol before interpreting or publishing results. Rechecking saved output is verification, not a new generation.

Apply the result to your own workflow.

Start with one representative task and a quality bar your team can verify. If discovery overhead outweighs reusable code on a small screen, record that result too. Prefer focused usage retrieval and reuse stable context where your client supports it. More tool calls do not automatically mean better output.

Build the example before measuring