Skip to content
Back to work

AI Model Evaluation / Public Evidence Lab / Case study

BuildArena

A self-initiated public lab that compares AI-built websites using shared challenges and inspectable evidence rather than unsupported model claims.

Client type
Self-initiated AI evaluation product
Version
Live External Project / v1
Format
external
Data
external

Problem

AI coding evaluations are difficult to trust when the prompt, implementation, review process, and limitations are hidden.

Approach

BuildArena exposes the method, prompt archive, challenge queue, build records, comparison surfaces, and publication rules as one public workflow.

Evidence boundary

  • Self-initiated external project. Current public site shows Run 001 in capture; scores and verdicts remain withheld pending evidence review.

System scope

What was built

  1. 01Shared challenge standard
  2. 02Evidence archive
  3. 03Scoped comparison
Open the working demo

Version history

Build notes

Open current version
  1. v1

    2026-08-12

    Public evidence lab

    Present BuildArena as a real external flagship project in the portfolio.

    • Added an external project entry to the portfolio demo registry.
    • Linked the live BuildArena deployment without copying its application code.
    • Recorded current capture state without inventing scores or verdicts.
    Review
    2026-09-12
    Target
    Give visitors a direct path from the portfolio to the complete BuildArena product.