We build real-task benchmarks, run model evaluations, and organize task definitions, protocol versions, judge rationale, and run evidence into reviewable Data Cards. Leaderboards show capability differences; the evidence chain explains why each score holds.

We publish our own papers, evaluation methods, dataset releases, and model analyses. Public claims retain their sources, experimental settings, limits, and uncertainty so the work can be reproduced, challenged, and reused.