Independent model evaluation

Compare model capability where real work happens.

Lens Frontier builds real benchmarks, runs models, and audits the evidence. We publish not only who leads, but why each score holds.

Explore all benchmarks
1 benchmark published

Benchmark model rankings

Choose one benchmark to see what it measures, who leads, and why the result is trustworthy.

All benchmarksCode-QA-Bench
Swipe through published evaluations; select one to switch
Model rankingHigher scores rank first4 models
Benchmark system

Every score resolves to a traceable Data Card.

The leaderboard reveals differences. The Data Card defines task scope, provenance, judgment, and status. They are two views of the same evaluation project.

Research & releases

Team papers, evaluation methods, benchmark releases, and model analysis.

View all research
Evidence standard

Trusted scores require a complete evidence chain.

Task, protocol, run, and judgment must remain mutually traceable.

View research publications
  1. 01
    Scope the task

    Define repository, environment, tools, and success.

  2. 02
    Freeze protocol

    Fix resources, constraints, judge, and version.

  3. 03
    Capture evidence

    Retain outputs, artifacts, traces, and exceptions.

  4. 04
    Review and publish

    Separate verified, reviewed, and pending results.