Independent model evaluation
Compare model capability where real work happens.
Lens Frontier builds real benchmarks, runs models, and audits the evidence. We publish not only who leads, but why each score holds.
Explore all benchmarks1 benchmark published
Benchmark model rankings
Choose one benchmark to see what it measures, who leads, and why the result is trustworthy.
Filter benchmarks and models1 Bench · 4 models
Filter benchmarks
1/1
Filter models
4/4
All benchmarksCode-QA-Bench
Swipe through published evaluations; select one to switch Selected4/6
No result
Every score resolves to a traceable Data Card.
The leaderboard reveals differences. The Data Card defines task scope, provenance, judgment, and status. They are two views of the same evaluation project.
Research & releases
Team papers, evaluation methods, benchmark releases, and model analysis.
Evidence standard
Trusted scores require a complete evidence chain.
Task, protocol, run, and judgment must remain mutually traceable.
View research publications- 01Scope the task
Define repository, environment, tools, and success.
- 02Freeze protocol
Fix resources, constraints, judge, and version.
- 03Capture evidence
Retain outputs, artifacts, traces, and exceptions.
- 04Review and publish
Separate verified, reviewed, and pending results.