<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Lens Frontier</title><description>把真实任务、验证协议和评测证据沉淀成可追溯的数据产品。</description><link>https://lens-frontier.github.io/</link><item><title>一个客观的代码重构评估体系是如何构成的</title><link>https://lens-frontier.github.io/zh/blog/refactor-bench-eval/</link><guid isPermaLink="true">https://lens-frontier.github.io/zh/blog/refactor-bench-eval/</guid><description>以行为保持和结构静态检查为核心，用真实开源代码提交做数据、隔离环境测试、用 tree-sitter 做语法树匹配，避开大模型当裁判带来的主观漂移，精准测量大模型的代码重构能力边界。</description><pubDate>Tue, 16 Jun 2026 00:00:00 GMT</pubDate></item><item><title>为什么 2026 年最强 coding agent 还在用 grep？</title><link>https://lens-frontier.github.io/zh/blog/agent-harness-search-research/</link><guid isPermaLink="true">https://lens-frontier.github.io/zh/blog/agent-harness-search-research/</guid><description>梳理 coding agent harness 的代码库检索路线：grep、语义索引、LSP、代码地图、Search Subagent 与工具检索如何走向混合检索。</description><pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate></item><item><title>题面写得再绕，模型照样秒解——怎么让搜索题真正变难</title><link>https://lens-frontier.github.io/zh/blog/browsecomp-query-hardening/</link><guid isPermaLink="true">https://lens-frontier.github.io/zh/blog/browsecomp-query-hardening/</guid><description>搜索智能体训练需要难度适中的题：模型不能秒解，但认真搜索可解。我们设计了 Hardener-Solver-Critic 多智能体闭环，用模型真实求解行为增强题目难度，在 BrowseComp 风格合成题上将 62% 的样本推入了可训练区间。</description><pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate></item><item><title>如何构建面向文生图模型的多模块协同评估框架</title><link>https://lens-frontier.github.io/zh/blog/t2i-multi-module-evaluation/</link><guid isPermaLink="true">https://lens-frontier.github.io/zh/blog/t2i-multi-module-evaluation/</guid><description>以 LLM-as-a-Judge 为核心，通过 OCR/VQA/CLIP/IQA/HP 五模块协同校准，突破单一指标 benchmark 局限，全面捕捉现代文生图模型的真实能力边界。</description><pubDate>Thu, 21 May 2026 00:00:00 GMT</pubDate></item></channel></rss>