Refine sets the benchmark for AI review
There are a growing number of AI tools for various components of the scientific process, but there are rarely careful, public evaluations of how they perform and compare. Last week Refine, an AI-powered review tool for economics research, published the kind of benchmark and evaluation study that I would love to see more of in this space. They ran ~1300 head-to-head matches on 150 economics preprints, comparing the quality of Refine’s reviews to both single-shot frontier LLM reviews and scaffolded review systems. Refine won roughly 90% of the match-ups (~95% against single-shot models and ~85% against scaffolds). Worth caveating that the benchmark’s unit of evaluation is a “paper-grounded atomic concern”, a specific verifiable issue tied to a location in the paper; this is a reasonable operationalization of review quality, but is also one that informed Refine’s design. Refine’s win rate is also strongest in matchups where it raises more unique concerns than the competitor (~90-95% when it has more concerns vs ~70% when it has fewer) - that could very well be signal, but it could also be an evaluator bias towards longer/more detailed reviews (a preference that has been observed in human evaluators), which would be both useful and interesting to check. But overall, it seems clear that one major takeaway is “Refine is very good at producing high-quality, technical feedback, and if you’re an economist you should consider giving it a try”, and another is “more tool-builders should run and publish studies like this.”