The Astera metascience essay contest
The most interesting thing I read recently was the winning entrants in the Astera metascience essay contest, for which I was one of the judges. While not every winning essay was related to AI (see these two, for example), the opportunities and challenges of AI dominated the contest:
It’s a bit of a truism that AI and data are strongly complementary; but what data should you collect? What if we could use scaling laws as an input into this decision process? Collect data and train models on it, until we can see the trendlines emerge, which will let us estimate how much data we would need to collect to achieve a given performance (more from Peter Koo’s winning essay).
A major problem in science is the “file drawer problem”, wherein null results don’t get published, in part because writing up any result takes time and energy, and the rewards to doing so are very low for academics (more here). But if AI agents are integrated into the research process, they have all the material they need to do this for us at almost no cost. Could we finally open all (well, more) file drawers? (more on this theme from Niveditha Iyer’s winning essay)
One of the most important problems in science is knowing what question to study. With AI that can finally parse text as good as most people, we are in a position to map out the enormous network of cause and effect that the collective literature has studied. With a map in hand, could we identify promising gaps we would otherwise miss? Or could we at least make progress on automating the identification of important questions? (more from Prashant Garg’s winning essay)
The use of benchmarks to assess the quality of different AI models has been a notable hallmark of AI progress, and one that seems likely to spread across science more generally. But how do you design good benchmarks? Shaamil Karim and Jaeeon Lee both proposed ideas related to this.
One more interesting AI-related challenge: when we ran the winning essays through Pangram, some of them were flagged as using AI (though the share that used AI was even higher among the ones that had been triaged out, prior to the AI check).
I think all future essay contests are going to need to think hard about an AI policy. In our case, we cared about the arguments and ideas, not the authenticity of the prose (or even the provenance of the idea?), and so we decided not to factor AI use into our decisions. But questions remain: what’s the right size for a prize, when the cost of writing an essay has declined? To ensure the prize is incentivizing ideas that outcompete what you would get by simply asking an LLM for ideas, maybe contest organizers of the future should use AI to generate essays in response to the contest announcement and covertly submit them; a prize is only awarded if the judges pick essays that were not generated this way.
- https://asterainstitute.substack.com/p/what-scientists-said-results-from
- https://matthewleighton.substack.com/p/catalyzing-multidisciplinary-frontier
- https://docs.google.com/document/d/1n4_f2O2DkHFUQ3dbbCxZnKw13A-lgwZAe9Fte3FUnRs/edit?usp=sharing
- https://koo-lab.github.io/before-the-next-pdb/
- https://www.newthingsunderthesun.com/pub/i0wovii0/release/10
- https://nivedithasi.github.io/post.html?slug=negative_result
- https://docs.google.com/document/d/1ZqPEIUN1CvD6YiTCOz_INvn8TKYIXkI1IsL-yBLRsVA/edit?usp=sharing
- https://docs.google.com/document/d/1K9hjokYgLuzP4LmsAhuEvJnD7t-SAu2lN4teWnkzFuE/edit?usp=sharing
- https://substack.com/home/post/p-195911734
- https://www.pangram.com/