Claude beats Codex at reproducing social science

AI

by Dylan Matthews · about work by Meysam Alizadeh, Mohsen Mosleh, Fabrizio Gilardi, Joshua Tucker

This year has seen AI agents break through in a serious way among social scientists I respect. So it was only a matter of time until social scientists designed an AI benchmark to see how good agents really are. Meysam Alizadeh, Mohsen Mosleh, Fabrizio Gilardi, and Joshua Tucker’s SocSci-Repro-Bench tested if Claude Code and Codex could reproduce 54 distinct papers. The most striking result to me is how different the models’ capabilities were: Claude could fully reproduce 78 percent of papers accurately, while Codex only got 35.8 percent. What’s more, the study found evidence the results were not memorized by the models; they were doing the work from scratch. Funny enough, OpenAI’s success rate here is similar to that of GPT 4/4o back when the Institute for Replication tested those models in 2024 (though it’s not clear to me if the replication tasks in the two papers are similarly challenging).