A longtime skeptic finds AI finally accelerating his research
Two years earlier, he and his collaborator found new models repeatedly disappointing on their actual mathematics and programming problems. More recently, parts of the workflow crossed a useful threshold.
Two mathematical bounty problems now appear substantially resolved with help from language models and Lean. Wentworth puts his confidence at about 80%, and explicitly says he has not yet examined one of the proofs himself. That qualification belongs beside the achievement.
For interpretability experiments, Claude Code now handles his routine coding needs well enough that he rarely has to write the code. Interpreting results and suggesting worthwhile next steps remain much less useful.
Zvi highlights this report because it describes a change in someone’s real research practice. The productivity gain is specific: implement experiments and help search for checkable proofs, while the researcher retains the work of judging significance and deciding where to look next.