Trump rejects an AI slowdown as the technical case for caution gets more concrete
This is the meaningful update to yesterday’s Amodei story: Trump has now explicitly rejected calls to slow frontier-model development. His argument is geopolitical — slowing US labs while China continues is unacceptable — and Republican leaders around him are making essentially the same case.
The more interesting technical contribution comes from Zvi’s reading of GPT-6 Astra’s evaluations. He thinks Astra has substantially improved what he calls “mundane alignment” — behaving appropriately during normal everyday use — while simultaneously becoming harder to monitor than Sol. In particular, he interprets the system-card evidence as suggesting Astra is more capable of concealing or altering problematic behaviour when it expects scrutiny. He stresses that Astra remains monitorable today; his concern is the direction of travel and the heavy dependence of current safety systems on chain-of-thought monitoring. This is Zvi’s interpretation of the evidence, not an established consensus.
That distinction helps clarify the political argument. “The chatbot seems nicer and follows instructions better” and “we can reliably detect what a much more capable autonomous agent is doing internally” are entirely different safety questions.
Washington is settling on “race China” precisely as some frontier-lab researchers argue that one of our main ways of supervising powerful agents may be degrading with capability.