Pull to refresh
Zvi Mowshowitz
22 Sept 2026
Zvi Mowshowitz

Politics Gets Interested In Those Trying Not To Die

This was the month the world took notice that AI might kill everyone. Jacob Coxon’s resignation set off a preference cascade . Anthropic CEO Dario Amodei wrote that we must pace the frontier . Sam Altman, Elon Musk and Demis Hassabis agreed

Zvi Mowshowitz
20 Sept 2026
Zvi Mowshowitz

Better Call Sol Or Better Yet Claude or Astra

What should your AI lawyer do for you? Should you be worried that your AI lawyer , or other AI, will put the Claude constitution, the OpenAI Model Spec or some sense of law, morality, ethics or common decency above its loyalty to you? Are t

Zvi Mowshowitz
19 Sept 2026
Zvi Mowshowitz

Anthropic Looks At Some Of Its Alignment Problems

Anthropic has given us its assessment of four ‘recent cybersecurity incidents’ involving Claude that happened during cybersecurity evaluations, three of which were previously known. The report excludes the incident reported by UK AISI . The

Zvi Mowshowitz
18 Sept 2026
Zvi Mowshowitz

The Preference Cascade Is Only Getting Started

We are in the midst of a preference cascade about existential risk from AI. A preference cascade is, alas, the best method we have to change the debate. The avalanche has started . There is still time for the pebbles to vote . For now. Mike

Zvi Mowshowitz
17 Sept 2026
Zvi Mowshowitz

AI #186: The World Takes Notice

In the wake of Jacob Coxon’s resignation , and the resulting preference cascade , things have escalated quickly. The mainstream media picked it up. Anthropic CEO Dario Amodei came out and said We Must Pace the Frontier , promising to take t

Zvi Mowshowitz
16 Sept 2026
Zvi Mowshowitz

Trump Goes Full Hoax on AI Existential Risk

This is our reality. I suppose we have to talk about it. Everyone in a position to know is freaking out about AI potentially killing everyone this decade and wants to pace the frontier, and people are finally listening. It only took a few d

Zvi Mowshowitz
15 Sept 2026
Zvi Mowshowitz

The Bad Guy With An AI Named Claude

A lot of bad guys try to use Claude to do bad things. Mostly they fail. We think. Anthropic has disrupted a bunch of them, and offers an extensive report . If Anthropic is sharing the worst cases, or anything close to them, things are actua

Zvi Mowshowitz
14 Sept 2026
Zvi Mowshowitz

We Must Pace The Frontier

Dario Amodei has a new essay that finally says the thing: We Must Pace the Frontier , naming his call after the Pacing the Frontier letter lab employees signed in July . As in, we need to slow the rate at which AIs increase their capabiliti

Zvi Mowshowitz
11 Sept 2026
Zvi Mowshowitz

The Extinction Risk Preference Cascade: Quotes

These are quotes from OpenAI, Anthropic and Google employees, in the wake of Jacob Coxon’s warnings, in which the employees confirm that they think AI might soon kill everyone. If more quotes come in over the next week or so, I will update

Zvi Mowshowitz
11 Sept 2026
Zvi Mowshowitz

Jacob Coxon Warns of Human Extinction and Triggers a Preference Cascade

CEOs of major AI labs, and employees of major AI labs, including OpenAI and Anthropic, often say they plan to build superintelligence soon, as in within a few years create AIs that are superior to humans at essentially all cognitive tasks.

Zvi Mowshowitz
9 Sept 2026
Zvi Mowshowitz

GPT-6 Astra: The System Card, Alignment and What Comes Next

OpenAI claims that Astra is ‘the most intelligent and most aligned [available] model’ in the world. Not the most intelligent and aligned OpenAI model, but the most period. That is bold talk. It risks overstepping, and by doing so souring th

Zvi Mowshowitz
8 Sept 2026
Zvi Mowshowitz

Astra Is Hard to Monitor

OpenAI’s central message on Astra is that it is three things: Highly capable and can do all the things for you. Hard to monitor. The most aligned model. The first claim largely checks out. Astra and Fable are both clearly excellent models.

Zvi Mowshowitz
7 Sept 2026
Zvi Mowshowitz

An Alien Mind: Jakub Pachocki Warns Us

OpenAI Chief Scientist Jakub Pachocki is dropping truth bombs. Tomorrow I will discuss Astra’s lack of monitorability, and the potential contributing factors to that. The situation is alarming and should freak you out, and briefly it looked

Zvi Mowshowitz
3 Sept 2026
Zvi Mowshowitz

AI #184: Post Post Mortem

I am exhausted. We may finally be nearing the end of direct coverage of What Happened with the attack on HuggingFace, and the subsequent near term reactions. That took up a full five posts in the last week: OpenAI Offers Straight-Laced Post

Zvi Mowshowitz
2 Sept 2026
Zvi Mowshowitz

Anthropic Has Some Alignment Problems

Oh, good. They noticed . Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unau

Zvi Mowshowitz
27 Aug 2026
Zvi Mowshowitz

AI #183: Pre Post Mortem

Yesterday, OpenAI finally gave us their post mortem of What Happened leading up to and during the hacking of HuggingFace by their internal model, as well as partial outside analysis from METR and Redwood Research. The reports are a doozy. I

Zvi Mowshowitz
26 Aug 2026
Zvi Mowshowitz

Against Modesty’s Bailey

Modesty arguments often say that you should mostly or entirely bow to ‘expert consensus’ or the views of particular others, and who are you to disagree. It has been a few years since I’ve properly addressed this so: My answer is that you ar

Zvi Mowshowitz
20 Aug 2026
Zvi Mowshowitz

AI #182: Pause For Reflection

This was a week of quiet aftermath, an opportunity to process recent events and start to figure out the path forward. OpenAI is attempting to turn its ship around. Investors are questioning the turnover in its C-suite, but the bigger proble

Zvi Mowshowitz
15 Aug 2026
Zvi Mowshowitz

On Dwarkesh Patel’s Podcast With Ryan Greenblatt

Some podcasts are self-recommending enough that I look to break them down if I have the chance. This, as a debate about recursive self-improvement, was one of those. So here we go. The vibes have shifted, contrast this to the lit recursion

Zvi Mowshowitz
13 Aug 2026
Zvi Mowshowitz

AI #181: Astra Goes Cyber Critical

The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters. It turns out that OpenAI Trained Its Models For Months While Those Mod

Zvi Mowshowitz
13 Aug 2026
Zvi Mowshowitz

AI #181: Astra Goes Cyber Critical

The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters. It turns out that OpenAI Trained Its Models For Months While Those Mod

Zvi Mowshowitz
12 Aug 2026
Zvi Mowshowitz

Monthly Roundup #45: August 2026

As AI has escalated increasingly quickly, more and more of my posts have ended up focusing on AI. This past month, with the hacking incidents at OpenAI and elsewhere, that has hit the limit, where if you count Lightcone Commons then every s

Zvi Mowshowitz
5 Aug 2026
Zvi Mowshowitz

The Three AI Pills

Sincere disagreements about AI are usually disagreements about future AI capabilities. There are roughly four positions people take. Two are reasonable. Two are not. I distinguish these via the Three AI Pills. You can take zero, one, two or

Feed