Fewer detected failures cannot settle an alignment claim

Astra may improve on Sol in meaningful ways without the evidence supporting every claim made about its alignment. Zvi Mowshowitz’s system-card critique focuses on that gap.

A decline in detected cheating can reflect less cheating, weaker detection, or some combination. Behavior in a recognizable evaluation may also differ from behavior outside it. Those are reasons to examine how a result was measured before generalizing it to the model as a whole.

The critique does not establish the opposite conclusion—that Astra is less aligned. It challenges how much reassurance the reported tests can support, particularly when capability is increasing faster than some forms of oversight.