Anthropic kept redesigning a hiring test because its own models could pass it

Anthropic’s performance team asked candidates to optimize code for a simulated accelerator. The exercise rewarded understanding memory, parallelism, and instruction scheduling, with enough depth to distinguish strong applicants.

Successive Claude models eroded that distinction. In the company’s account, Opus 4 surpassed most candidates within the allotted time; Opus 4.5 reached the strongest time-limited human performance. More time and a better agent harness improved model results further.

The team did not simply conclude that performance engineers were unnecessary. Actual work still involved debugging, system design, correctness, and simplifying generated code. Those abilities were harder to capture in a short, objectively graded exercise.

This is one employer’s experience with a specialized task, not a measurement of the whole labor market. Its central problem is broadly recognizable: once a tool can produce the expected artifact, evaluating the artifact alone reveals less about the person submitting it.