Someone finally tested whether AI labs were optimizing for the pelican

The experiment combined eight animals with six vehicles, asked seven models for three SVGs per combination, and produced 1,008 images. That gives the famous pelican some company: herons, whales, cats, and other subjects riding things that are not always bicycles.

Castillo inspected the drawings and used model-assisted judging to evaluate them. He then adjusted for the difficulty of different animals and vehicles, looking for an unusual advantage on the exact pelican-and-bicycle combination.

He found no statistically significant advantage on that combination. GLM-5.2 showed the largest apparent boost, but it was small and uncertain. Models that drew good cycling pelicans generally drew other combinations well too.

This is a more useful test than staring suspiciously at one especially polished bird. It remains a limited experiment, with a small number of repetitions and imperfect model judges. Small advantages could escape detection, and a negative result cannot reveal what was in a model’s training data.