OpenAI says its top agent priority is automating the research that builds the next models
The Information’s Rocket Drew reports that OpenAI researcher Noam Brown described automating AI research as the company’s “number one priority” for agents. OpenAI’s GPT-6 Astra improved across varied professional tasks, including game design and sheet-music transcription, but Brown’s emphasis is on a more consequential use: helping conduct the research and engineering that produce later models.
That creates a compounding mechanism. An agent that can generate hypotheses, write experimental code, run evaluations and interpret results could reduce the time between model generations. Even partial automation matters if it expands the number of experiments researchers can test. But “recursive” improvement should not be mistaken for a system independently redesigning itself without humans: today’s bottlenecks include trustworthy evaluation, compute allocation, data quality and deciding which apparent gains are real.
The strategic risk is also the mechanism’s attraction. If labs rely on models to perform more of the work used to assess successor models, independent checking becomes harder at the same time the development cycle accelerates. Broad benchmark gains do not establish reliable scientific judgment, especially when an agent can produce plausible but subtly wrong analyses.
The most valuable agent may not be the one that replaces a user task; it may be the one that increases a frontier lab’s rate of technical progress, concentrating advantage in organizations that already possess models, compute and experimental infrastructure.