Bringing my LED display to life with GPT-Live-1 and Codex

For my roommate’s birthday, I’d bought him a little LED display that shows flights passing near our window. It was a pretty cool gift, but after watching planes pop up on the screen, I kept wondering whether a display like that could do more than track airplanes. Could it show the weather or my calendar? And when I was heading out, could I just ask, “When’s the next train?” and see the answer on the display? Your browser does not support the video tag. This quickly turned into a side project involving a Raspberry Pi, a 128 × 64 HUB75 LED panel, Codex, and an assistant I eventually started calling Jack. The conversation runs on GPT-Live-1, the full-duplex voice model we recently launched in the API. Another model helps with research, and a renderer takes care of the pixels, although we definitely didn’t start with anything that well thought out. This post documents how that process went. How Codex helped me get a HUB75 RGB LED panel working At first, Codex and I had to figure out how to get anything useful onto the display. I was working with a 128 × 64 HUB75 RGB panel, but I didn’t know much about the controller, the wiring, or what it would take to send it something other than the content it was already showing. I started sending Codex pictures and asking it to explain what I was looking at. It helped identify the important connections, make sense of the ESP32-based controller, and work out how a HUB75 panel actually receives display data. It even marked up the pictures to show me which parts mattered, which was useful because, on my own, I mostly saw a circuit board and a lot of opportunities to guess wrong. Once we understood the basic hardware, Codex helped me get the controller running WLED-MM and confirm that the panel was configured for the right 128 × 64…

Rethinking skills and prompts for GPT-6 Astra

Coding agents have come a long way, and best practices are changing fast. With more capable models, what used to require a lot of handholding and scaffolding no longer does. If you’ve been using agents like Codex for your projects over the last year, you’ve likely accumulated a lot of instructions as you worked to steer the models toward good outcomes. With each release, it’s been worth revisiting those assumptions, but with GPT-6 Astra, it’s more important than ever. These instructions can take many forms: skills, AGENTS.md, and your task prompts are all shaping how the model gets work done. Better skills These instructions can be in the form of skills, which are essentially prompts stored as Markdown files that can also be packaged with resources and bundled scripts. Generally, they are most useful for guidance around a specific workflow, or when using certain apps. People now default to packaging a lot of skills into their projects, and each skill comes with a name and description that are loaded into the model’s context so it knows when to use them. But many descriptions are far too long, and when you add too many skills, Codex starts shortening their descriptions to fit. The model ends up seeing less of each description, making it harder to know which skill to pick. What’s worse is that descriptions can often contradict each other or over-emphasize when skills should be used, leading the model to load instructions that don’t actually help the task. A common workflow to create skills is to use the $skill-creator skill. We recently updated its guidance to help mitigate many of the failure modes we’ve seen in practice. First, skill descriptions should be as short as possible while making it clear when the model should use them: Be clear about when it applies Bad…

Building games with Astra

Astra has become very good at translating what I have in mind into gameplay and art direction. I’ve been using it in Codex to build Void Explorer, a space exploration game where you can travel from distant stars down to alien landscapes. A gameplay supercut from Void Explorer. The game has 2,048 star systems and more than 10,000 procedurally generated planets, including worlds the size of Earth. Every visible star belongs to the universe and can be targeted. You can pick a point of light in the distance and travel to it. You can approach a planet from space, pass through its atmosphere, and keep descending until you’re flying over a coastline. You can land, get out of the ship, walk around, and take off again. Getting that journey to work meant solving the scale, terrain, controls, and rendering together. I’ve included a few prompts from the build, edited for length and clarity. They show how I described what I wanted and how the technical work developed from there. Start with the experience My initial brief described what the player should be able to do: Everything I can see should be reachable. Keep the distances real, then make travel work through scale and speed. I want to fly from space into a planet’s atmosphere and down to the ground. Planets can be as large as Earth, so we’ll need procedural terrain and a chunked renderer. That gave Astra some concrete constraints. A star couldn’t just be a dot painted into the background. A planet couldn’t become a separate level when I approached it. Travel had to cover real distances, with fictional pulse travel and hyperdrive making those distances practical. I also used image generation to work out the look before building much of the game. The first concepts were too realistic. The next direction was too simple. This was…

Architectural visualization with Astra

I started with a simple brief for a house: minimalist but detailed furniture, a garden, and a cinematic atmosphere. I asked Astra in Codex to turn that brief into an editable 3D scene in Blender. We developed it into a furnished family home, explored a version in Unreal Engine 5, and kept refining the Blender scene until I could direct a camera tour through it. The first scene centered on a furnished living pavilion. We later developed a floor plan for a larger family home before building its rooms. Along the way, I worked with Astra on the architecture, furniture, materials, and lighting, using each version to decide what to develop next. The latest pass added the everyday objects and warmer lighting that made the house feel lived in. A camera tour of the house, rendered in Blender using Cycles. I’ve included a few prompts from the project, edited for length and clarity. They show how the brief evolved as I saw what Astra could build. Start with the house and the atmosphere My first prompt described the result I wanted, without supplying a floor plan or a furniture catalog: Design a beautiful house with minimalist but highly detailed furniture and a nice garden. Make the scene cinematic, with considered lighting, materials, and time of day. Astra built an editable scene through the Blender Python API (bpy). It created the architecture, joinery, furniture, planting, materials, lights, and cameras. The first result already had the low roof, open living space, warm timber, pale stone, and connection to the garden that I was looking for. We named the house Solace. The original Courtyard House, delivered from the first design prompt. That house came from one design prompt, but Astra still iterated on its own work before handing it back. It inspected preview renders,…

Meet Rosalind Workbench: Empowering every scientist to be their own research team

Fragmented data, disconnected tools, and unclear experimental records can get in the way of scientific work: the data may be in one place, the analysis in another, and the record of how a result was produced somewhere else. Following the evidence means keeping those pieces connected as you decide what to investigate next. Your browser does not support the video tag. That’s why we built Rosalind Workbench – to bring that work into one environment. It provides life science users with a central environment to leverage their favorite scientific tools, explore new specialized biology models, and define their most used data analysis workflows. This virtual workbench offers a guided experience that helps scientists make the most of frontier models and state-of-the-art scientific tooling to accelerate their research. Rosalind Workbench is available in research preview through the ChatGPT app, where researchers can try guided scientific tasks and adapt them to their own data, methods, and research goals. Rosalind Workbench builds on GPT-Rosalind, our dedicated life sciences model, which combines frontier reasoning with specialized tool orchestration across medicinal chemistry, genomics, wet-lab assistance, and other scientific applications. Over time, we envision teams of agents working together across these domains, giving researchers access to broader expertise and the ability to pursue more ambitious scientific questions. Exploring scientific workflows Different disciplines use different data and methods, but the underlying need is often the same: connect a biological question with the right evidence and make the path to the next decision clear. Research rarely fits neatly inside one tool – consider a question about a disease mechanism: it might begin with a genomic signal,…

Automating repetitive work at OpenAI with Codex

I’ve spent most of my career as a software engineer either turning a crank—deploying and operating software—or building software to turn that crank for me. My first job at OpenAI was on the cloud infrastructure team, bringing up new Kubernetes clusters for application teams. I’d spend a week getting a batch of clusters ready, working through issues with private links, quota, and Terraform. As soon as those clusters were ready, I’d start on another batch. My next job was on the API team, running evaluations against our latest models. I’d work through problems with graders, quota, configuration, and PyTorch. When the evaluations ran successfully and the model shipped, I’d start again with the next model. Now I use Codex to help with that repetitive work. I still build software—or, more accurately, Codex helps me build it—but instead of creating a separate automation for every task, I’m building Runme to: Collect and curate the context around a workflow. Keep the right review and approval boundaries in place. Improve future Codex runs with what earlier runs learned. Running evaluations with Codex At OpenAI, we run evaluations when shipping new models and features to check that they work as expected. To run an evaluation, I create a Runme notebook and write a short outline of the goal: # Goal: Run the evaluation against the current model - Review a previous run to understand the workflow. - Write a detailed plan in this notebook. - Wait for me to review and approve the plan before beginning. - Document the commands you run, their output, and how you interpret the results. Then I ask Codex to use that notebook cell as its goal: Read the goal cell in the Runme notebook open in the browser. Treat it as the goal, write your plan in the notebook, and wait for my approval before…

Meet the winners of OpenAI Build Week

OpenAI Build Week started with a simple challenge: build something real with Codex and GPT‑5.6. Thanks to all of you for taking on that challenge and transforming the week into a vibrant celebration of the builder mindset. Nearly 47,000 builders from 186 countries took part, making this our biggest hackathon yet. Over eight days, they submitted more than 8,000 projects and joined seven digital and 60 in-person community events to learn from one another, test new tools, and ship. Build Week also offered a glimpse of where building is headed. A veterinarian with no coding background built a triage tool now being piloted in her short-staffed clinic. Between hospital shifts, a cardiologist in Cairo built a research prototype for cardiac-arrest response. Other winners drew on personal experience to build tools for speech accessibility and Vietnamese pronunciation, while developers took on MCP security, spatial-audio design, and interactive reconstructions of ancient machines. Together, their projects show how the builder community is expanding: developers can take on more ambitious systems, while people with expertise in other fields can turn what they know into working software. Today, we’re announcing eight winners across four categories: Apps for Your Life, Work & Productivity, Developer Tools, and Education. Judges evaluated each project on technical implementation, design and user experience, potential impact, and the quality of the idea. The eight winners will share $100,000 in cash prizes. First-place teams will also receive passes to OpenAI DevDay, time with the Codex team, and one year of ChatGPT Pro. Choosing only two winners in each category was difficult. Explore the full gallery to see what builders made during OpenAI Build Week. Apps for Your Life First-place:…

Scaling cyber defenders with Daybreak

If you work through a security backlog, finding another possible issue is only the start. You still need to work out whether it affects your software, gather evidence, and land a safe fix. That gets harder as the code, alerts, and vulnerability reports keep coming in. We’ve recently added more ways to work through that process with ChatGPT, Codex Security, and the open-source Codex Security CLI. You can review a pull request before it merges, investigate a repository or an existing vulnerability backlog, and add recurring checks to CI. These capabilities are part of OpenAI Daybreak, which brings together models, security tools, responsible access, and the security ecosystem for approved defenders. In previously reported results, Codex Security cloud had analyzed more than 30 million commits across more than 30,000 codebases. Here, I want to walk through where the available workflows fit and how I’d choose a starting point. The aim is the same throughout: turn findings into evidence and reviewed fixes, while keeping access scoped and people responsible for consequential decisions. These workflows are suggested starting points, not a one-size-fits-all deployment pattern. Developers should tailor them to their organization, use case, risk profile, and data-handling practices, and determine the appropriate configuration, safeguards, and deployment for their environment. Start with an investigation in ChatGPT If you already have a log excerpt, an advisory, or an incident timeline, ChatGPT is a useful place to start reasoning through it. A few things to try: Investigate a suspicious log excerpt and identify what evidence is still missing. Summarize a vulnerability advisory and map its likely impact on your systems. Reconstruct an incident timeline or draft a detection rule.…

Codex as a platform: build on the open agent harness

Most people know Codex through the App, Command-Line Interface, or IDE Extension. Those experiences are important, but they are only a few of the ways the same underlying system can be used. The open-source Codex harness is what powers all these experiences. It helps models gather context, reason through tasks, use tools, operate within configured boundaries, request approval, and carry work forward. That changes what developers can build. Instead of asking every team to move its work into a general-purpose coding assistant, you can bring the agent into software designed around the actual job: an engineering workflow, an operations dashboard, a security investigation, a customer-support console, or an internal application built for one specialized team. The reusable part is the agent loop A capable agent is more than a prompt and a model response. It needs a way to understand a task, maintain context over time, inspect relevant information, call tools, expose progress, handle failures, request human approval when necessary, and return a useful result. That surrounding execution system is the harness. Harness design can materially change results: on ARC-AGI-3, retained reasoning and context compaction raised GPT-5.6 Sol’s score from 13.3% to 38.3% while reducing output tokens sixfold. We built the Codex harness to manage conversation state, stream execution, use tools, enforce configured sandbox and approval policies, and carry work across turns. With Codex app-server, we expose those capabilities through a documented client protocol: applications can create threads, start turns, receive events, and handle approval requests. If you are building software that needs an agent, you can start with Codex instead of inventing a new runtime, then decide what the surrounding…

Custom Code Review rules for Codex

When doing code reviews with Codex, some comments keep coming back. It could be about preserving an older API contract, keeping customer data out of logs, or avoiding a rename that would break another service. These checks are important, but they are easy to miss when the context lives with a handful of reviewers. Codex Code Review can now use custom repository rules in AGENTS.md to catch those issues and point authors to the guidance behind a finding. If you already use AGENTS.md to guide coding tasks, the same file can help guide reviews, too. This is especially useful when contributors or coding agents are working in an unfamiliar part of a repository and may not know its history yet. In this post, we’ll show where repository rules fit and how to write them well, including what we learned while testing them. Shipping more code Coding agents can take on larger changes and work over longer horizons, helping teams move more of their ideas into code. At OpenAI, weekly PR volume has more than doubled since Q4, and we’re seeing similar trends for many of our customers. More code is good: it helps teams ship new features and solve more problems. It also means more pull requests waiting for someone who knows what to look for, and code review can quickly become the bottleneck. Review gets harder when several changes arrive at once. A diff can look completely reasonable and still break an older client or cross a boundary the author did not know about. Someone has to remember that context and share it while the author can still act on it. The review bottleneck When more pull requests land, reviewers have less time to work out what each change is trying to do and gather the relevant context before leaving feedback. Once an author moves on to something else, even a small…

Making private MCP servers reachable without making them public

We built Secure MCP Tunnel because the MCP servers teams care about most are often the ones they least want to expose to the Internet. We wanted to share how we approached that constraint: keeping private servers private while still giving ChatGPT, Codex, and other OpenAI products a normal MCP request path. The Model Context Protocol has made it easier for AI systems to connect to external tools and data. But many of the most valuable MCP servers run inside enterprise networks, private service meshes, developer laptops, and other environments designed to reject inbound public traffic. Connecting these servers to hosted AI products has often required teams to create public endpoints, deploy additional proxy infrastructure, or introduce new network operators into sensitive paths. Secure MCP Tunnel provides a simpler approach: Customers run a small client inside their private environment that establishes an outbound HTTPS connection to OpenAI. The client: Receives MCP requests Forwards them to an approved local server Returns responses and notifications through the same connection. OpenAI products can use the standard MCP request and response model, while the underlying server remains behind the customer’s existing network controls. Making that work reliably and securely meant solving several engineering problems at once: preserving the server’s private network boundary, supporting MCP’s streaming and authentication flows, and giving teams a client they can inspect and operate. This post walks through those decisions. We designed the tunnel around a small set of principles: outbound-only connectivity, explicit destination configuration, compatibility with MCP streaming and notifications, and a customer-run client that teams can inspect and operate themselves. Together,…

Mastering remote engineering work from your phone

Remote in the ChatGPT mobile app is easy to underestimate. At first glance, it looks like a way to check on a coding chat from your phone. That is useful, but it misses the bigger idea. The real power of Remote is that it lets you start, direct, review, and organize work running on your development machines without pretending that an iPhone should be a tiny terminal. Over the last two months, we have added a surprising amount of depth to that loop: remote host connections, worktrees, goals, side chats, inline code review, queued and steering prompts, attachments, skills and plugins, archived chats, security controls, and a long tail of small details that make the app useful for serious work. This is the field guide I wish every new power user had. The right mental model: Your phone is the control plane The code still runs where it belongs: on your Mac, Windows machine, devbox, or other connected host. In the ChatGPT mobile app, Remote gives you a native interface for controlling that work. That distinction matters. The goal is not to reproduce every terminal affordance on a small screen. The goal is to make the decisions that unblock an agent easy from anywhere: What repository and workspace should it use? Should this run in the current branch or a fresh worktree? Should my next message wait, or should it redirect the active turn? Is this command safe to approve? What changed, and do I agree with it? Should this become a durable goal, a separate chat, or a quick side question? Once you use the app this way, it stops feeling like remote desktop software and starts feeling like an engineering control plane. 1. Start the chat with the right boundaries Good agent work starts with a correctly scoped environment. Remote lets you choose the connected host and workspace before…

How Perplexity Brought Voice Search to Millions Using the Realtime API

At Perplexity, we care a lot about building products that feel amazing to use. For Perplexity Comet, our agentic browser, and Perplexity Computer, our powerful, general-purpose digital worker, a big part of that was making these fully usable through voice. There is something uniquely satisfying about being able to just say what you want, hand off the task, and watch it go. We are bullish on voice as an interface because it makes the actual interaction feel a little closer to magic. We used Realtime-1.5 in production to bring that magic to the millions of voice sessions Perplexity manages every month. Watching the growth of voice through Computer’s interface has been incredibly exciting and educational. We’ll share a few of the surprising things we’ve learned so far. We encourage you to try Realtime-1.5 and share what you learn with us too. Your browser does not support the video tag. 1. Figure Out Your Context Management Strategy Long-form content, especially dense multi-hour podcasts, was one of our clearest tests of context management. We wanted to make podcast transcripts usable through voice so a user could jump in, ask what was happening at a specific moment two and a half hours in, and get a coherent answer. You can’t fit the whole transcript into context. Our first pass was to send it in large chunks. We quickly found that large updates fail in an all-or-nothing way. If you try to send a 10,000-token update into a window that only has room for 5,000 more tokens, the model will lose all of the preceding history. That made large chunks much riskier because one oversized update could wipe out a whole block of context instead of letting the system forget more gracefully. So we changed the approach. Instead of large updates, we started breaking everything into much…

Designing delightful frontends with GPT-5.4

GPT-5.4 is a better web developer than its predecessors—generating more visually appealing and ambitious frontends. Notably, we trained GPT-5.4 with a focus on improved UI capabilities and use of images. With the right guidance, the model can produce production-ready frontends incorporating subtle touches, well-crafted interactions, and beautiful imagery. Web design can produce a large surface area of outcomes. Great design balances restraint with invention—drawing from patterns that have stood the test of time while introducing something new. GPT-5.4 has learned this wide spectrum of design approaches and understands many different ways a website can be built. When prompts are underspecified, models often fall back to high-frequency patterns from the training data. Some of these are proven conventions, but many are simply overrepresented habits we want to avoid. The result is usually plausible and functional, but it can drift toward generic structure, weak visual hierarchy, and design choices that fall short of what we visualize in our heads. This guide explains practical techniques for steering GPT-5.4 toward crafting the designs you envision. Your browser does not support the video tag. Video quality Auto 1080p 720p 480p Model Improvements While GPT-5.4 improves across a range of axes, for front-end work we focused on three practical gains: stronger image understanding throughout the design process more functionally complete apps and websites better use of tools to inspect, test, and verify its own work Image understanding and tool use GPT-5.4 was trained to use image search and image generation tools natively, allowing it to incorporate visual reasoning directly into its design process. For best results, instruct the model to first generate a mood board or several…

From prompts to products: One year of Responses

One year ago, we introduced the Responses API — a foundation for developers and enterprises to build useful and reliable agents. Equipping models with a set of hosted tools allowed AI to evolve from chat assistants to systems that can take action on your behalf. Today, the Responses API supports a number of tools to power agentic workflows and a new set of features and primitives specifically designed for building with more capable models. Thousands of developers are building with the Responses API today to accelerate productivity across industries like customer support, legal, life sciences, travel, and more. Having shared many success stories from those industries, today we’re celebrating five lesser-told stories of the developers who have built on the Responses API for the past year. Detecting and fixing failures in AI agents By Alexis Gauba and Ben Hylak from Raindrop AI Tools: Custom built tools Models: GPT-5.2 (testing GPT-5.4) Raindrop is the monitoring platform behind the world’s most ambitious AI companies to catch when their agents go off the rails in production. As agents have gotten more complex, these failures have become more critical. Without the Responses API, building this kind of monitoring system would have been much harder and a lot less reliable. The system runs background analysis using the Responses API (via the Vercel AI SDK) to share tools across different model providers and keep their system portable across environments. These workflows surface unusual behavior. When something goes wrong, the system alerts developers and assists with diagnosing the underlying issue. Your browser does not support the video tag. The platform focuses on three core systems: Agent behavior monitoring Failure detection and alerting Developer investigation and…