Jalapeño’s first results show industry-leading speed and efficiency in AI inferenceJalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.openai.com
How we built a realtime system for responsive voice AI in six monthsGPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.openai.com
How GPT-5.6 fuses frontier intelligence with frontier efficiencyGPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.openai.com
Core dump epidemiology: fixing an 18-year-old bugOpenAI engineers used large-scale core dump analysis to debug rare infrastructure crashes, uncovering both a hardware fault and a long-standing software bug.openai.com
Building self-improving tax agents with CodexSee how OpenAI, Thrive, and Crete built a self-improving tax agent with Codex, automating filings, improving accuracy, and accelerating workflows.openai.com
Building a safe, effective sandbox to enable Codex on WindowsLearn how OpenAI built a safe, effective sandbox to enable Codex on Windows with controlled file access and network limits.openai.com
Unlocking large scale AI training networks with MRC (Multipath Reliable Connection)OpenAI introduces MRC (Multipath Reliable Connection), a new supercomputer networking protocol released via OCP to improve resilience and performance in large-scale AI training clusters.openai.com
How OpenAI delivers low-latency voice AI at scaleHow OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.openai.com
An open-source spec for orchestration: SymphonyLearn how Symphony, an open-source spec for Codex orchestration, turns issue trackers into always-on agent systems—boosting engineering output and reducing context switching.openai.com
Speeding up agentic workflows with WebSockets in the Responses APIA deep dive into the Codex agent loop, showing how WebSockets and connection-scoped caching reduced API overhead and improved model latency.openai.com
From model to agent: Equipping the Responses API with a computer environmentHow OpenAI built an agent runtime using the Responses API, shell tool, and hosted containers to run secure, scalable agents with files, tools, and state.openai.com
Beyond rate limits: scaling access to Codex and SoraHow OpenAI built a real-time access system combining rate limits, usage tracking, and credits to power continuous access to Sora and Codex.openai.com
Harness engineering: leveraging Codex in an agent-first worldBy Ryan Lopopolo, Member of the Technical Staffopenai.com
Unlocking the Codex harness: how we built the App ServerLearn how to embed the Codex agent using the Codex App Server, a bidirectional JSON-RPC API powering streaming progress, tool use, approvals, and diffs.openai.com
Inside OpenAI’s in-house data agentHow OpenAI built an in-house AI data agent that uses GPT-5, Codex, and memory to reason over massive datasets and deliver reliable insights in minutes.openai.com
Unrolling the Codex agent loopA technical deep dive into the Codex agent loop, explaining how Codex CLI orchestrates models, tools, prompts, and performance using the Responses API.openai.com
Scaling PostgreSQL to power 800 million ChatGPT usersAn inside look at how OpenAI scaled PostgreSQL to millions of queries per second using replicas, caching, rate limiting, and workload isolation.openai.com
How We Used Codex to Ship Sora for Android in 28 DaysOpenAI shipped Sora for Android in 28 days using Codex. AI-assisted planning, translation, and parallel coding workflows helped a nimble team deliver rapid, reliable development.openai.com
How we built OWL, the new architecture behind our ChatGPT-based browser, AtlasA deep dive into OWL, the new architecture powering ChatGPT Atlas—decoupling Chromium, enabling fast startup, rich UI, and agentic browsing with ChatGPT.openai.com