OpenAI's WebSocket Overhaul Makes Agents 40% Faster
OpenAI rebuilt the Responses API around persistent WebSocket connections, cutting agent latency by up to 40% and letting GPT-5.3-Codex-Spark hit 1,000 tokens per second in production, with bursts to 4,000 TPS.