Model launchesOpenAI unveils ChatGPT 'Ultrafast' mode to accelerate GPT-5.6 SolOpenAI has introduced an "Ultrafast" mode for ChatGPT that accelerates the GPT-5.6 Sol model, claiming a 14-fold increase in processing speed and output rates up to 750 tokens per minute.2 min read
Model launchesUltrafast model with Cerebras accelerator: up to 750 tokens/second performanceThe Ultrafast language model runs on Cerebras hardware and can produce up to 750 tokens per second, providing low-latency intelligence for real-time speech processing, customer service, commerce, coding, design, financial research, and security response.1 min read
Model launchesDeepSeek launches V4‑Pro GA and open-source Harness while shifting API to peak/off‑peak pricingDeepSeek released the official DeepSeek‑V4‑Pro model and an MIT‑licensed open‑source agent harness called DeepSeek Harness (dsh) on Aug.5 min read
Model launchesWriter launches Palmyra X6 and harness upgrades to cut token costsWriter introduced Palmyra X6, a post-training variant of Z.ai’s open-source GLM-5.2, together with upgrades to its agentic harness intended to reduce token costs for customers.3 min read
Model launchesGPT‑5.6 cuts agent costs with model selection and new API primitivesThe GPT‑5.6 model family improves agent performance while substantially reducing inference costs by enabling smaller models to handle longer-horizon tasks and by introducing new Responses API primitives.4 min read
Model launchesMusk and Zuckerberg Close In on AI Leaders with New Models and Big InvestmentsElon Musk's SpaceX and Mark Zuckerberg's Meta have released new AI models that narrow the gap with frontier labs such as OpenAI and Anthropic, combining competitive performance with aggressive cost and scale strategies.3 min read
Model launchesGoogle DeepMind integrates sign-language-to-text into Pixel 11 input toolsGoogle DeepMind has added a sign-language-to-text mode called SL2T to Gboard and Live Transcribe on the Pixel 11, initially supporting American Sign Language (ASL) to English.2 min read
Model launchesTwo Budget Models Narrow the Frontier: Price and Speed Undercut High-End LLMsTwo recently released models, DeepSeek V4 Pro and Grok 4.6, arrived within hours of each other and challenged the performance monopoly of leading systems by attacking price and speed respectively.2 min read
Model launchesAlibaba releases Qwen3.8-2.4T-A95B open weights; optimized multinode inference on NVIDIA GB300 NVL72Alibaba published the weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), a 2.4‑trillion-parameter open-weight Mixture-of-Experts model with 95B activated parameters per token, hybrid attention and support for up to one million tokens context.3 min read
Model launchesSpaceXAI releases Grok 4.6 targeting long-running agents with competitive performance and mid-range token pricingSpaceXAI (formerly xAI) has launched Grok 4.6, an updated frontier model optimized for long-running agents, coding and knowledge work.6 min read
Model launchesNVIDIA releases Nemotron 3.5 Lightning: a small, fast MoE model for high-volume, always‑on agentsNVIDIA introduced Nemotron 3.5 Lightning, an open 30B mixture‑of‑experts (MoE) model with 3B active parameters optimized for high-frequency execution in long‑running AI agents.4 min read
Model launchesLTX releases LTX-2.5 with native ComfyUI integration, open weights and faster video/world generationLTX, spun out of Lightricks, launched LTX-2.5 — an open-weights video and world model — with day-one native support in ComfyUI.6 min read