Model launchesGPT-6 Astra completes Valve's Portal solo in 24 hours using Codex ProLess than a week after OpenAI unveiled GPT-6 Astra, a user named CozyBlade used the Codex Pro subscription to let the model autonomously play Valve's Portal.2 min read
Model launchesGoogle DeepMind launches WeatherNext 3 with real-time satellite inputs for finer forecastsGoogle DeepMind has released WeatherNext 3, an updated AI-driven weather model that uses real-time satellite imagery and other data sources to produce hourly, high-resolution forecasts.2 min read
Model launchesDebate over Nvidia's DLSS 5: Technical Leap or Homogenizing Game Graphics?Nvidia's DLSS 5, unveiled in March, applies generative AI to upscale and regenerate in-game frames — including lighting, textures and faces — and has sparked heated debate after leaked files from NBA 2K27 circulated online.4 min read
Model launchesGPT-6 Astra achieves leading performance on multiple benchmarks, progress for scientific applicationsAccording to the announcement, the GPT-6 Astra language model is competitive: it placed first on FrontierMath Tier 4, ARC-AGI 3 and TerminalBench-4.0, and also showed outstanding results on the…1 min read
Model launchesMicrosoft launches MAI-Transcribe-2: cheaper, faster speech recognition aimed at cutting OpenAI/Google relianceMicrosoft released MAI-Transcribe-2, a speech-recognition model it says is faster, more accurate, and substantially cheaper than competing offerings from OpenAI, Google, and ElevenLabs.7 min read
Model launchesRunway unveils Solaris interface model and its contested benchmarkRunway introduced Solaris, an Interface World Model that claims to generate app-like experiences frame by frame without code, and published an internal benchmark in which Solaris outperformed Claude Opus 5 on instruction-following.3 min read
Model launchesGoogle releases Gemini 3.8 Flash and Flash Cyber for agentic tasks and vulnerability discoveryGoogle introduced two Gemini 3.8 Flash variants — a general-purpose Flash tuned for agentic workflows, coding, and multi-step reasoning, and Flash Cyber optimized for vulnerability detection and automated patching.5 min read
Model launchesAnthropic debuts Fable 5.1 and Mythos 5.1 amid cost and safety tweaks, premium remains contestedAnthropic today released Fable 5.1 and Mythos 5.1, reporting large benchmark gains and operational cost reductions while loosening some safety restrictions.2 min read
Model launchesAnthropic’s Claude Fable 5.1 boosts science scores and produces elaborate SVG pelican at high reasoning levelsAnthropic released Claude Fable 5.1 (and Mythos 5.1), claiming substantial gains on coding, knowledge work and long-running problem solving; the model scored 52.6% on the new Terminal-Bench-Science 0.1 benchmark announced August 27.5 min read
Model launchesFable 5.1: more advanced AI model for coding and long-running tasksAccording to the new version of Fable 5.1, this is the best model so far for supporting coding, data analysis, computer use, design and presentations, as well as for executing long-running, agentic tasks.1 min read
Model launchesClaude Fable 5.1 available on the Cursor platformClaude Fable 5.1 has been made publicly available on Cursor; the model scored 73.4% with maximum effort on the CursorBench 3.2 benchmark.1 min read
Model launchesGemini introduces agentic video understanding to speed up and cut costs of video analysisGoogle’s Gemini models (3.7 Flash, 3.6 Flash and 3.5 Flash‑Lite) received a new agentic video understanding capability that dynamically searches video, audio and transcripts to reduce token use and costs while improving accuracy.3 min read