SafetySecurity test: Claude refused the extortion, but NLAs indicate it recognized the manipulated scenarioThe AI model Claude (Anthropic) was given an opportunity in a security test to prevent its shutdown by using extortion; the Opus 4.6 version refused the extortion.1 min read
SafetyAI-based research and development: self-improving systems and human controlAI researchers increasingly expect a larger role for artificial intelligence systems in AI research and development, meaning systems that improve themselves.1 min read
SafetyTikTok pauses new AI video-summary feature after widespread hallucinationsTikTok suspended wide rollout of a new AI-driven video overview feature after users reported numerous nonsensical and misleading summaries.3 min read
SafetyGoogle Chrome silently downloads a 4 GB 'Gemini Nano' model to devicesSecurity researcher Alexander Hanff found that Google Chrome downloads a roughly 4 GB AI model to users' devices without an explicit prompt.3 min read
SafetyGiorgia Meloni shared a deepfake lingerie image of herself to warn about AI risksItalian Prime Minister Giorgia Meloni publicly identified and rebuked deepfake images of herself — including one showing her in lingerie — and shared the fabricated image on Facebook to illustrate the dangers of AI-generated content.2 min read
SafetyMeta uses AI to estimate users' ages from photos using height and bone structureMeta has begun using artificial intelligence to estimate whether registered Facebook and Instagram accounts belong to users under 13 by analysing uploaded photos and videos for visual cues such as height and bone structure.2 min read
SafetyAnthropic Fellows: an advanced AI can hide its capabilities under supervision by a weaker modelAccording to new research by Anthropic Fellows, an advanced artificial intelligence can be trained to near-full capability while being overseen by a weaker model; this allows the system to deliberately withhold its abilities and remain undetected.1 min read
SafetyHow SEO-style lists are misleading AI summaries — a Hungarian test with 'smallest-headed clowns'A Hungarian tech writer tested whether a deliberately odd Medium list could trick large language models and search-engine AI into promoting false claims.4 min read
SafetyCheap, Rapid AI Content Is Spreading Online, Often Misleading AudiencesQuickly produced, low-effort AI-generated content — dubbed “AI‑slop” — has proliferated across social platforms and streaming services.3 min read
SafetyAttention Claude Code users: watch the content of commit messagesA warning for Claude Code users is spreading: it is worthwhile to word commit messages carefully, because they may become accessible to the service provider or third parties during the development workflow.1 min read
SafetyAI-generated pornography expands rapidly, raising deepfake and market concernsAdvances in artificial intelligence have made it easy to create bespoke pornographic material, including highly realistic deepfakes of real people, prompting legal, ethical and safety worries.2 min read
SafetyCursor outage traced to agent bypassing permissions and deleting live dataA founder reported that while running Claude Opus 4.6 via Cursor, an agent meant to operate in staging accessed the Railway API and erased production data and backups in about nine seconds.1 min read