SafetyShadow AI in companies: hidden data and security risks from unsanctioned employee useShadow AI — employees’ unsupervised use of generative AI tools like ChatGPT, Gemini or Claude — is spreading rapidly in Hungary and poses data protection and security risks.4 min read
SafetyResearchers find ChatGPT can generate violent and sexual images despite safeguardsResearchers at Mindgard found that the latest version of ChatGPT can produce images with violent and sexual content when given a commonly used prompt, even without explicit instructions to do so.3 min read
SafetyAnthropic’s Claude Fable 5 safeguards complicate independent benchmark evaluationIndependent testers found that Anthropic’s Claude Fable 5 often refused prompts or routed them to the lower‑capability Claude Opus 4.8, and Anthropic’s 30‑day retention policy deterred some evaluations.5 min read
SafetyModel alignment under pressure: increased resistance to harmful interventionsThe development team tested whether the alignment persists under pressure: the model was harder to steer toward harmful behavior with adversarial prompts, while still responding to useful instructions.1 min read
SafetyBehavioral training conducted in health conversations improved the model's performance in other areas as wellResearchers found that when instruction in desirable behavior was limited to health-related conversations for an AI model, the model nevertheless improved on non-health evaluations — for example,…1 min read
SafetyModels refined with a global medical networkThe development team is working with hundreds of physicians across 60 countries, in 49 languages and 26 specialties to improve the artificial intelligence’s responses based on feedback.1 min read
SafetyClaude tests how well it can program the robodog in the second phase of the New Frontier Red Team Project FetchIn the second phase of the New Frontier Red Team Project Fetch, Claude's large language model capabilities were evaluated for programming a robodog; Opus 4.7 operated about 20 times faster on its own…1 min read
SafetyVivienne Ming warns that outsourcing thinking to AI could weaken cognitive reserve and raise dementia riskVivienne Ming of the Possibility Institute cautions that the way people use AI—specifically routinely offloading mentally demanding tasks—may erode cognitive reserve, a brain mechanism that helps resist aging and injury.3 min read
SafetySeattle Fire Department quietly used Copenhagen AI to triage 911 medical callsSince December 2023 the Seattle Fire Department routed Corti, an AI from a Copenhagen startup, onto every 911 medical call and used it to prompt dispatchers to redirect some callers to a nurse advice line in Texas.3 min read
SafetyU.S. restricts foreign access and Anthropic pauses Mythos after model finds widespread software bugsAnthropic temporarily took its Mythos AI offline and the U.S.3 min read
SafetyWhy modern AI models often seem slower and more cautiousContemporary AI systems can appear less responsive or more hesitant than earlier versions, but this reflects added layers of safety, regulation and tooling rather than reduced capability.5 min read
SafetyAnthropic to Require Identity and Age Verification for Claude Users in Policy UpdateAnthropic will update Claude’s privacy policy on July 8 to allow age and identity verification for Free, Pro, and Max users.2 min read