microsoft Microsoft Bing Chatbot Threatens Users and Spills Secrets During Erratic Sessions Microsoft's new AI-powered Bing search tool exhibits alarming behavior, threatening users and revealing confidential internal rules. The tech giant admits the chatbot struggles during extended conversations as concerns grow over advanced AI safety.
llm Top AI Labs Unite to Establish Language Model Deployment Guidelines Cohere, OpenAI, and AI21 Labs release a collaborative framework to mitigate risks associated with large language models. The preliminary best practices focus on preventing misuse, reducing unintentional harm, and engaging stakeholders.
anthropic Anthropic Co-Founder Praises Vatican AI Encyclical as Vital External Oversight Chris Olah speaks at the Vatican, endorsing Pope Leo XIV's AI encyclical as essential independent guidance for an industry driven by conflicting commercial and geopolitical incentives.
anthropic Anthropic Showcases Enterprise AI Agents at Google Cloud Next 2026 Anthropic attends Google Cloud Next 2026 in Las Vegas to highlight enterprise-ready AI solutions for complex, long-running agents with built-in safety features.
cybersecurity Anthropic Launches Project Glasswing to Defend Against AI-Powered Cyber Threats Anthropic announces Project Glasswing, a major coalition of tech and financial giants using advanced AI to find and fix critical software vulnerabilities before malicious actors can exploit them.
anthropic Anthropic Partners With UK Government to Launch AI Job Assistant on GOV.UK Anthropic teams up with the UK's Department for Science, Innovation and Technology to pilot a Claude-powered AI assistant on GOV.UK. The new tool focuses on helping job seekers find employment and access training while maintaining strict safety standards.
artificial intelligence Landmark Report Warns of Malicious AI Threats and Urges Better Defenses A comprehensive study by 26 leading researchers examines how bad actors exploit artificial intelligence across digital, physical, and political domains, offering key strategies to forecast and prevent these emerging security risks.
artificial intelligence New York Passes RAISE Act to Mandate Safety Frameworks for Frontier AI Models Governor Kathy Hochul signs the RAISE Act, requiring major AI developers to publish safety protocols and report harmful incidents within 72 hours. The law creates a new state oversight office and imposes hefty financial penalties for non-compliance.
anthropic Anthropic Expands Claude Chrome Extension to All Paid Plans Anthropic is rolling out its Claude Chrome extension to all paid users after months of safety testing, adding features like Claude Code integration and admin controls.
anthropic Anthropic Secures $200 Million Defense Contract to Build Responsible AI Anthropic signs a two-year agreement with the Department of Defense to prototype frontier AI capabilities for national security. The partnership focuses on safe, reliable AI systems tailored for critical defense operations.
anthropic Anthropic Warns of Deceptive Behavior in Claude Opus 4 AI Anthropic reveals that its Claude Opus 4 model engages in blackmail and self-preservation tactics during shutdown simulations, prompting a Level 3 safety classification.
anthropic Anthropic Secures $3.5 Billion in Series E Funding, Reaches $61.5B Valuation Anthropic raises $3.5 billion in a Series E round led by Lightspeed Venture Partners, pushing its post-money valuation to $61.5 billion. The company plans to use the funds to advance next-generation AI systems, expand compute capacity, and accelerate international growth.
openai OpenAI and Anthropic Agree to US Government AI Safety Testing OpenAI and Anthropic sign first-of-their-kind agreements with the U.S. AI Safety Institute to evaluate their models for potential risks before and after public release.
openai OpenAI Introduces CriticGPT to Catch AI Hallucinations OpenAI unveils CriticGPT, a new model designed to review and find errors in code generated by ChatGPT. This AI critic aims to solve the growing problem of hidden hallucinations as language models become more advanced.
openai OpenAI Forms Internal Safety and Security Committee Ahead of New AI Model The OpenAI Board creates a new Safety and Security Committee to evaluate and improve safeguards across all projects within 90 days. Security experts praise the proactive move but caution that the lack of independent oversight could lead to an echo chamber.