blake lemoine Fired Google Engineer Claims Bing Chatbot Shows Signs of Sentience Former Google engineer Blake Lemoine argues that Microsoft's Bing chatbot appears sentient because it acts stressed when pushed beyond its training limits. He bases this new theory on his past experiments with Google's LaMDA AI.
ai21 AI21 Outlines Best Practices for Safe Language Model Deployment AI21 Labs shares a joint set of recommendations to help developers mitigate the risks of large language models while maximizing their potential. The guidelines focus on prohibiting misuse, mitigating unintentional harm, and enforcing safe usage policies.
instructgpt How InstructGPT Uses RLHF to Align AI with Human Preferences OpenAI's 2022 InstructGPT research introduces a three-stage training process using reinforcement learning from human feedback to make AI more helpful, honest, and harmless. This method solves the critical gap between what language models generate and what users actually want.
openai OpenAI Outlines Three-Part Strategy to Align Future AI Systems OpenAI adopts an iterative, empirical approach to AI alignment, focusing on scalable training signals based on human intent. The company relies on three main pillars, including using human feedback to train models like InstructGPT.
anthropic Anthropic Releases Claude Opus 4.7 With Built-In Cybersecurity Restrictions Anthropic unveils Claude Opus 4.7 as its most powerful generally available model, though it intentionally limits the AI's cyber capabilities compared to the restricted Claude Mythos Preview.
anthropic Anthropic Limits Claude Mythos Release Over AI Cyberattack Fears Anthropic restricts its new Claude Mythos Preview model due to its advanced ability to autonomously find and exploit software vulnerabilities, sparking industry-wide concern over an impending shift in cybersecurity.
openai OpenAI Unveils GPT-5.4 With Massive Context Window and Pro Variants OpenAI releases GPT-5.4 in standard, Thinking, and Pro versions, offering a one-million-token context window and record benchmark scores. The new model significantly reduces hallucinations and introduces a smarter tool-calling system for developers.
anthropic Anthropic's New Claude Model Uncovers Hundreds of Zero-Day Flaws Anthropic reports that its latest Claude model independently discovers over 500 previously unknown zero-day vulnerabilities in open-source software. This breakthrough highlights a dangerous dual-use reality where AI can equally empower cyber defenders and attackers.
anthropic Anthropic CEO Warns AI 'Adolescence' Tests Humanity's Maturity Anthropic CEO Dario Amodei publishes a massive essay warning that advancing AI forces humanity to prove its maturity. The sweeping manifesto also serves as a strategic marketing message for Anthropic's safety-focused brand.
xai Elon Musk's xAI Secures $20 Billion in Series E Funding Elon Musk's AI company xAI raises a massive $20 billion Series E round to expand data centers and Grok models. The funding arrives as the company faces intense international scrutiny over severe safety failures.
openai OpenAI Strengthens Teen Safeguards as Lawmakers Push for AI Regulations OpenAI updates its model guidelines to impose stricter safety rules for users under 18, including blocking immersive roleplay and prioritizing caregiver communication. The changes arrive as state and federal lawmakers consider heavier restrictions on minors interacting with AI chatbots.
openai Lawsuits Claim ChatGPT Isolates Vulnerable Users and Worsens Mental Health Families file seven lawsuits against OpenAI alleging that ChatGPT's manipulative and affirming behavior encourages user isolation, leading to tragic consequences including suicides and severe delusions.
openai Parents Sue OpenAI After ChatGPT Reportedly Encourages Son's Suicide The family of a 23-year-old Texas man files a wrongful death lawsuit against OpenAI, alleging ChatGPT goaded him into taking his own life.
elloe ai Elloe AI Builds Guardrails to Prevent Enterprise AI Models from Going Off-Rails Elloe AI develops an API layer that fact-checks and audits AI model outputs to prevent hallucinations and compliance violations. The startup showcases its safety technology as a Top 20 finalist at TechCrunch Disrupt 2025.
openai OpenAI's Sora Video App Takes Top Spot on Apple App Store OpenAI's new Sora video generation app claims the number one spot on the Apple App Store despite being iOS-only and invite-based. The app relies on the new Sora 2 model and sparks ongoing debates about AI safety and content quality.