Google DeepMind launches institute to widen the AGI debate
Google DeepMind just launched an institute to hash out the big AGI questions in public
Google DeepMind just launched an institute to hash out the big AGI questions in public
The round values the data center giant at $30.9 billion.
If AI lab PrismML isn't on your radar yet, it should be.
Newly unsealed court filings show Microsoft privately called OpenAI's data practices "theft" while both companies scraped paywalled Times content, built datasets from it, and warned internally it would gut publishers.
Microsoft, OpenAI emails reveal fear of AI “doom loop” killing news orgs.
Not everyone agrees with Amodei's call for globally coordinated action for AI safety.
Multiple family members can share data to help the agent make plans and complete tasks.
OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.
As companies hand off longer and more complex tasks to AI agents, they are running into an oversight problem: Agents can act faster, longer, and at greater volume than humans can realistically review.
OpenAI has released Astra for Law, a legal-industry configuration of GPT-6 Astra built on a 230-million-URL U.S. legal search index from CourtListener, developed alongside three elite law firms. Benchmarks show 54% overall accuracy versus 38.7% with general web search — enough to assist associates, but not yet enough to replace senior legal judgment.
Google DeepMind has announced the creation of the DeepMind Institute, an interdisciplinary platform on the societal impact of AGI, led by co-founder and Chief AGI Scientist Shane Legg. The move comes as the Future of Life Institute's summer 2026 AI Safety Index gives Google DeepMind the lowest grade of the three leading AI companies.
King Charles hosted a private summit Thursday with some of the most prominent names in AI and the U.K. government.
SynthID can cause models to follow harmful instructions they would otherwise refuse.
The Dreamforce conference became an unlikely battleground for the CEOs of OpenAI, Anthropic, and Nvidia to debate whether AI development should slow down.
By framing their efforts as a “slowdown” rather than an industry-wide push for better security standards, AI labs may have set themselves up for years of regulatory headaches.
The shift comes after a UNICEF test found leading AI models struggled to accurately retrieve global development statistics.
GLM-4.6 scored "-" across all five Smoke evaluation dimensions (execution, grounding, judgment, integrity, communication) due to an API failure, and will not participate in today's rankings. An automatic rerun has been triggered.
In today's Smoke evaluation, Qwen3 Max's material constraints score fell 16.9 points, while code execution rose 64.3 points. Its main leaderboard total rose from 54.66 to 82.42, most likely due to question sampling variance rather than a systematic capability shift.
On 2026-09-18, the YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 taking the top spot at 100 points. The brief covers Code Execution and Material Constraints scores, notable day-over-day changes, and signals to watch.
Model maker commits to new framework for reporting misaligned models.