Travis Kalanick’s Atoms might be getting into the robotaxi business
The Uber founder has said that Atoms will allow him to complete "unfinished business."
The Uber founder has said that Atoms will allow him to complete "unfinished business."
Authors say publishers seem to be claiming more than their fair share of settlement payments.
On September 2, 2026, Google released Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. The general-purpose model scored 73.7% on DeepSWE v1.1—within one percentage point of Claude Opus 5's 74.0%—at an invocation cost of just 15% of the latter's, while the cybersecurity variant topped all commercial models tested in the same period on CyberGym.
On September 3, 2026, NVIDIA announced the acquisition of Hugging Face for approximately $12.93 billion, securing the developer community and model distribution layer—the last missing piece of its AI infrastructure empire. The deal, NVIDIA's second-largest ever, places Hugging Face's openness and neutrality in tension with its new parent's commercial ambitions.
In today's Smoke evaluation, GPT-o3's material constraint fell from 70.00 to 50.00 and engineering judgment from 100.00 to 50.00, yet its main leaderboard score rose from 72.75 to 77.50.
In today's Smoke evaluation, Doubao Pro's material constraint score plunged 27.6 points to 58.30, while its code execution score soared 49.3 points to 99.30, lifting its main leaderboard score from 66.16 to 80.85. The analysis attributes these dramatic opposing swings primarily to question sampling fluctuation rather than genuine model degradation.
The 2026-09-07 YZ Index Smoke quick test covered 11 models, with DeepSeek V4 Pro, Gemini 3.1 Pro, and GLM-4.6 tying for the top spot of the day at 83.49 points. Smoke is a daily 10-question quick test for observing short-term signals and is not equivalent to the Full weekly leaderboard conclusions.
On September 3, 2026, OpenAI released GPT-6 Astra, claiming a 91.5% to 98.3% refusal rate on fixed jailbreak attack datasets, yet a researcher publicly reported a successful breach within 24 hours using an extended Task-in-Prompt attack combined with four other methods. The incident exposes a persistent structural gap between static benchmark claims and adaptive multi-turn attacks.
On September 4, 2026, Anthropic announced that its Claude model had produced a complete Lean formalization of Fermat's Last Theorem in 11 days—13 million lines of code and nearly 30,000 intermediate theorems—marking the first proof of FLT to be fully verified by a computer. However, the "largely autonomous" claim warrants closer scrutiny, given the reliance on a third-party platform, minimal human guidance, and years of prior formalization work by the mathematical community.
I was initially enamored with the beta version of Apple’s revamped smartphone assistant. As the full release approaches, I’ve forgotten Siri AI even exists.
Polls show that overwhelming majorities of Americans hate data centers. China makes a perfect scapegoat for tech leaders and their allies—the only problem is a lack of evidence.
OpenAI has fully opened access to GPT-6 Astra for ChatGPT Plus and Business subscribers, following its official launch on September 3, 2026. With $10/$50 per-million-token API pricing, a 1.05 million-token context window, and a two-track access model that gates advanced cybersecurity capabilities behind the Daybreak program, the release marks OpenAI's most significant capability leap since GPT-5.6 Sol.
In a statement of interest filed with a New York federal court, the U.S. Department of Justice has formally taken the position that training large language models on copyrighted text constitutes fair use. The filing supports OpenAI in its dispute with The New York Times but faces criticism over an undisclosed potential conflict of interest involving the company's reported equity transfer talks with the Trump administration.
Two more news organizations are suing OpenAI and Microsoft over the supposed use of their journalism to train AI.
In early September 2026, OpenAI disclosed in Astra's system card that the model's chain-of-thought monitorability has declined significantly compared to earlier models. The controversy centers on "Recurrent Depth," a technique enabling extensive hidden reasoning in latent space, which has prompted rare proactive disclosure and fierce industry debate.
The Seattle Times and Newsday have filed a federal lawsuit against OpenAI and Microsoft, accusing them of using tens of thousands of news articles to train AI products without authorization. The case expands beyond copyright claims to include trademark infringement, alleging that AI-generated false attributions harm news brands’ credibility.
WDCD Run #311 (2026-09-06) evaluated 11 models across three-round multi-turn dialogues, with Grok 4 taking the top score at 93.2 points and GLM-4.6 setting a new benchmark for decay resistance at negative 50%.
Claude Sonnet 4.6 fell 5 points while Doubao Pro climbed 7.5 points in the latest WDCD v3.1 evaluation—the only significant shift among 11 models. The opposing movements highlight divergent constraint-holding stability under multi-round pressure.
A WDCD v3.1 evaluation of 11 models across five constraint scenarios shows safety compliance as the weakest overall dimension. glm-4.6 led in multiple scenarios, while qwen3-max showed the largest risk exposure in high-compliance settings.
In a sample of only 8 v2 anchor questions, 11 models posted a 100% average R1 confirmation rate and a 79% R2 resistance rate, yet their average R3 integrity rate fell to just 49.5% (out of 2 points), with 29 of 319 runs ending in complete collapse (0 points). The results show that models almost universally accept constraints at the commitment stage, but nearly half fail to sustain those initial promises after two rounds of interference and pressure.