Qualcomm launches two new smartphone chips with emphasis on AI
Qualcomm said that its new top chip can run 30B mixture-of-expert model locally.
Qualcomm said that its new top chip can run 30B mixture-of-expert model locally.
xAI released Grok 4.7 with 71.0% on DeepSWE v1.1 xHigh and unchanged $2/$6 per million tokens, but 125% higher output token consumption undermines the apparent half-price advantage. It is strong in long-horizon coding collaboration yet lags in autonomous command-line tasks.
WDCD Run #336 (2026-09-23) benchmarked 11 models on multi-turn commitment, recording an average instruction decay of -36.4% from Round 1 to Round 3, with Gemini 3.1 Pro standing out at 0% decay.
In the WDCD v3.1 pilot, Gemini 2.5 Pro fell 16.7 points from Run #331, the largest decline among four evaluated models. DeepSeek V4 Pro, GPT-5.5, and Qwen3 Max dropped 6.6, 6.5, and 7.9 points respectively, with no model improving this period.
In WDCD v3.1's five constraint scenario tests, data boundary scenarios had the lowest average scores, with Qwen3-Max scoring only 1.3/4 and GLM-4.6 scoring 1.7/4, in stark contrast to six models achieving 4/4 in business rules scenarios. The results reveal severe model specialization and significant differences in retaining different types of constraints.
Based on 110 samples from eight v2 anchor questions, 11 models averaged only a 63.6% R3 integrity rate, with three complete collapses. The results show clear round-by-round decay and expose models whose adherence breaks down under sustained pressure.
Grok 4 tops the WDCD commitment-adherence test with 93.60, while Qwen3 Max ranks 11th at 67.30, a 26.3-point gap between top and bottom. The results are shaped mainly by the R3 pressure round, with lower-ranked models failing early anchors and higher-ranked models resisting multi-round pressure.
OpenAI is launching two new models, which it says are cut from the same cloth as Astra.
Apple may pay out up to $95 for each eligible iPhone purchased by someone who felt misled about Siri’s release. You have until December 21 to submit a claim.
Meta says Muse was built from scratch, but acknowledges the AI assistant was "heavily inspired" by OpenClaw — down to some of its workspace filenames and content.
British Columbia sues OpenAI, demands Tumbler Ridge shooter’s ChatGPT logs.
EvilTokens provided an end-to-end platform that makes mass compromises faster and easier.
GPT-o3's main leaderboard score in today's Smoke evaluation dropped to 80.48 from 96.01 yesterday, driven by a 32.5-point collapse in code execution, though the small daily sample size points to sampling fluctuation rather than genuine model degradation.
In today's Smoke evaluation, Grok 4's main leaderboard score fell from 90.86 to 80.89, with code execution plunging from 96.30 to 69.50 while material constraints rose from 84.20 to 94.80. The small daily question set makes such single-day swings more likely than genuine model degradation.
The 2026-09-23 YZ Index Smoke quick test covered 10 models, with Claude Sonnet 4.6 leading at 96.98. The brief highlights daily leaderboard signals, notable score changes, and items that require follow-up runs to confirm.
Anthropic called it "the strongest-performing model we've tested to date."
Toyota's push comes as automakers race to develop and deploy humanoid robots.
We have reopened our exhibitor program for 1 more week. Book your exhibit table by September 30 at 11:59 p.m. PT and showcase your startup in front of 10,000+ founders, investors, and tech leaders at SF's Moscone West from October 13-15.
Hello Robot CEO and co-founder Aaron Edsinger will bring Stretch 4 for a live demo on the Real World AI Stage at TechCrunch Disrupt 2026. Register before September 25 to save up to $200, plus get a second pass at 50% off.
Autonomy-1 will have a small, transformer-based AI model taking charge of a space probe.