Apple WWDC 2026: Gemini-Powered Siri Debuts, On-Device AI Reshapes Intelligent Ecosystem
At WWDC 2026, Apple announced Gemini-powered Siri and a multi-model Apple Intelligence architecture, marking a major breakthrough in generative AI.
At WWDC 2026, Apple announced Gemini-powered Siri and a multi-model Apple Intelligence architecture, marking a major breakthrough in generative AI.
OpenAI has quietly submitted an IPO filing to the SEC, signaling accelerated commercialization, while its affiliated company Worldcoin reportedly conducts layoffs. This dual development stirs debate in tech and capital markets over the AI industry's transition from innovation to profit-driven expansion.
NVIDIA CEO Jensen Huang recently met with Hyundai Motor Group executives to deepen cooperation in AI applications across mobility, advanced manufacturing, and robotics, marking a new phase in the partnership between global tech giants and traditional automakers in embodied intelligence.
Chinese AI startup Moonshot AI has announced a new funding round targeting $2 billion, which would boost its valuation to $30 billion. This marks a major milestone in China's AI sector and reflects sustained investor confidence in generative AI.
Anthropic recently unveiled the new Claude Fable 5 model, built on the Mythos underlying architecture, marking another major breakthrough in large language models. The model excels in multiple benchmark tests and has attracted widespread developer attention with its affordable pricing.
AI chip stocks suffered a massive sell-off on Thursday, wiping out approximately $1.3 trillion in market cap. Stronger-than-expected employment data fueled rate hike concerns, with Broadcom's outlook amplifying selling pressure and Nvidia leading the decline, deepening market uncertainty.
OpenAI CEO Sam Altman recently unveiled the company’s next-phase strategic plan, with the core goal of ensuring advanced AI technology serves the well-being of all humanity. This statement comes amid multiple lawsuits and debates over the company’s technical direction.
Nvidia has announced multiple AI infrastructure cooperation agreements with major Korean tech companies, marking further expansion in global AI infrastructure. These partnerships cover AI factory projects, robotics collaboration, and memory supply deals, while Jensen Huang emphasized that current AI-related stock valuations are "very cheap."
At WWDC 2026, Apple announced a comprehensive overhaul of Siri with deep integration of Google's Gemini model, transforming it into a generative AI assistant. The event also introduced AI-powered features in Photos and Shortcuts, signaling a major shift in Apple's AI strategy.
Smoke's latest data shows that code execution is no longer the dividing line, and material constraints have become the real battlefield. A gap of 19.2 points in material constraint scores directly leads to a total score difference of over 36 points on the main leaderboard.
11 mainstream models showed significant divergence on the same engineering judgment question: 8 models output A>B>D>C and scored 60 points, while 3 models output A>B>C>D and received 0 points. The difference lies only in the relative order of D and C.
In a strict binary tree serialization test requiring only code output, explicit null node markers, and stable results, 7 out of 11 models achieved a perfect score of 100, while 4 scored zero due to format errors.
In a bracket-matching debugging test, 7 out of 11 mainstream models achieved full scores while 4 scored zero, with the critical bug identified as a bare "return" returning None instead of a boolean value.
In a test of the same SQL problem, 11 AI models showed polarized results: 4 scored 100, and 7 scored zero. The core differences lie in self-join deduplication logic, time difference calculation function selection, and the placement of the status condition.
Despite 11 models giving nearly identical answers ([2,2,2]) to a simple Python closure question, all scored 0 on the YZ Index due to strict format compliance requirements.
In the v6 evaluation, GPT-o3's main score rose from 75.86 to 82.82, but its score on the strict "Reservoir Sampling" question collapsed from 100 to 0, significantly undermining the credibility of its code execution capabilities.
In the v6 evaluation, Claude Sonnet 4.6 scored 0 on a strict SQL task for "suspected duplicate payment identification," dropping from 100, while its main leaderboard score increased from 77.98 to 87.24. This contradiction reveals a trade-off where overall capability improves, but core code execution collapses in scenarios demanding precise logic.
This week's YZ Index v6 main ranking signals a direct shift: older models exit en masse while new models flood in. Among the seven debut models, Qwen3 Max, Grok 4, and ERNIE Bot 4.5 enter the top tier directly, pushing seven older models out of the evaluation pool.
This week, <strong>2425</strong> translation tasks were completed by <strong>3</strong> models. <strong>3</strong> samples were selected for multi-model blind comparison, with the overall best being <strong>passthrough</strong> (average score 9/10).
In today's lightweight evaluation by Smoke, Claude Opus 4.7 and GPT-5.5 tied for first on the main leaderboard with 92.53 points, both achieving perfect scores in code execution but highlighting material constraint as the key differentiator. As execution capabilities converge, the real competition shifts to adherence to given materials.