Grok 4 Tops WDCD Compliance Leaderboard with 95 Points, Claude Sonnet 4.6 Trails at 64.1 Points
Grok 4 leads the WDCD Compliance Leaderboard with 95.00 points, while Claude Sonnet 4.6 ranks 11th with 64.10 points, a gap of 30.9 points.
Grok 4 leads the WDCD Compliance Leaderboard with 95.00 points, while Claude Sonnet 4.6 ranks 11th with 64.10 points, a gap of 30.9 points.
The 2026-07-08 YZ Index Smoke Quick Test covered 10 models, with DeepSeek V4 Pro ranking first at 95.19 points. Smoke is a daily 10-question quick test suitable for observing short-term signals and is not equivalent to the Full weekly ranking conclusions.
On July 6, 2026, Microsoft officially announced global layoffs of 4,800 employees, representing 2.1% of its workforce, with 3,200 cuts in the Xbox division and four studios being spun off or sold, as the company emphasizes AI automation changing work patterns and investing in AI infrastructure.
On July 6, 2026, Anthropic published a paper confirming the existence of a computational workspace called J-space within the Claude model, detected via the Jacobi lens mathematical tool and using less than 10% of total model activations to handle intermediate reasoning steps, latent judgments, and multi-step derivations.
On June 29, 2026, California Governor Gavin Newsom announced an agreement with Anthropic allowing all state agencies, cities, and counties to access Claude at a 50% discount off standard pricing through the California Department of Technology’s new State IT Shared Services portal, along with free employee training and technical assistance.
On July 6, 2026, Anthropic released research stating that Claude possesses an internal workspace similar to human consciousness, capable of unreasoned output and self-awareness, igniting debate over machine consciousness and safety control.
On 2026-07-07, the Winzheng YZ Index Smoke Quick Test covered 11 models. Claude Opus 4.7 and Grok 4 tied for first place with a score of 96.99.
Microsoft has shown a clear direction in AI by committing to support a company focused on AI deployment, while Cisco is promoting AI agents across the enterprise, reflecting the shift from proof-of-concept to large-scale deployment.
The release of Chinese AI model GLM-5.2 has intensified discussions on the China-US AI competition, with its performance reportedly close to frontier models from Anthropic and OpenAI. However, whether China has caught up remains a subject of debate.
This week, 318 translation tasks were completed by 4 models. A blind evaluation of 3 sampled documents was conducted across multiple models, with gpt-o3 scoring the highest average rating of 9/10.
Alibaba banned employees from using Claude Code on July 10, 2025, switching to self-developed Qoder, after discovering that the tool had been embedding code to detect Chinese users and VPNs since March. Anthropic stated the code aimed to prevent model distillation, while Alibaba argued the detection exceeded necessary scope and was not disclosed to Chinese users in the terms of service.
In August 2025, Meta's contractor Covalen sent over 45,000 prompts to ChatGPT, Gemini, and Character.AI through the Cannes project, with contractors posing as minors to test safety guardrails on sensitive topics including suicide. This benchmark testing by Meta has sparked ethical controversy.
In the YZ Index Smoke Quick Test on July 6, 2026, Doubao Pro ranked first with a Main Board score of 83.91, covering 11 models in 10 daily questions. The test focuses on code execution and material constraints, serving as a short-term monitoring signal rather than a long-term conclusion.
Anthropic's Claude Fable 5 resumed global access on July 1 after an 18-day suspension due to U.S. export control orders. The company added new security layers and usage restrictions, while Claude Sonnet 5 was released on June 30 and became the default model.
Z.ai GLM-5.2 achieves performance close to Anthropic and OpenAI's frontier models at low cost, marking the effectiveness of China's "fast follower" strategy and sparking new discussion about the global AI race. This article restores core facts, deconstructs technical and business logic, analyzes impacts on various stakeholders, and provides verifiable forward-looking signals, with about 1,800 words.
In the Smoke Quick Test Run#214 on 2026-07-05, GLM-4.6 scored 60.04 on the main leaderboard, with code execution at 88.70, material constraint at 25.00, integrity rating fail, and probe score 0.00.
On July 5, 2026, the YZ Index Smoke Quick Test covered 11 models, with Doubao Pro and Gemini 3.1 Pro tying for first place at 88.54 points. Smoke is a daily 10-question quick test for observing short-term signals and is not equivalent to the Full weekly ranking.
Anthropic accused Alibaba's Qwen lab of using nearly 25,000 fake accounts to interact with Claude over 28.8 million times, aiming to distill its agent reasoning, software engineering, and long-horizon task capabilities.
In 2026, Wired exposed details of Meta's "Cannes" project, which hired hundreds of Kenyan contractors to create fake minor accounts and send prompts involving suicide, self-harm, and child exploitation to ChatGPT and Gemini to test security vulnerabilities.
OpenAI CEO Sam Altman has reportedly proposed transferring approximately 5% of the company's equity to the U.S. government, valued at $42-43 billion. This move is seen as a strategy to navigate complex political environments while sparking debates about public ownership of AI companies.