Claude 3.5 Sonnet Leads GPT-4o in Programming Benchmark: 49% Accuracy Ignites Developer Community

Anthropic's Claude 3.5 Sonnet achieves 49% accuracy on the SWE-bench software engineering benchmark, surpassing OpenAI's GPT-4o (33.2%) for the first time in real programming tasks. This breakthrough has sparked tens of thousands of reposts on X platform and heated discussions in the programmer community, with developers sharing real-world cases demonstrating its human-like debugging capabilities.

Claude 3.5 Sonnet Anthropic SWE-bench
1,346

The Battle Over AI Agent Autonomy and Personhood Rights: Silicon Valley's X Platform Ignites the 21st Century's Ideological Battlefield

On February 10, 2026, discussions about AI agent autonomy, personhood rights, and ideological impact surge on X.com, becoming the day's fastest-rising and most controversial AI topic. The debate reflects growing concerns about how autonomous AI systems might reshape ethical boundaries and evolve into the 21st century's greatest ideological battleground.

AI Agents 人格权 自主性
748

Alibaba's Open-Source Qwen2 Model Outperforms Llama3 on Multiple Benchmarks, Bilingual Capabilities Spark Community Buzz

Alibaba Cloud officially released the Qwen2 series of open-source large models in June 2024, with Qwen2-72B-Instruct surpassing Meta's Llama3-70B-Instruct on multiple authoritative benchmarks, achieving an MMLU score of 84.2%. The series' breakthrough in Chinese-English bilingual capabilities has caused a sensation in the open-source community, with reposts on X platform's Chinese community quickly exceeding 30,000.

Qwen2 阿里云 Open Source AI
697