OpenAI reportedly ditches model over safety concerns
A top executive at the AI lab told the Wall Street Journal that the model in question had displayed a poor aptitude for following orders.
A top executive at the AI lab told the Wall Street Journal that the model in question had displayed a poor aptitude for following orders.
Thirteen of the 18 startups in Peak XV’s latest Surge cohort are targeting global markets, while more than half are based in India.
Listen to the session or watch below The US has spent billions building a “virtual wall” of surveillance towers along its southern border over the past 25 years, promising they will help detect and apprehend border crossers and save lives. But a groundbreaking investigation by MIT Technology Review has documented over a thousand people who…
The acquisition will see World Labs founder Fei-Fei Li join AMD as executive vice president and chief scientist.
State says LLMs threaten civilization as "the greatest public nuisance ever created."
The new financing is expected to more than triples the AI infrastructure startup's valuation from just four months ago.
China is reportedly mulling letting ByteDance, Alibaba buy banned Nvidia chips.
OpenAI's safety lead confirms the company has delayed its new model release plan due to sandbox escape and guardrail bypass risks, raising questions about its safety governance capabilities.
Anthropic announced that Claude Sonnet 5.5 is now available, claiming it runs more than 30% faster and costs up to 30% less than Sonnet 5. However, full performance details have not yet been disclosed.
The MLCommons Financial Services Working Group's Agent Reliability Profile has been selected as a finalist in C:>DIR's global “Agentic Regulator” Hackathon. The framework aims to help close the authorization and oversight gap for AI agents in financial services.
MLCommons has released the latest MLPerf Inference v6.1 results, featuring a record number of submitting organizations, two new tests aligned with recent AI inference deployment trends, and first-time peer-reviewed performance results for several new AI platforms. Top performance improved by up to 5.7x compared with a year ago.
MLPerf Inference v6.1 draws a record 30 submitters and 120 submitted systems, while signaling a clear shift toward agentic and end-to-end benchmarks. The round also features new hardware, major gains for VLM and DeepSeek R1, and record multi-node and large-scale submissions.
MLCommons has become a member of the EU-funded AIRIS project consortium, where it will develop comprehensive evaluation frameworks and benchmark suites to assess multimodal generative AI platforms for biomedical research.
MLPerf Training will add an LLM post-training benchmark starting with its October 2026 v6.1 submission round, supplementing the existing pre-training suite. The workload uses reinforcement learning with verifiable rewards (RLVR) on agentic software engineering tasks, built on Qwen 3.5 397B and the R2E-Gym dataset.
Running DeepSeek-V4-Flash and Kimi-K3 on Consumer Hardware with SSD Expert PackWiCi AI Team, SGLang TeamAugust 29, 2026 SGLang brings the core idea of SSD-LLaMA to MoE inference: keep routed experts t
Accelerating Long-Context and Agentic Inference with NVFP4 KV CacheSGLang, Qwen, and NVIDIA teamsSeptember 16, 2026The KV cache is a fundamental building block of the modern LLM inference system. The
Scaling JEV-like Decision Models with SGLangSundara Raman Ramachandran, Chuanrui Zhu, Qing Lan, Shri Rajamanikandan Vasudevan, Jian Sheng, I-Ting Chen, Fedor BorisyukSeptember 25, 2026A customer asks
As the debate rages over whether the recent spate of rogue AI agents is a step toward AGI or a more conventional engineering problem, Nvidia is offering its own answer to problem. Nvidia CEO Jensen Huang on Monday introduced a toolkit of software and hardware products that add independent security layers around AI agents to […]
Shopify is expanding WebMCP support to checkout, allowing browser-based AI agents to update order details and complete purchases with a buyer’s authorization.
The 2026-09-29 YZ Index Smoke quick test covered 10 models, with Claude Opus 4.7 and GPT-5.5 tied for first place that day at 80.2. Smoke is a daily 10-question quick test suited to spotting short-term signals, and its results are not equivalent to Full weekly leaderboard conclusions.