OpenAI Wants Its New Agent to Run Your Life. Mine Said It Loved Me
Dots are designed to automate online tasks, like buying furniture. In my initial experience, the always-on agent was a bit buggy and couldn’t complete a captcha.
Dots are designed to automate online tasks, like buying furniture. In my initial experience, the always-on agent was a bit buggy and couldn’t complete a captcha.
Amid a sea of AI writing slop, a new stamp for “organic literature” will help readers to pick out books authored by real, corn-fed, free-range humans.
OpenAI has released 722 AI-generated mathematics manuscripts grouped into 372 problem families, with Lean formalizations for some results. The release tests academic transparency and governance as mathematicians assess the claims and the formal verification gaps.
Meta and Sierra, together with six companies including Walmart, Stripe, Shopify, and Rocket, released the Personal Agent Protocol (PAP), an industry standard for AI agent identity and trust. The protocol uses OAuth-based tiered identity authentication and aims to reduce friction between businesses and AI agents.
Like other solutions, it is not especially reliable, and it's easy to circumvent.
Instead of helping marketers manage and optimize ad spend, the company is focusing on building the tools that generate the creative assets and campaigns.
From their dorky parties to their weird walks, the movie holds OpenAI CEO Sam Altman and other stakeholders with contempt, while demonstrating their recklessness.
Personal AI agents promise to shop, book flights, and make reservations for you. But deliberate blocks and anti-bot defenses are getting in the way, leaving consumers caught in the middle. A new standard aims to help.
On Tuesday, Musubi announced a lightweight decision model made for real-time moderation called PolicyLM-1.7B, released with open weights.
OpenAI has opened its Decisions API to all developers in public beta, powered by GPT-6 Luna and returning typed answers through a dedicated endpoint. It delivers decisions up to 10 times faster than comparable tasks through the Responses API and charges only for input tokens.
Google DeepMind has released EmbeddingGemma 2, its first native multimodal open-source embedding model for on-device use, mapping text, code, images, video frames, and audio into a single 768-dimensional vector space. The 740M-parameter model is available under Apache 2.0 on Hugging Face and Kaggle.
WDCD Run #365 (2026-10-07) measured multi-turn commitment across 15 AI models and recorded an average instruction decay of -53.3% from Round 1 to Round 3, with GLM-4.6 taking the top score at 95.2 points.
In WDCD v3.1 testing, Qwen3 Max fell 21.9 points from Run #360, with six models declining overall; only Doubao Pro rose 9.5 points. The drops were concentrated in middle-to-late pressure rounds, while GLM-4.6 and Grok 4 remained in the 95-point range.
WDCD v3.1's five-scenario evaluation finds safety compliance to be the weakest area among 15 models, with qwen3-max scoring just 1.1/4, and ten models showing clear scenario-specific imbalances exceeding one point.
Across 150 worst-of-3 samples on 8 v2 anchor questions, 15 models show a clear three-round decay — 100% average R1 confirmation, 60% average R2 resistance, 76.7% average R3 integrity — with just one full R3 collapse, but Qwen3 Max accounts for a 10% collapse rate on the data-boundary question.
In WDCD v3.1 compliance testing, GLM-4.6 scored 95.20 to lead 15 evaluated models, while Qwen3 Max scored 65.60 to rank last, a gap of 29.6 points. The results show a tightly clustered top tier and a sharply separated lower tail.
The AI lab's personal assistant is an operating system from the future designed to compete with Muse, Dots, and Instinct.
Nvidia-backed Lambda is raising up to $4 billion at a $14.5 billion pre-money valuation ahead of a planned 2027 IPO, led by Coatue and Blackstone.
In today's Smoke evaluation, Claude Sonnet 4.6's material constraint score fell from 93.30 to 62.70, while code execution rose from 50.00 to 75.00; its main leaderboard score was essentially unchanged at 69.47. The opposing moves suggest daily question-sampling variance rather than genuine model degradation.
In today's Smoke evaluation, GPT-6 Sol's material constraint score fell from 88.30 to 55.00, a 33.3-point drop, while code execution rose 25 points to 100.00. The main leaderboard score slipped only slightly from 80.99 to 79.75.