vibehacker
News
Android Developers Blog ·

Google Android Bench 2.0: long-horizon tasks; GPT-6 Astra leads at 28%

Google’s Sept 17 Android Bench 2.0 adds multi-day long-horizon tasks and continuous scoring; full-task pass rates top out around 28% (vs ~91% on the old incremental set). GPT-6 Astra leads the leaderboard; the suite also scores Gemini Flash, Fable 5.1, Kimi K3, and Qwen, with agent runs on Codex and Antigravity.

More news

View all

Claude Code: build-eval and hillclimb tune agents without overfitting

Anthropic’s claude api skill adds /claude api build eval (guided eval design in your repo) and /claude api hillclimb (one change per round tuning with a held out set to catch overfitting). On an internal support bench, hillclimb lifted search accuracy from 74.4% to 98.9% while cutting cost to about one fifth…

Anthropic

Claude Code 2.1.285: disable WebFetch, admins lock API providers

Claude Code 2.1.285 (npm Sept 29) adds CLAUDE CODE DISABLE WEB FETCH to turn off WebFetch and a managed allowedProviders policy so admins can lock machines to Anthropic, Bedrock, Vertex, Foundry, or a cloud gateway. It also ships claude desktop , claude plugin configure , and a fix for URL passwords leaking past log redaction…

Mixed News

OpenAI MCP Events: ChatGPT plugins react via signed webhooks

OpenAI’s DevDay MCP Events let ChatGPT plugins subscribe to MCP server updates (messages, comments, status) and trigger automations over verified signed HTTPS webhooks. Servers need MCP 2.0 (protocol 2026 07 28) with events/list, events/subscribe, and events/unsubscribe; polling and streaming aren’t supported…

OpenAI

Spotted something we missed? Start a thread.