vibehacker
News
RuntimeWire ·

Artificial Analysis Coding Agent Index now reports safety refusals

Artificial Analysis’s Coding Agent Index v1.5 (announced Sept 18) now reports safety refusals, splitting blocked zeros from recoverable fallbacks when an agent switches models or continues. The chart sits beside DeepSWE, Terminal-Bench 4.0, and SWE-Atlas-QnA scores so builders can see how often security or terminal work dies on a policy gate.

More news

View all

MCP Python SDK: malicious servers could steal OAuth credentials

Cycode found the official MCP Python SDK (1.9.1–1.29.1 / 2.0.0–2.1.1) could skip issuer checks on OAuth discovery fallbacks, letting a malicious MCP server harvest client secrets, auth codes, and PKCE keys (GHSA qx49 fqc8 xw99, CVSS 7.5). Fixed in 1.30.0 / 2.2.0—machine to machine providers also need an explicit issuer= argument…

Cycode

Claude Code: build-eval and hillclimb tune agents without overfitting

Anthropic’s claude api skill adds /claude api build eval (guided eval design in your repo) and /claude api hillclimb (one change per round tuning with a held out set to catch overfitting). On an internal support bench, hillclimb lifted search accuracy from 74.4% to 98.9% while cutting cost to about one fifth…

Anthropic

Claude Code 2.1.285: disable WebFetch, admins lock API providers

Claude Code 2.1.285 (npm Sept 29) adds CLAUDE CODE DISABLE WEB FETCH to turn off WebFetch and a managed allowedProviders policy so admins can lock machines to Anthropic, Bedrock, Vertex, Foundry, or a cloud gateway. It also ships claude desktop , claude plugin configure , and a fix for URL passwords leaking past log redaction…

Mixed News

Spotted something we missed? Start a thread.