vibehacker
News
Hugging Face ·

Prefix cache decides: Qwen3.8-27B NVFP4 serves agents on two RTX 5090s

Scalably’s Sept 19 field report: two RTX 5090s ran Unsloth’s Qwen3.8-27B NVFP4 via vLLM 0.27.0 for 14 days of real agent traffic—82.6% prefix-cache hits, 28k requests, zero engine errors. A faster MoE lost the job on tool recovery (1/20 vs 12/20), so they kept the dense model.

More news

View all

Claude Code: build-eval and hillclimb tune agents without overfitting

Anthropic’s claude api skill adds /claude api build eval (guided eval design in your repo) and /claude api hillclimb (one change per round tuning with a held out set to catch overfitting). On an internal support bench, hillclimb lifted search accuracy from 74.4% to 98.9% while cutting cost to about one fifth…

Anthropic

Claude Code 2.1.285: disable WebFetch, admins lock API providers

Claude Code 2.1.285 (npm Sept 29) adds CLAUDE CODE DISABLE WEB FETCH to turn off WebFetch and a managed allowedProviders policy so admins can lock machines to Anthropic, Bedrock, Vertex, Foundry, or a cloud gateway. It also ships claude desktop , claude plugin configure , and a fix for URL passwords leaking past log redaction…

Mixed News

OpenAI MCP Events: ChatGPT plugins react via signed webhooks

OpenAI’s DevDay MCP Events let ChatGPT plugins subscribe to MCP server updates (messages, comments, status) and trigger automations over verified signed HTTPS webhooks. Servers need MCP 2.0 (protocol 2026 07 28) with events/list, events/subscribe, and events/unsubscribe; polling and streaming aren’t supported…

OpenAI

Spotted something we missed? Start a thread.