vibehacker
News
The Verge ·

Microsoft Copilot app adds Code tab and Autopilot agents

Microsoft unveiled a redesigned Copilot desktop app with Home, Code, and Autopilot tabs: Code uses GitHub Copilot tech to build sandboxed, tenant-hosted internal apps, and Autopilot (ex-Scout) is a persistent enterprise agent with its own identity and cloud computer. Home and Code roll out to Frontier users in coming weeks; Autopilot enters private preview later this month.

More news

View all

Anthropic: open-weight GLM-5.3 near Mythos on end-to-end exploits

Anthropic finds Z.ai’s freely downloadable GLM 5.3 builds end to end exploits at rates close to Claude Mythos Preview (50/410 vs 56/410 on ExploitBench), while simple jailbreaks and weight abliteration bypass its safeguards 64–100% of the time—unlike safeguarded Claude models in the same tests…

Anthropic

CodeScene: agents refactor 300K-line C game for ~$4k in three weeks

CodeScene’s agents (mostly Claude Code + Opus) refactored a 300K line Street Fighter III decompilation in three weeks for $4k—2,903 commits, Code Health 5.6→10.0—guided by a CodeHealth MCP score and frame by frame replay checks; they also accumulated a 22 recipe refactoring playbook…

InfoQ

OpenClaw Enterprise: free MIT control plane for persistent AI agents

OpenClaw shipped OCE, a free MIT licensed, self hostable control plane (Docker/Kubernetes) for multi tenant agent deploy with permissions, sandboxing, and audit—framed as “Kubernetes for agents.” The project started at OpenAI and now lives under the OpenClaw Foundation with Red Hat and Nvidia; OpenAI and Red Hat are already piloting it (still pre 1.0, aimed at internal pilots)…

VentureBeat

MLC ships TIRx Harness: open compiler harness for agentic GPU kernels

MLC’s TIRx Harness pairs a thin PTX level compiler with a kernel zoo, sync/race diagnostics, and a remote benchmark server so coding agents can iterate on GPU kernels without measurement noise. On Blackwell, agents hit family geo mean speedups from 1.33× to 6.84× vs baselines (KDA forward 2.94× over FlashKDA; backward 6.84× over FLA)…

MLC Blog

OpenAI Decisions API: Luna picks from answers you define in ~150ms

At DevDay (Sept 29), OpenAI opened a limited preview of Decisions API on Luna: you supply a question, finite answer set, and text/image context; it returns one choice with a confidence score in about 150ms—aimed at classification, routing, and picking an agent’s next step. Broad rollout is promised in the coming days; pricing and full docs are not public yet…

The New Stack

Spotted something we missed? Start a thread.