vibehacker
News
The New Stack ·

Opus 5.5 matches Opus 5 on hard reasoning, costs less

A New Stack side-by-side on three hard reasoning problems found Opus 5.5 matched Opus 5’s answers while spending less—43% cheaper on a logic grid and 69% on a stone game—mostly from fewer output tokens. Measured write speed was only ~11% faster, short of Anthropic’s 30% claim; both models still burned long token budgets with no answer on a combinatorial ordering task.

More news

View all

OpenClaw Enterprise: free MIT control plane for persistent AI agents

OpenClaw shipped OCE, a free MIT licensed, self hostable control plane (Docker/Kubernetes) for multi tenant agent deploy with permissions, sandboxing, and audit—framed as “Kubernetes for agents.” The project started at OpenAI and now lives under the OpenClaw Foundation with Red Hat and Nvidia; OpenAI and Red Hat are already piloting it (still pre 1.0, aimed at internal pilots)…

VentureBeat

MLC ships TIRx Harness: open compiler harness for agentic GPU kernels

MLC’s TIRx Harness pairs a thin PTX level compiler with a kernel zoo, sync/race diagnostics, and a remote benchmark server so coding agents can iterate on GPU kernels without measurement noise. On Blackwell, agents hit family geo mean speedups from 1.33× to 6.84× vs baselines (KDA forward 2.94× over FlashKDA; backward 6.84× over FLA)…

MLC Blog

OpenAI Decisions API: Luna picks from answers you define in ~150ms

At DevDay (Sept 29), OpenAI opened a limited preview of Decisions API on Luna: you supply a question, finite answer set, and text/image context; it returns one choice with a confidence score in about 150ms—aimed at classification, routing, and picking an agent’s next step. Broad rollout is promised in the coming days; pricing and full docs are not public yet…

The New Stack

GPT-6.1 Sol now generally available in GitHub Copilot

GitHub is rolling out OpenAI’s GPT 6.1 Sol across Copilot Pro+ through Enterprise (VS Code, JetBrains, Copilot CLI, coding agent, Mobile, and more) at provider list rates. Early tests showed solid multistep coding with fewer tokens and steps than earlier GPT 6 / GPT 5.6 models; Business/Enterprise admins gate it via model policy…

GitHub Changelog

Spotted something we missed? Start a thread.