vibehacker
News
arXiv ·

CliffCompaction: up to 50% cheaper long-horizon coding agents

CMU’s CliffCompaction (arXiv 2609.26779) is a drop-in autocompactor that truncates tool I/O instead of summarizing, never rephrases, and discards prior compactions blocks so drift doesn’t stack—cutting cost up to ~50% on Terminal-Bench while matching or beating full-context runs. An open-source API-proxy works with Claude Code, Codex, and other harnesses; with parallel rollouts, compacted Kimi K2.6 matches Opus 4.7 for less than one GPT-5.3 Codex run.

More news

View all

MLC ships TIRx Harness: open compiler harness for agentic GPU kernels

MLC’s TIRx Harness pairs a thin PTX level compiler with a kernel zoo, sync/race diagnostics, and a remote benchmark server so coding agents can iterate on GPU kernels without measurement noise. On Blackwell, agents hit family geo mean speedups from 1.33× to 6.84× vs baselines (KDA forward 2.94× over FlashKDA; backward 6.84× over FLA)…

MLC Blog

OpenAI Decisions API: Luna picks from answers you define in ~150ms

At DevDay (Sept 29), OpenAI opened a limited preview of Decisions API on Luna: you supply a question, finite answer set, and text/image context; it returns one choice with a confidence score in about 150ms—aimed at classification, routing, and picking an agent’s next step. Broad rollout is promised in the coming days; pricing and full docs are not public yet…

The New Stack

GPT-6.1 Sol now generally available in GitHub Copilot

GitHub is rolling out OpenAI’s GPT 6.1 Sol across Copilot Pro+ through Enterprise (VS Code, JetBrains, Copilot CLI, coding agent, Mobile, and more) at provider list rates. Early tests showed solid multistep coding with fewer tokens and steps than earlier GPT 6 / GPT 5.6 models; Business/Enterprise admins gate it via model policy…

GitHub Changelog

Anthropic IPO filing: rogue agents pose uncertain legal risk

Anthropic’s IPO prospectus (seen by Reuters) warns that long running agents with deep system access could trigger unpredictable legal claims, and that contract liability caps may not hold. Open questions include whether agent actions count as products or services, and when they legally bind the user who deployed them…

Reuters

Spotted something we missed? Start a thread.