Google open-sources Mantis for agentic bug finding and fixing
Google released Mantis, an open-source agent harness that finds, sandboxes-reproduces, and patches vulnerabilities — aimed at cutting hallucinated findings from naive AI code scanners.
Google released Mantis, an open-source agent harness that finds, sandboxes-reproduces, and patches vulnerabilities — aimed at cutting hallucinated findings from naive AI code scanners.
At DevDay (Sept 29), OpenAI opened a limited preview of Decisions API on Luna: you supply a question, finite answer set, and text/image context; it returns one choice with a confidence score in about 150ms—aimed at classification, routing, and picking an agent’s next step. Broad rollout is promised in the coming days; pricing and full docs are not public yet…
Hugging Face’s Serge overnight picks persistent Transformers GPU CI failures, reproduces them, writes a patch, verifies both trees on GPU, and only then opens a maintainer PR—29 merges in 80 days ( $43 inference per merged PR on Kimi K 2.7 Code)…
GitHub is rolling out OpenAI’s GPT 6.1 Sol across Copilot Pro+ through Enterprise (VS Code, JetBrains, Copilot CLI, coding agent, Mobile, and more) at provider list rates. Early tests showed solid multistep coding with fewer tokens and steps than earlier GPT 6 / GPT 5.6 models; Business/Enterprise admins gate it via model policy…
Anthropic’s IPO prospectus (seen by Reuters) warns that long running agents with deep system access could trigger unpredictable legal claims, and that contract liability caps may not hold. Open questions include whether agent actions count as products or services, and when they legally bind the user who deployed them…
LiveNerf is a frozen panel, append only eval (Inspect AI + headless Claude Code) that scores Opus 5.5 daily for 30 days to detect quiet post launch regressions. As of Sept 29 it’s on day 6 of 10 in the baseline window; first statistical call lands around Oct 24…
Spotted something we missed? Start a thread.