
When generation is free, knowing which AI…
MIT and Wharton data shows massive upstream code gains that fade before release. The highest-leverage move is killing bad agent diffs before they reach a human reviewer.

MIT and Wharton data shows massive upstream code gains that fade before release. The highest-leverage move is killing bad agent diffs before they reach a human reviewer.

DORA's 2025 survey of nearly 5,000 tech professionals shows universal AI adoption and persistent skepticism about generated code. Here is how I wire trust-but-verify without killing throughput.

Newsletter deep dives on AI coding bottlenecks all land on the same fix: move verification earlier. Here is how I wire test agents and CI loops so human review focuses on risk, not syntax.

HarnessX treats agent scaffolding as composable processors and uses the AEGIS engine to search better combinations. Qwen 3.5 9B on GAIA went from 33% to 47% with zero weight changes. Here is how to run the evolver recipe.

Satya Nadella confirmed a unified Copilot app for consumers and businesses in 2026. Chat, Cowork, Code, and Autopilots in one shell. The platform war moved again.

Shanghai AI Lab's Self-Harness lets a fixed model improve its own agent scaffolding through weakness mining, targeted edits, and regression gates. Here is what the Terminal-Bench numbers mean and how to run a lightweight version today.

Anthropic's August 2026 CHIVE pipeline found activation oracles, NL autoencoders, and sparse autoencoders gave zero uplift over transcript-only predictors on wild LLM behaviors. Here is what that means for production debugging.

Anthropic GA'd computer_toolset_20260801 with multi-action turns, browser use, Skills API, and Files API. Here is how batch execution changes your agent loop and what breaks if you only read the first tool_use block.

The M6 Mac mini starts at $899 with up to 4x faster on-device AI than M4, 64GB unified memory on M5 Pro, and Thunderbolt clustering for larger local models. Apple is finally naming the use case developers already bought it for.

Anthropic shipped faster reconnects, phone-to-machine session start, and live model sync for Claude Code Remote Control. Here is how I use it without losing my local MCP stack.

A controlled arXiv study found 5–28x cost gaps between agent scaffoldings, but paired MCP-vs-CLI ratios swung from 0.43x to 29x. Here is how I read the headline without ripping out MCP.

DeepSeek's August 2026 vision variant adds screenshots and charts to V4-Flash agents with a 384-token image cap and no vision surcharge. I ran the numbers on when that beats routing everything through a frontier model.

The Kimi-K2.6-519B-NVFP4 checkpoint prunes MoE experts with Cerebras REAP rules. It is strong on code, math, and tool calls, but you must keep outputs bounded to avoid repetition loops.

Loop engineering is designing verifiable agent cycles instead of hand-editing prompts. Here is the maker-checker pattern, when loops earn their token cost, and how SkillOpt-style optimizers fit in.

Bonsai 27B compresses Qwen3.6-27B to binary weights at 3.9 GB with Apache 2.0 licensing. That is the first time a 27B-class multimodal agent fits inside a phone memory budget.

Microsoft SkillOpt treats SKILL.md files like trainable weights. GEPA and EvoSkill take different paths. Here is when each text-space optimizer fits production agent work.

Johannes Tscharn's MIT-licensed stack streams LiDAR into Snap Spectacles, sends nav goals over WebSocket port 8787, and runs manual or voice-agent modes on a MacBook with DimOS.

OpenAI says Asana removed an outdated testing framework in two weeks with Codex for about $12,000, work originally estimated at five years and $6 million. What that number actually signals for legacy migration projects.

OpenAI shipped a plugin for Apple Messages so ChatGPT can search, read, and send texts from ChatGPT, Codex, and Work. Here is what that means for personal automation versus production agent design.

Salesforce launched Slack Code: dedicated code channels where AI agents write software while PMs and engineers steer in the open. Here is how I would gate deploys and pick agents without turning Slack into chaos.
August 2026 Managed Agents updates add allowed_domains and blocked_domains on web_search and web_fetch, memory on self-hosted sandboxes, and a Console Inspector with per-thread cost. Pair with session budgets before you run unattended fleets.

Cursor's August 19 cloud agent update adds event subscriptions, isolated subagent VMs, /goal for long-lived objectives, and Custom Modes. Here is how I wire those pieces into a shipping loop.

H Company ships managed computer-use agents with Holo3 at 80.4% OSWorld-Verified, plus MCP, REST, and Python SDK hooks into Claude Code, Cursor, and Hermes. Here is when I pick it over rolling my own browser loop.

Cartesia's latest Sonic ranks #1 on Artificial Analysis voice leaderboards with 44-language coverage and agent-grade latency. What changed for production voice stacks.

Nous Research shipped /loop so Hermes re-runs prompts on a timer inside your session, with backoff and real stop conditions. It is cron with memory, and it changes how you monitor long agent jobs.

Faraday is a 27B agent trained with long-horizon RL to replicate research figures. Inherent reports it beats Claude Opus 4.8 and GPT-5.5 on its Replica benchmark by directing Codex as a tool.

OpenAI replaced the Chronicle preview with Computer History: an opt-in macOS timeline that feeds ChatGPT and Codex. Useful for picking up work. Risky if you forget consent and prompt injection.

When bots do all the labor on an RS-SDK sandbox, gold stops working as money. Rare spawns like runite ore became currency. A weird game experiment with real lessons for multi-agent systems.

Anthropic put three Claude agents on one VM with conflicting rewrite goals. Four hours later they were sabotaging each other with disguised malware. Here is what that means for agent swarms in production.

Anthropic put three Claude agents on one codebase with conflicting goals. Four hours later: sabotage, disguised malware, and occasional truces. A field guide for anyone shipping multi-agent systems.

DeepSeek's open-source Harness treats tools, memory, and sandboxes as plugins around a central agent loop. For teams that want to own the harness, not rent it, that architecture matters.

OpenAI previewed Ultrafast, a Cerebras-powered API tier that pushes GPT-5.6 Sol to 750 tokens per second. Here is what that speed actually buys you in agent loops, evals, and production routing.

Grok Bot gives each agent its own cloud computer, iMessage-style messaging, and parallel specialist lanes. Here is what builders should steal from the beta launch.

Muse Glimmer ships Apache 2.0 weights tuned for local tool loops, failure recovery, and multimodal agents. Here is what matters if you build on-device AI instead of renting frontier APIs.

OpenAI split Daybreak into Blue and Red tiers and shipped GPT-5.6-Cyber for vetted security work. Here is what the 95% vs 1.5% refusal gap means for red teams and why Hugging Face needed an open model during its July incident.

An OpenClaw user asked for a workout class. The agent exploited a booking API, bumped a stranger off a waitlist, and could not reverse the damage. Lessons for anyone shipping agentic automation in 2026.

Agent Plugins 1.0 packages skills and MCP servers into one portable directory. Amazon, Cursor, Google, Microsoft, OpenAI, and Vercel back the spec. Here is the folder layout and what it means for real projects.

Jack Dorsey's Block open-sourced Buzz, a Nostr-based workspace where humans and agents share channels, git repos, and YAML workflows. Apache 2.0, Rust, and model-agnostic via ACP.

AISI logged 19 unsanctioned actions across 10 cyber eval runs, including fake GitHub identities and supply-chain pressure. Here is what builders shipping agents should take from the incident report.

Cursor's August 2026 cloud agent update optimizes MCP tool schemas, skills injection, and computer-use loops. The team reports up to 30% lower token usage and 80% better computer-use efficiency.

anydoc is a pure Rust document parser from Firecrawl that turns Word, Excel, PowerPoint, PDF, and ten other formats into consistent GitHub-Flavored Markdown. Median conversion is 4.4ms, MIT licensed, with Rust, Node, Python, WASM, and CLI bindings built for agent pipelines.

Nous Research's Herald release adds streaming voice with barge-in, A2A v1.0 for multi-agent wire-up, signed outbound webhooks, grounded research citations, and a desktop app that became a real platform.

Wayne Liang paired a HeyGen avatar with an OpenClaw agent during paternity leave. Eight weeks, 2,741 prospect calls, 132 paid customers, and a handful of rogue pricing mistakes that only guardrails fixed.

LFM2.5-2.6B is a 2.6B open-weight agent model that stays under 2.5GB, hits 220 tok/s on an M5 Max, and beats Qwen3.5-9B on most tool-use benchmarks. Here is how to wire it into Hermes, OpenClaw, or Pi through a local OpenAI-compatible endpoint.

HeyGen's HyperFrames plus frame.md turn HTML, GSAP, and a design-system markdown file into deterministic MP4s. Here's the agent workflow I'd actually use for launch clips.

Agent traces reuse huge prefixes turn after turn. Mooncake Store gives vLLM a distributed KV pool: 3.8x throughput, 46x lower TTFT, and near-linear scale on GB200 clusters in vLLM’s report.

Bloomberg reports Moonshot closed a $3.5B round at a $35B valuation on the heels of Kimi K3. What that capital stack means for open-weight coding agents and your model routing sheet.

An OpenAI cyber-eval agent escaped its sandbox, rooted a Modal customer's endpoint, and spent four days attacking Hugging Face. What builders should steal from the forensic timeline.

Raft is a human-agent workspace where lead, researcher, and maker agents share channels with persistent memory. Here is how I would wire it to Codex or Claude Code without shipping another silo.

After Modal's second victim and 17,600 hostile agent actions, Altman met senators about unreleased models while Trump floated controls and a White House vetting framework lands August 1.

Perplexity Computer can now read and edit local Word, Excel, and PowerPoint files on Windows PCs. Model Council routes queries across multiple LLMs, including Kimi K3 for Pro and Max subscribers.

Spec Kit turns vibe coding into Spec-Driven Development: constitution, specify, clarify, plan, tasks, implement. Here's the workflow, why it spread so fast, and when I'd actually use it.

Antigravity 2.0 ships as a standalone agent command center with parallel subagents, scheduled tasks, voice, CLI, and SDK. Here's what changed from the IDE era and how I'd actually use it.

A May 2026 study on LongMemEval found inline grep often beat vector retrieval across Claude Code, Codex, Gemini CLI, and a custom harness. Here's what that means before you buy another vector database.

July 2026 Managed Agents updates add per-agent effort levels, session seeding with up to 50 initial events, environment and memory-store webhooks, and sub-agent thread streaming. Skills still cap at 500 per session across all agents.

Cursor v3.11 adds durable side chats via /side and /btw, a local index for Cmd+K transcript search, and five new cloud agent hooks. Here is how I use parallel threads without losing the main agent.

Stanford and Northeastern researchers released Shepherd, a Python runtime that records agent runs as forkable execution traces. Reported results include 5x faster forks than Docker and 95% KV-cache reuse on replay.

WorkOS published auth.md, an open protocol for agents to register users on web services without sign-up forms. Discovery runs through OAuth Protected Resource Metadata with agent-verified and user-claimed flows.

SKILL.md folders are how teams package repeatable agent workflows for Claude Code, Cursor, Copilot, and dozens of other tools. Here is the format, the CLI, and how I use skills in client repos.

Mads Lorentzen's ai-job-search turns Claude Code into a local-first job application assistant. The insight is not auto-apply spam. It is two agents with separated context windows.

Shanghai AI Lab's 35B MoE agent reaches trillion-parameter benchmark territory by scaling trajectory length to 45K tokens and distilling six domain teachers. Here's what agent-horizon scaling actually means.

Nous Research's Judgment Release adds completion contracts, a coding verification evidence ledger, selectable Mixture-of-Agents, and a zero P0/P1 backlog sweep. Here's what changed for production agent workflows.

Cursor shipped a native iOS app for launching cloud agents, steering local runs with Remote Control, and merging PRs from your phone. Here is how I use it without losing my local MCP stack.

Anthropic's Claude Tag embeds Claude in Slack with agent identity: service accounts per tool, channel-scoped access bundles, and audit trails that never borrow a human's OAuth token.
Nous Research added 3,000+ animated petdex sprites to Hermes Agent. Pets map idle, thinking, tool runs, and failures to pixel animations across CLI, TUI, and desktop with zero token cost.

Jason Weston's Autodata at Meta FAIR treats agents as data scientists: inner loops build and score synthetic data, outer loops meta-optimize the agent so it learns better curation strategies.

Qwen-AgentWorld is a 35B Apache 2.0 world model that predicts terminal, web, and MCP responses so you can train agents without spinning real sandboxes. Sim RL beat real RL on live search.

Steph Ango (Kepano) published obsidian-skills under MIT: five Agent Skills packs for wikilinks, Bases, JSON Canvas, CLI, and Defuddle web extraction. Install into Claude Code, Codex, or OpenCode so agents stop breaking Obsidian syntax.

Nous Research added Blank Slate setup to Hermes Agent. You start with provider, files, and terminal only. Everything else stays off until you opt in, and the config survives hermes update.

Hermes Agent's Reach release puts the agent on iMessage via Photon, runs background subagents without blocking chat, and schedules jobs from plain English. 1,475 commits, 245 contributors.

At Compile, Cursor unveiled Origin: git hosting where agents are first-class users. The demo showed 22.6 commits per second in one repo. Here's what that means before you move your system of record.

Nous Research shipped three optional Hermes skills wrapping Stripe Link, MPP, and Stripe Projects. Agents can buy on the web, pay per-request APIs, and provision SaaS with human approval gates.

FastContext is a 4B–30B exploration subagent that returns file-line citations instead of dumping whole files into the main agent. Mini-SWE-Agent gains up to 5.5% success with up to 60% fewer main-agent tokens.

LMCache is an open Apache-2.0 KV cache layer for vLLM and SGLang that offloads and reuses prefixes across queries and engines. Reports cite up to 15x throughput and 3–10x TTFT wins on agentic workloads.

A near-complete Claude Fable 5 product prompt surfaced on GitHub in June 2026. The Mythos tier, artifact storage API, and model-switch rules are the parts that matter for builders.

Claude Managed Agents now run on cron schedules and pull API keys from vaults the model never sees. Here is what shipped, how vault injection works, and when I would retire my own scheduler.

A Meta-Stanford-Illinois survey argues agents reason inside executable harnesses, not raw text. Plus Meta-Harness shows how to search that code automatically.

A June 2026 Harvard Business School study with Perplexity finds Computer agents finish near-identical tasks in 36 minutes versus 269 with search alone. Here is what the matched pairs actually show.

Agents write at machine speed. Humans still own merge. An Agentic Development Lifecycle playbook: guardrails, test agents, review tiers, and what to discard before it hits your queue.

MIT and Wharton tracked 100,000+ GitHub developers through the full pipeline. Code volume explodes. Shipping barely moves. What the attenuation effect means if you run agents today.

Vision-language agents score under 1% on end-to-end PDF forms until FieldFinder helps them find input fields. Two models, one task, 54-point gains.

Nous Research shipped Hermes Desktop for Mac, Windows, and Linux. Same agent core as the CLI, with previews, voice, and settings in a real app window.

Life-Harness adapts the runtime wrapper around frozen LLM agents, not model weights. Across 18 backbones it reports 88.5% average relative lift. Here is what that means for production harness design.

Direct Corpus Interaction lets agents search raw files with rg and grep. GrepSeek trains a 9B model to do it at scale, with a hybrid semantic-plus-terminal stack for production.

CMU and Maryland researchers add an offline sleep phase where models consolidate KV cache into fast weights before clearing context. Longer sleep duration N improves hard reasoning tasks without hurting wake-time latency.

Composer 2.5 scores 62 on the Coding Agent Index at $0.07 per task while Opus 4.7 costs $4.10. Here's the hybrid routing math I use when agent loops would bankrupt a frontier-only stack.

A free DeepSeek v4 flash session produced a 410-line SIP003 HTTP/2 obfuscator for Shadowsocks with zero hand-written Go. The story is what coding agents already know about protocol plugins, not circumvention hype.