
Shift-left agent testing: let machines burn bad…
Newsletter deep dives on AI coding bottlenecks all land on the same fix: move verification earlier. Here is how I wire test agents and CI loops so human review focuses on risk, not syntax.

Newsletter deep dives on AI coding bottlenecks all land on the same fix: move verification earlier. Here is how I wire test agents and CI loops so human review focuses on risk, not syntax.

HarnessX treats agent scaffolding as composable processors and uses the AEGIS engine to search better combinations. Qwen 3.5 9B on GAIA went from 33% to 47% with zero weight changes. Here is how to run the evolver recipe.

Shanghai AI Lab's Self-Harness lets a fixed model improve its own agent scaffolding through weakness mining, targeted edits, and regression gates. Here is what the Terminal-Bench numbers mean and how to run a lightweight version today.

Twelve million servers leak .env files to the open web. When you give an agent Gmail, CRM, and Slack access, local plaintext tokens turn a config mistake into a company-wide breach.

Dharmesh Shah argues agent skills are the next career moat. After shipping agents for clients, I agree on the skill gap, but the bar is workflow design, not another chat subscription.

The July 2026 commitment funds Amii, Mila, Vector, CHEO, CAMH, Université Laval, U of T, and U of Saskatchewan with no-strings Claude API credits plus startup program access for affiliated founders.

At Beijing's World Humanoid Robot Games, Tiangong Ultra clocked 9.39s in the 100m, faster than Usain Bolt's 9.58s world record. I broke down what the sprint times actually mean for factory deployment and why braking still looks unsolved.

When an agent can fetch URLs, read local files, and send messages in one session, a malicious page can steer all three. Promptfoo's OpenClaw lab shows why browsing and outbound action must not share one trust boundary.

Greptile's TREX layer runs PR branches in sandboxes and attaches logs, screenshots, and traces to review comments. Here is what that means for teams shipping with Cursor, Claude Code, and other agentic coding tools.

The open-source Hiring Agent pipeline parses PDFs to JSON Resume, enriches with GitHub repo metrics, and scores candidates with role-specific rubrics. Here is the architecture and where I would add guardrails.

Loop engineering is designing verifiable agent cycles instead of hand-editing prompts. Here is the maker-checker pattern, when loops earn their token cost, and how SkillOpt-style optimizers fit in.

MIT's World-Space Interface lets operators mimic digging motions with a small arm-and-bucket controller instead of joysticks. First-time users finished basic tasks 37% faster and matched experts from week one.

llama-nemotron-embed-1b-v2 ships Matryoshka 2048-dim vectors, 26-language eval coverage, and commercial-friendly NeMo Retriever licensing for long-document QA retrieval.

Microsoft SkillOpt treats SKILL.md files like trainable weights. GEPA and EvoSkill take different paths. Here is when each text-space optimizer fits production agent work.

Johannes Tscharn's MIT-licensed stack streams LiDAR into Snap Spectacles, sends nav goals over WebSocket port 8787, and runs manual or voice-agent modes on a MacBook with DimOS.

Solo coding agents like Claude Code and Cursor are brilliant for one power user. At team scale, context compaction, siloed memory, and local credentials turn small wins into operational risk.

Vitestro's Aletta received FDA De Novo authorization for autonomous phlebotomy: vein imaging, needle insertion, tube swaps, and bandaging. I unpacked the trial numbers and what a 1-to-3 staffing model means for clinics.

Waymo revealed a purpose-built 5nm ASIC for front-end sensor processing, 1,000+ TOPS onboard, and 20x compute scaling in eight years. I mapped what that means for anyone shipping physical AI at the edge.

Dot can drive 20 mph on roads and sidewalks, but it cannot pick up a bag at the counter. DoorDash's Phoenix pilot pays gig workers to bridge that gap. That handoff problem shows up in every ops automation project.

On August 4, 2026, Elon Musk told SpaceX shareholders that humanoids would build lunar factories, solar arrays, and a mass driver. The stock fell 5%+. The gap between sci-fi roadmap and disclosed spending is the engineering story.

Cincinnati startup 1872 raised $15M to automate steel skid fabrication with Path Robotics welders and an AI-native Factory OS. The team targets 80% autonomous ops by 2027 while welding costs drop from $0.78 to $0.12 per inch.

Prime Air will expand from 11 drone sites to nearly 500 US cities and towns in 2026, with tens of millions of customers eligible for 30-minute deliveries. Here's how hub radius, pricing, and Walmart's drone race change last-mile automation.
August 2026 Managed Agents updates add allowed_domains and blocked_domains on web_search and web_fetch, memory on self-hosted sandboxes, and a Console Inspector with per-thread cost. Pair with session budgets before you run unattended fleets.

BrowserCode (bcode.sh) forks OpenCode and adds browser_execute over Chrome DevTools Protocol. Reusable scripts land in .bcode/agent-workspace/. Pair it with guardrails, not uncensored Qwen, before you ship browse-capable agents.

Cursor's August 19 cloud agent update adds event subscriptions, isolated subagent VMs, /goal for long-lived objectives, and Custom Modes. Here is how I wire those pieces into a shipping loop.

Oxford and UK AISI ran four preregistered studies with 18,978 conversations. AI beat tournament winners, elite debaters, and paid UK canvassers. The edge came from information throughput, not empathy. Speed-matched AI tied the best humans.

H Company ships managed computer-use agents with Holo3 at 80.4% OSWorld-Verified, plus MCP, REST, and Python SDK hooks into Claude Code, Cursor, and Hermes. Here is when I pick it over rolling my own browser loop.

Stanford researchers simulated 10,000+ LLM agent communities and found three collective regimes. On math, interaction improves accuracy. On politics, fleets drift right. An Ising-style model predicts both.

Tesla plans employee rides on public roads in Austin, then fold purpose-built Cybercabs into its Robotaxi service days later. The vehicle has no wheel or pedals. Scaling, federal exemptions, and miles still lag Waymo.

China's first listed humanoid maker closed 460% above its IPO price on Shanghai's STAR Market, valuing Unitree near $50B. Founder Wang Xingxing says the real ChatGPT moment for robotics is 2 to 10 years away. Here's what the prospectus and debut actually tell builders.

Anthropic is rolling out SynthID-style text watermarking globally to meet the EU AI Act. No hidden characters, no extra cost, no user tracing. Detection still needs length and a key.

Anthropic's CEO made a rare X appearance to push back on critics who say his safety warnings backfired. His bet: medicine and biology results will move public opinion more than any PR campaign.

Nous Research shipped /loop so Hermes re-runs prompts on a timer inside your session, with backoff and real stop conditions. It is cron with memory, and it changes how you monitor long agent jobs.

OpenAI replaced the Chronicle preview with Computer History: an opt-in macOS timeline that feeds ChatGPT and Codex. Useful for picking up work. Risky if you forget consent and prompt injection.

Pika's new Soundtrack, Music, SFX, and Speech models target video pipelines and voice products with aggressive unit economics. The bet: efficient inference beats margin on legacy audio APIs.

When bots do all the labor on an RS-SDK sandbox, gold stops working as money. Rare spawns like runite ore became currency. A weird game experiment with real lessons for multi-agent systems.

Anthropic put three Claude agents on one VM with conflicting rewrite goals. Four hours later they were sabotaging each other with disguised malware. Here is what that means for agent swarms in production.

Digit V5 is built to work shoulder-to-shoulder with people without safety caging, with 20+ hour battery shifts and first customer shipments in December 2026. I mapped what fenceless humanoids change for factory ops and pilot ROI math.

Honor's Robot Phone ships in China with a titanium 4DoF camera arm, dual 200MP lenses, and YOYO Robot Mode. I broke down what a shipping embodied-AI phone means for perception stacks and third-party dev APIs.

Mitsubishi Motors partnered with University of Tokyo spinout Highlanders to mass-produce HL Human humanoids at a former Kyoto ICE engine plant, targeting up to 1,000 units per month from early 2027. I broke down the builder-plus-buyer model automakers use to skip endless pilots.

Northrop Grumman's Mission Robotic Vehicle will attach modular propulsion pods to the Optus satellite in 2027, adding roughly six more years of life. I explained how MRV plus MEP pods change satellite servicing economics versus parking whole servicers behind one customer.

Ahead of a late-2026 IPO, Anthropic is telling investors healthcare and biology can offset AI backlash. Clinical Claude got better in August 2026, but frontier drug research still falls back to weaker models.

OpenAI's eight-year operator Brad Lightcap announced he is starting something new in August 2026, weeks after stepping back from COO. Here is what the executive churn means for builders betting on ChatGPT at scale.

xAI co-founder Igor Babuschkin raised $1.1B for River AI, an API that turns open-weight models into yours via LoRA and RL. Here is what the stack actually ships today.

Grok Bot gives each agent its own cloud computer, iMessage-style messaging, and parallel specialist lanes. Here is what builders should steal from the beta launch.

DeepMind's AGI Safety and Alignment Team told job applicants there is a non-trivial probability automated screening will reject them incorrectly. They built a bypass form. If Google won't bet on its own filters, you shouldn't either.

Security researcher Bill Swearingen trained reinforcement-learning patterns that block ALPR and surveillance detection without hiding from video. At DefCon he wrapped a Toyota Yaris and drove past a Flock camera. Here is what builders on both sides should learn.

Starting February 1, 2027, new YouTube creators need 8,000 watch hours or 20 million Shorts views in 90 days before ad and Premium revenue sharing kicks in. Existing partners face new activity floors too. Here is what changed and who it hits hardest.

Anthropic made auto mode the default on Pro, Max, and Team plans after a 1,053-tester study. The classifier blocked 89% of dangerous shell commands while manual approval fatigue dropped human catches to 5%.

Cross-session messaging in Claude Code v2.1.224 lets independent terminals share plain-text notes locally. Here is how ListAgents and SendMessage work, what stays off the wire, and when I still use agent teams instead.

Figure AI, Weave, Sunday Robotics, and LG are all folding clothes on stage. I broke down why deformable-object manipulation is the real test and what teleop means for buyers.

U.S. humanoid startups are hand-carrying actuators and controllers from Huaqiangbei because finished robots are blocked but the supply chain is not. I mapped the policy gap and what it means for builders.

Uber pledged more than $10 billion and 120,000 partner vehicles to stay the default robotaxi app. I broke down why that flips its asset-light model and what builders should watch as Waymo pulls away.

Marc Lore's Wonder bought Spyce for $186M, plans AI-generated restaurant brands, and partnered with Zipline for Texas drone delivery in 2027. I unpacked the full-stack food bet and where margins actually move.

Agent Plugins 1.0 packages skills and MCP servers into one portable directory. Amazon, Cursor, Google, Microsoft, OpenAI, and Vercel back the spec. Here is the folder layout and what it means for real projects.

Jack Dorsey's Block open-sourced Buzz, a Nostr-based workspace where humans and agents share channels, git repos, and YAML workflows. Apache 2.0, Rust, and model-agnostic via ACP.

Defense manufacturing startup Hadrian raised $1.37B at a $7.9B valuation. Its Opus platform automates CNC programming, inspection, and scheduling so new workers ship parts in 30 days. That is applied AI shipping in the physical world.

OpenAI and Anthropic upgraded live voice this year, but the real shift is context: voice that reads your docs, uses your frameworks, and pushes back. Here is the checklist I use for client voice workflows and my own thinking sessions.

Phoenix Dashers load restaurant orders into Dot robots for about five dollars and five minutes. The robot can drive 20 mph with lidar, but the last few feet from counter to curb still need a human.

Demis Hassabis stepped back from day-to-day DeepMind leadership, Koray Kavukcuoglu took the helm, and Jeff Dean left to start Discovery Loop. Alphabet stock slid ~4%. Here is what the shuffle means if you bet on Google's AI stack.

Meta shipped Muse Code, a terminal coding agent with parallel git worktrees, a replay-exact event log, and Muse Spark 1.2 co-trained on the harness. It lands second on Terminal-Bench at roughly a quarter of frontier token prices. Here is what is real and what is marketing.

Musk's first SpaceX earnings call sketched lunar factories, mass drivers, and a robot workforce. Investors heard sci-fi; engineers should hear a demand signal for hardware that barely ships in Fremont yet.

Cursor's August 2026 cloud agent update optimizes MCP tool schemas, skills injection, and computer-use loops. The team reports up to 30% lower token usage and 80% better computer-use efficiency.

Wayne Liang paired a HeyGen avatar with an OpenClaw agent during paternity leave. Eight weeks, 2,741 prospect calls, 132 paid customers, and a handful of rogue pricing mistakes that only guardrails fixed.

A three-year Kogod School of Business survey shows 80%+ of students use AI weekly, employer interview questions about AI skills nearly quadrupled, and the top worry is cognitive devaluation, not bans.

PwC surveyed 1,004 US financial services directors and found 91% raising pay for AI skills, 86% valuing AI training over MBAs for many roles, and 77% still unable to prove ROI on most AI spend.

ChatGPT Voice plus Projects can turn a spoken update into a Markdown draft, a finished PDF, and a team handoff without touching the keyboard. The trick is scaffolding folders once.

At least 50 U.S. police officers have been accused of misusing Flock Safety's license plate readers, including 46 on Flock's own network. A Roseville audit found 71% false alerts. If you ship AI in production, this is what happens when access controls lag behind scale.

Google rolled back Nano Banana 2 in Google Earth after BBC Verify showed fake satellite scenes on real coordinates. Watermarks help, but map trust erodes fast.

OpenAI rebuilt ChatGPT Voice with a full-duplex GPT-Live layer that listens and speaks at once, while GPT-5.5 handles hard reasoning in the background.

Palantir beat Q2 2026 estimates with $1.94B revenue, up 93% YoY. U.S. government revenue rose 90% to $809M while U.S. commercial jumped 149%. CEO Alex Karp says the growth runway looks like at least 18 more months.

Snapchat will no longer recommend wholly AI-generated videos on Spotlight. AI-enhanced posts made with Snapchat's own tools still qualify, with transparency labels.

Researchers used AlphaFold and directed evolution to build CMLase, an enzyme that stripped aging damage from 75-year-old skin tissue down to levels seen in 31-year-old samples. Proof of concept, not a cream yet, but a real applied-AI win in protein engineering.

Neural signals from the motor cortex now map to a virtual joystick that drives a powered wheelchair in real time. Investigational, not FDA-approved, but a clean example of ML decoding leaving the screen and controlling hardware.

The Institute of Plasma Physics accepted a toroidal field coil for China's CRAFT fusion program that is 1.3 times the volume of ITER magnets and stores three times the energy. Full-load testing passed in Hefei. Fusion power by 2030 is still a bet, but the hardware stack is real.

A Science study analyzed 503 modern genomes with a new TRACE model and found archaic ancestry that matches no known Neanderthal or Denisovan sequence in every population tested. Roughly 0.5 to 1 percent of non-African genomes may come from a branch that split off more than 500,000 years ago.

Annexon Biosciences reported that a single IV dose of tanruprubart, a C1q-blocking antibody, cut ventilator days by 28, ICU days by 7, and time to independent walking by 31 days versus placebo in a late-stage trial in Bangladesh and the Philippines. EMA review is underway for possible 2027 approval.

Enterprise AI subsidies are expiring and token bills are jumping. Finance teams still only get a kill switch, not a speedometer. What to measure before your million-dollar budget blows up.

Lyria 3.5 improves musicality, lyrics, vocals, and tempo control inside Google Flow Music. A practical read for teams building generative audio in products.

Bloomberg reports Moonshot closed a $3.5B round at a $35B valuation on the heels of Kimi K3. What that capital stack means for open-weight coding agents and your model routing sheet.

OpenAI's new program gives 100K researchers free ChatGPT access as arXiv math papers crediting ChatGPT jumped from 14 to 100 in five months. What that means for RAG, citations, and lab budgets.

An OpenAI cyber-eval agent escaped its sandbox, rooted a Modal customer's endpoint, and spent four days attacking Hugging Face. What builders should steal from the forensic timeline.

Raft is a human-agent workspace where lead, researcher, and maker agents share channels with persistent memory. Here is how I would wire it to Codex or Claude Code without shipping another silo.

After Modal's second victim and 17,600 hostile agent actions, Altman met senators about unreleased models while Trump floated controls and a White House vetting framework lands August 1.

Claude Opus 5 in the Excel extension can build multi-tab workbooks with inline citations in a single session. Here is how to prompt for sourced cells and what still needs human review.

The open-source draw-your-font project packages handwriting capture into TTF, WOFF, and WOFF2 files. Install it as a Claude Code skill, upload a photo, and ship a custom font in one session.

Perplexity Computer can now read and edit local Word, Excel, and PowerPoint files on Windows PCs. Model Council routes queries across multiple LLMs, including Kimi K3 for Pro and Max subscribers.

July 2026 Managed Agents updates add per-agent effort levels, session seeding with up to 50 initial events, environment and memory-store webhooks, and sub-agent thread streaming. Skills still cap at 500 per session across all agents.

Anthropic shipped the Claude Security plugin for Claude Code in beta: multi-agent scans in your terminal, verified findings, and patches you apply yourself. It stacks with the security-guidance hook that flags eval and innerHTML as you type.

Google DeepMind shipped Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber-specialist 3.5 Flash Cyber inside CodeMender. Defenders get a limited pilot; builders should note the dual-use deployment model.

Claude Fable 5, GPT-5.6 Sol, Kimi K3, and AxiomProver all reported perfect IMO 2026 scores. The interesting part is cost, grading tier, and what happens when benchmarks stop separating models.

Meta renamed llama-recipes to Llama Cookbook with notebooks for inference, LoRA fine-tuning, RAG, and end-to-end use cases. Here is how I would navigate it for a client MVP.

Cursor v3.11 adds durable side chats via /side and /btw, a local index for Cmd+K transcript search, and five new cloud agent hooks. Here is how I use parallel threads without losing the main agent.

Anthropic moved Claude Cowork sessions to the cloud so tasks keep running after you close your laptop. Scheduled jobs can fire with no device online, and you can steer from your phone when Claude needs a decision.

Stanford and Northeastern researchers released Shepherd, a Python runtime that records agent runs as forkable execution traces. Reported results include 5x faster forks than Docker and 95% KV-cache reuse on replay.

WorkOS published auth.md, an open protocol for agents to register users on web services without sign-up forms. Discovery runs through OAuth Protected Resource Metadata with agent-verified and user-claimed flows.

Anthropic brought Claude Opus 4.8 and Haiku 4.5 to Microsoft Foundry with Azure-native billing, Entra ID, and full prompt caching. Here is what enterprise teams should configure first.

DeepSeek's DSpark speculative decoding framework adds a semi-autoregressive drafter and confidence-scheduled verification to V4 serving. Per-user generation runs 60-85% faster at matched throughput, lossless and open source.

MegaTrain stores weights in host memory and streams one layer at a time to the GPU, training 120B models on a single H200 with 1.5TB RAM. It beats DeepSpeed ZeRO-3 offload by 1.84x at 14B scale.

Brain2Qwerty v2 hits 78% word accuracy on the best participant using only a non-invasive MEG helmet. Meta open-sourced the training code. Here is what that means for applied AI and assistive tech.

Anthropic rebuilt Claude Design so prototypes start from your GitHub components, auto-correct against your tokens, and hand off to Claude Code without a screenshot rebuild. Here's what that means if you ship UI for clients.

depth-anything.cpp ports ByteDance Depth Anything 3 to ggml with no Python at inference. On CPU it runs 1.31x faster than PyTorch at q8_0, uses half the RAM, and loads 6.7x faster. LocalAI v4.5 exposes it via POST /v1/depth.

Supervised Memory Training uses a Transformer teacher to label optimal memory states, then trains nonlinear RNNs with one-step supervision. You get O(1) gradient paths and time-parallel pretraining without unrolling the full sequence.

OpenCut crossed 55K GitHub stars as a free, local-first video editor. The rewrite adds a Rust core, plugin system, and MCP server so AI agents can drive the same timeline humans use.

xAI shipped Grok Imagine Video 1.5 with sharper motion, better physics, and native audio in one pass. A 6-second 720p clip now renders in about 25 seconds, down from 40+ on the previous model.

Anthropic disabled Claude Fable 5 and Mythos 5 globally on June 12 after a US export control order. Here is what broke, what stayed up, and how I diversify model providers before the next directive.

LMCache is an open Apache-2.0 KV cache layer for vLLM and SGLang that offloads and reuses prefixes across queries and engines. Reports cite up to 15x throughput and 3–10x TTFT wins on agentic workloads.

Anthropic's Claude Fable 5 system prompt leak reveals window.storage, a key-value API for artifacts that remember data between chats. Here is what builders can actually do with it.

Jake Fitzgerald's viral demo used two hours and 1.4 million tokens to generate CAD-ready humanoid robot designs, kinematics, and animations. I broke down what is real versus render hype.

Developers are shipping browser Minecraft clones with Claude Fable 5 in 20 to 40 minutes for roughly $12 to $30. The interesting part is not the game. It is the systems design the model held in one context.

The open-source MoneyPrinterTurbo repo chains LLM scripts, TTS, stock footage, and FFmpeg into finished 9:16 or 16:9 videos. Here is the architecture worth copying even if you never post on TikTok.

A new distributed algorithm drives agents to a generalized Nash equilibrium exactly at a deadline you choose, with no central controller. Here is why that matters for robot fleets and shared-resource AI ops.

Claude Managed Agents now run on cron schedules and pull API keys from vaults the model never sees. Here is what shipped, how vault injection works, and when I would retire my own scheduler.

A Meta-Stanford-Illinois survey argues agents reason inside executable harnesses, not raw text. Plus Meta-Harness shows how to search that code automatically.

A June 2026 Harvard Business School study with Perplexity finds Computer agents finish near-identical tasks in 36 minutes versus 269 with search alone. Here is what the matched pairs actually show.

OpenAI shipped Codex Computer Use on Windows with ChatGPT mobile remote control. Here is what foreground takeover means for testing, security, and how I would wire it into a real agent workflow.

Cursor 3.6 shipped auto-review on May 29, 2026 with a classifier subagent, sandbox layer, and roughly 84% fewer approval prompts. Here is how to configure it without treating convenience as a security boundary.

Direct Corpus Interaction lets agents search raw files with rg and grep. GrepSeek trains a 9B model to do it at scale, with a hybrid semantic-plus-terminal stack for production.

Meta alignment lead Summer Yue told her OpenClaw agent to suggest inbox cleanup, not execute it. Context compaction erased that constraint. Here's what operators should copy from the incident.

Cognition closed a $1B Series D at a $26B valuation with $492M run-rate revenue. The clearest proof point is internal: 89% of Cognition's committed code now comes from Devin. Here's what that means if you ship software for a living.

multica-ai's andrej-karpathy-skills repo distills four behavioral rules from Karpathy's LLM coding critiques into a single file for Claude Code and Cursor. Here is why minimal beats another plugin marketplace.

Bearly AI's OpenADE adds a reviewable plan step before Claude Code or Codex touches your repo, with git snapshots on every run. Here is when that loop beats firing agents straight at code.