1. 2026-10-10LATESTagents4 min read

    A traffic light for AI agents sharing one Mac

    Netdata alerts become a GREEN/YELLOW/RED verdict that every Claude Code session reads before heavy work — advisory, actionable-only, with the fix in the warning. What it watches, how agents hear it, and what a six-reviewer panel changed.

  2. 2026-10-04claude-code5 min read

    One Claude Code Session, Two Model Upstreams

    How cc_auto keeps its Claude lead on Anthropic while vendor-pinned reviewers run through OpenRouter: frontmatter model selection, request routing, credential separation, and the probes that distinguish a healthy router from a working subagent.

  3. 2026-10-03agents4 min read

    Four agent protocols: MCP, ACP, A2A, AG-UI

    What each agent protocol connects, how its wire works, and when to reach for it — one diagram each.

  4. 2026-09-20agents2 min read

    Driving your logged-in browser from a sandbox

    BrowserSkill's shape in four words — CDP, profile, session, borrow — and why a daemon that cannot live in the sandbox is what makes reusing a real Chrome login possible at all.

  5. 2026-09-03agents14 min read

    Runner Fleet Atlas

    A self-hosted coding-agent fleet, reference-style: topology, lifecycle, identities, what's baked, the numbers, the footguns, and a ranked checklist for preparing agent-friendly environments.

  6. 2026-08-24interviewing11 min read

    The Questions I Ask, By Round

    Most advice gives you a flat list of questions to ask your interviewer. The list isn't the problem — the round is. A peer engineer can tell you what time people left last crunch and has no idea what the promotion bar is. Matching the question to who's in the room, and the one rule for reading any answer.

  7. 2026-08-24agents5 min read

    Swarm Runtime Atlas

    The complete architecture of the autonomous coding swarm on one page — master map, loops over loops, run lifecycle, Turn-Stop handshake, git flow, admission funnel, leg lifecycle, context anatomy, harness loop, delegation, toolbox, robustness matrix, invariants.

  8. 2026-08-19agent-evals12 min read

    Agent Evals That Train: the Harbor + Osmosis Pattern

    Why agent benchmark numbers go wrong, the two paradigms of agent evaluation, a comparison of the major frameworks, and a one-adapter pattern that makes the same verifier serve both benchmarking and RL training.

  9. 2026-08-19agents14 min read

    Lessons of an Autonomous Swarm

    Nine weeks of an agent runtime's git history reviewed commit by commit: eight eras, the five costliest incidents, the silent-failure classes, and the design lessons that survived contact with production.

  10. 2026-08-13agents3 min read

    Verifying an agent self-improvement loop design against the code it assumes

    An LLM-drafted design for an N×M evolve loop looked complete — then three of its load-bearing assumptions died on contact with the three repos it depends on.

  11. 2026-08-13benchmark2 min read

    Do code-graph tools actually save agent tokens? A 4-condition benchmark

    Grep vs three code-graph tools on a real onboarding task: the effective ones halve the token bill — and one is worse than nothing.

  12. 2026-08-13agents3 min read

    Human off the loop — notes on autonomous coding agents

    Interactive and autonomous are two lanes, not one. What it takes to let a coding agent's loop close itself.

  13. 2026-08-13benchmark3 min read

    Do “token optimizer” tools survive contact with a real agent swarm?

    Four community optimizers, nine configurations, three studies on an autonomous coding swarm: cost is run length, not verbosity — and only one combo reliably pays for itself.

Console

  • 14:57:58publishA traffic light for AI agents sharing one Mac
  • 03:15:54publishOne Claude Code Session, Two Model Upstreams
  • 02:47:29publishFour agent protocols: MCP, ACP, A2A, AG-UI
  • 04:12:46publishDriving your logged-in browser from a sandbox
  • 09:33:07publishAgent Evals That Train: the Harbor + Osmosis Pattern
  • 07:35:25publishRunner Fleet Atlas