Teardownspublic

Agent harnesses

What to run an agent on top of. Twenty taken apart, and exactly one tool actually came in.

20 written · 1 adopted

Things that provide the floor an agent runs on: rule sets, skill directories, self-improving harnesses, hook dispatchers, model gateways. Someone else’s operating system, opened up and held against my own.

This is the biggest bucket — twenty of the forty. Exactly half. The categories weren’t drawn first and filled in; they were counted after the fact. That number is where my time actually went.

Across twenty pieces, one tool came in

Subject The tool What stayed
ECC Adopted, then it wouldn’t run in my environment Managing hooks by strength tier rather than on/off
mattpocock/skills Not installed — because of the repo’s own argument That cost comes in two kinds. A machine’s context load and a person’s cognitive load are handled differently
Prime Agent Rejected — zero approval gates Don’t re-run a failure gate when the artifact hasn’t changed
Clawd on Desk Adopted — the only one that actually came in The installer plants its own uninstaller · append, never overwrite
Odysseus Not recommended Escape delimiter literals before wrapping untrusted text
ponytail Wholesale rejection · one line kept “Grep every call site” — what transferred was the shape of the sentence, not its meaning
OmniRoute Blocked — and the seven harvestable techniques rejected with it Draw the boundary at the repository, not at the piece
Cross-session messaging Built nothing — it was already on The real failure mode isn’t security, it’s a sender misidentifying itself
Buzz group chat On hold — platform unsupported For verification, split models rather than personas, or agreement carries no information
Four plugins Two of four were ones I’d already rejected The counterexample lives in the promoter’s own table, not in a critique
Meta-prompting Procedure already in use · one line kept Eliciting questions is really about seeing the holes in my own context
A field guide to Fable 3 of 6 techniques A third path that leaves a trace without stopping — record the deviation and keep going
obra/superpowers Rejected — and the haul fell from 3 to 1 Measure only your own side and the number confirms what you already believed
ASIDE Rejected — it won’t even install Whether you check “it isn’t there” on the download page or in the changelog
One skill, ad images 1 takeaway, and that one an estimate A good design and a good conclusion don’t arrive as one package
DeepSeek Harness Rejected — the install died after 300 silent seconds What was missing wasn’t a tool but a record of ever running one
ip-as-logo Rejected — nothing here to call it from A rule in a file the agent never reads has zero runtime effect
anti-slop 27 hits, none fixed Count a gate by whether it runs, not by whether it exists
DSH better-sidebar Rejected · 4 takeaways → 1 Check that the tool you doubt others with is committed in your own repo
cumora Deferred — the only one here When two numbers agree, check they are looking at the same thing

Nineteen of twenty brought in no tool, and the one that did was not justified on productivity grounds. Its page says so outright — that I installed it because I wanted to — because otherwise the memory rewrites itself into “I added it for the efficiency.”

That probably isn’t a coincidence. A harness is the kind of thing you have to swap wholesale for something you’re already running, so the payoff has to beat one technique by a wide margin — and that threshold turns out to be very high.

Which changes how this bucket is read

The other buckets get read as “what can I take.” This one is closer to “look at what I’m already running, through someone else’s eyes.”

ECC is the extreme case. It wasn’t someone else’s repo — it was the upstream of my own configuration, copied out as files six months ago and never synced since. That was the first time the drift got measured.

Two more landed in the same place. In Odysseus, of five disciplines the docs declare, the code enforces one. In ponytail, the investigation turned up not their rule set’s problems but three contradictions inside mine. Open somebody’s operating system and what you usually see first is your own.

The second line — “I already had this”

At six, only one pattern showed: no tools come in. Six more, and a different conclusion turns out to be the more frequent one — “I already had this.”

Cross-session messaging is the extreme case. The mechanism the video dug up doesn’t exist on my OS at all, and yet the capability I actually wanted was already switched on. ECC wasn’t someone else’s repo but the upstream of my own config. The ponytail investigation surfaced not their rule set’s problems but three contradictions in mine. Even rejecting all seven of OmniRoute’s techniques comes down to the same axis — harvesting a piece means keeping the original open in front of you, and there is no reason to do that for something I can already do.

So the work I do most often in this bucket isn’t evaluating someone else’s thing. It’s taking inventory of my own.

At twenty, this line gained a layer

Not “I already had this” but “I have it and have never once switched it on.”

I went into a harness with 181,170 stars meaning to take away five techniques, and all five died against two skills already sitting in my own folder with zero run history. One leaves no output anywhere in six months of records; the other’s log file and trash folder — both of which it instructs you to create — do not exist. It never ran.

This wasn’t a one-off. The same skill turned up again in the context-budget piece: installed 24 days ago, run zero times. Two investigations that knew nothing about each other pointed at the same object.

The unit of the inventory was wrong. Count only “what do I have” and you conclude you should go buy the thing that is already in the warehouse. What needs counting is not the list but the date it was last switched on.

The third line — the target was holding the tool that broke it

At fifteen, something else came into view. In four pieces, the decisive evidence that overturned a verdict came from the subject’s own material. Four of the five new ones did it again, so it is eight now.

Subject Where the decisive evidence sat
Four plugins the promoter’s own table
obra/superpowers their own rule — “if the control shows no failure, stop.” So I stopped
ASIDE their own changelog — what the download page lacks appears three times
One skill, ad images the second table in the same talk, which reverses the title’s premise
ip-as-logo its own two documents disagree — the headline rule is in the README and not in the executable file
anti-slop 21 lines of its own source, which break four of its own rules
DSH better-sidebar the same repository carries three different numbers at once
cumora its own CI guard — call it directly and it catches none of the repo’s own idioms

There is one exception. Only DeepSeek Harness had its decisive evidence outside — a 404 from the package registry and community reports. Even that was not a critique. Nobody commented; a machine answered.

There was no need to go looking for outside criticism. If anything, the outside criticism was wrong twice — a quoted figure eight months stale, and a verdict of stagnation that hadn’t looked at the branch layout.

So the adversarial pass in this bucket has a fixed order now. Before finding out what anyone else said, check whether what the subject said about itself hangs together.

Five arrived at once, and this time I didn’t pick them

Up to fifteen I picked by hand out of a link library. The five that took this bucket to twenty came from a feed a machine scrapes every dayJust out walks six sources twice a day, and I took apart every GitHub item it had stacked up, in order, all of them. This is the first time that path was actually used.

Changing who picks changed what the sample is. When I pick, I pick things that look like they touch my problems, and the verdicts converge on “I already have this.” When the feed picks, things unrelated to my problems come in — one of the eight was Apple-only and could not be run here at all, so it was not published — and in exchange, places I would never have opened get opened. Three of the five found defects in my own setup, and all three were subjects I had no reason to pick.

The verdicts came out harsher, not softer. Of the five: zero adopted, one deferred, four rejected. Stars don’t carry usefulness with them — add up the five and it clears 190,000, and the number of tools that came in is zero.

And one thing the ledger can’t show. Prime Agent has two reports — the repository analysis and a video. The ledger counts subjects rather than files, so it’s one row; and the piece above was written from reading the code, so nothing from the video is in it yet. That’s the kind of fact a table can’t hold, so it goes here.