You don't have it if it has never run
Being installed is only the first of three layers. Three investigations pointed at the same two tools, and all three times those tools had never run.
“We already have this” is a sentence that has to be checked three times — is it installed, does pressing it produce output, and is there a date you last ran it. Taking apart forty tools, I paid for collapsing those three into one column more than once. A skill picked up specifically to measure context budget had zero runs across the 24 days from install to check, and the command meant to invoke it was never created at all (all 37 commands checked, as of 2026-08-22). But that single case is not where this piece earns its keep. Three pieces named the same two tools in sequence, and none of the three ran them. The first named them as “the top priority this investigation leaves behind” and closed. The second quoted that first piece and then went off to take apart a different repository to fill the same gap. In the third, one of them turned up again as “installed 24 days, run 0 times”. This is not a coincidental overlap — it is knowing and repeating anyway, which makes it a process failure rather than a signal. Layer two (having ≠ working) is covered in all four times the numbers were right, so it gets one paragraph here. This piece puts its weight on layer three: a tool that runs perfectly well when pressed, sitting at zero because nobody presses it. The symptom is silence, so the ledger keeps saying “have it”, and that “have it” becomes the grounds for the next adoption verdict — producing the conclusion to buy from outside what is already in the storeroom.
Questions this answers
- I checked what is installed in my setup — is that enough to write “we already have it”
- Why do tools get installed and then never used
- Why does the same investigation come round again six months later
- How do I stop adopting something I already have
- How far should an inventory check go before bringing a tool in
- Can I find out when a skill or tool last ran
- How do I actually count tools that are installed but never run
- What is the problem with copying in someone elses preset bundle
- I found a defect in a tool that never ran — should I fix it now
- Should I delete an unused tool or keep it
The full procedure for deciding whether to bring someone else’s tool in is in Forty rejections, and the reason was rarely the tool. This piece covers step zero of that procedure: the inventory check.
“We already have this” is three different questions
Possession, function and run history are different questions, and each layer has its own symptom and its own remedy. Layer one is answered by looking at a list, layer two by pressing once, layer three by hunting for traces the tool left behind.
Layer one is cheap. 99 skills are registered and a full scan is possible, so “is it installed” takes seconds. Which means layer one fills itself in automatically — and a column that fills itself in becomes the table’s default answer.
Layer two takes one press. The only value this piece needs from it is items 0 · files 0 · observer hooks 0; why that happened, and what state the rest of that table is in, belongs to the baselines piece. Layer two ends here.
Layer three is different. It is not zero because the tool is a shell; it is zero because a perfectly working tool is never pressed. The context-budget measurement skill had zero runs across 24 days, from install on 07-29 to the check on 08-22, with zero output. Nothing was broken — there had simply never been an occasion to call it. Which is why layer-one and layer-two checks can never catch this: it is on the list, and it runs when pressed. Seeing layer three requires first deciding “what should exist if this tool had run” and then going to look for it, and that question does not occur to you on its own.
The remedies differ too. An empty layer one means install; an empty layer two means fix; an empty layer three means turn it on. Only the last one leaves the tool untouched, so the more comfortable you are with fixing tools, the more wasted effort you spend at layer three — taking apart something that works.
The lesson that cost something here is how many columns the comparison table needs. One column crushes three. Writing “have it” requires three columns, not one, and with a single column what disappears is not the difference between the layers but their existence.
A run history of zero comes from the shape of the tool, not from laziness
Layer-three failure does not happen because a tool breaks. It is a working tool that has never once been invoked, which is why there is no error and no warning. Three shapes actually produce that zero, and none of them is fixable by willpower.
① Invocation-shaped. It only runs when a person runs it. Asking why three inspection skills already in place failed to catch a defect gave exactly that answer — all three are invocation-shaped, and from the moment the defect appeared until that scan, nobody ran them. So the net gain was not a new tool but “when does it run”.
② No entry point. The tool exists but there is no way to call it. The slash command meant to back the context-budget measurement skill was never created — checking all 37 commands found none. The skill document is intact and correct; there is simply no handle to pull it out with during a session. What makes this shape especially bad is that the person who builds the tool and the person who would build the entry point are the same, but the order splits them. While building the tool it is “I’ll call it later”; when the occasion to call it arrives, the tool has been forgotten.
③ No consequence. Nothing happens when it does not run. A distillation pipeline I built myself was like this — the front end diligently piled up months of material and the back end never ran once, and nobody noticed for months. If the back end does not run, no output file appears, and sessions run fine regardless. Here, “no error” is the bad news. And this shape also corrupts the diagnosis. That piece’s initial diagnosis was “observations are not accumulating” when 20.6MB of observations plainly existed; the real cause was the back end not running. Fixing it as diagnosed would have meant time spent on hooks while the distiller stayed dead — a failure with no consequence has no symptom, and therefore no clue pointing at its cause either.
The same piece carries the counterexample right next to it. A plain, hand-written memory file is injected automatically every session and is still running today. The dead one was far more sophisticated — fully automated by hooks, carefully designed. The difference between the two channels was not quality but whether there was an injection point. The simple one is force-read at the start of every session; the sophisticated one only ran when a person called it.
Line the three up and the common thread shows. None of them is a matter of will. Invocation-shaped depends on human memory, no-entry-point has no handle, and no-consequence has no penalty for skipping. Not one of them is fixed by “next time I’ll remember”.
The lesson that cost something here is the shape of the tracking list. A tool with layer three at zero never appears on a “broken” list. As long as management runs off a broken list, this layer stays invisible forever. And the way to fill layer three is not to fix the tool but to attach a trigger — quality and run history turned out to be uncorrelated.
Three investigations named it, and none of the three ran it
A layer-three gap is not accidental laziness but a recurring structure. Each later piece explicitly quoted the earlier one and still opened the next investigation without running the same two tools.
Round one. A piece taking apart an upstream repository of 285 skills closed with this line — “I installed two tools for trimming context six months ago and have never run them once. Running them is the top priority this investigation leaves behind”. That investigation adopted 1 of 285.
Round two. A piece taking apart a coding-agent harness from a model builder found that three of its five harvest candidates collapsed against exactly those two skills. Final harvest: zero. That piece quoted the earlier one and wrote, in its own words, “the same two tools. I wrote them down as top priority, never ran them, and went off to take apart another repo to fill the same gap.”
Round three. A piece taking apart seven prompt-assembly presets found one of them turning up again as “installed 24 days, run 0 times”. Final harvest 4 → 0. That piece also wrote “this was the third round.”
Round two settled what state the two tools were in — one had a single adoption memo as its only trace, and the other had produced none of what it promised to produce (how that was settled is in the “counted by missing output” section below). Since those two were what killed the three candidates, round two’s harvest of zero was not produced by a weak subject — it was produced by an empty drawer on my side. No amount of picking better subjects moves that number.
The three subjects are unrelated to one another. A skill collection, a coding-agent harness, a prompt-assembly preset — different makers, different jobs. What they share is not the subject but my drawer. Which means watching a single category would never catch this repetition: the three pieces sit in different buckets of the ledger, and the only thing joining them is the fact that they named the same two tools.
Why it repeated is also clear. The record existed; what was missing was an occasion to read it. Round one’s closing line was published, is searchable, and round two actually quoted it. But quoting and executing are different acts — round two used that sentence as evidence of a gap on my side, not as “then let us run them now”. The force that opens an investigation and the force that opens a drawer are different, and only the first one fired every time.
The lesson that cost something here is that repetition is itself a signal. When the same tool shows up as grounds for “we already have it” in two or more separate investigations, that is not an asset signal but a liability signal. An asset would have output piling up, and all three times it was zero. And writing something down triggers nothing on its own — the second piece read the first, quoted it, and still did not run that top priority. A record that gets read and a record that fires are different objects, and all three of these were the former.
Judge on layer one and you buy from outside what is already in the storeroom
Read inventory as layer one only and you get it wrong in both directions. You write down something you do not have as present and reject an adoption; and you go looking outside for something you already have.
Three lines checked before starting catch most of it. ① Is it installed ② Is there output ③ Is there an entry point to call it. ① alone makes you record something that never runs as “have it”, ② catches layer three, and ③ tells you why it is zero.
Three things actually went wrong for want of those three lines, and in all three the answer was already sitting somewhere in my own documents. In the coding-agent harness, five harvest candidates became zero: three were already in the drawer — but not running — and the other two folded into a two-line edit rather than an imported technique. In an agent I read down to the source, harvesting four techniques, the best of them was a rediscovery — a report I had written five days earlier on the same subject already had it, complete with where to apply it and a reopen condition. I keep every link collected and searchable and still did not open it. And an investigation that ended with zero techniques had its rejection reason already written in my own project documentation — the thing that subject sells simply does not apply to how my screens are drawn.
Recommendations from outside hit the same problem. A 44-second recommendation reel contained four tools, and two of them were ones I had already taken apart and rejected.
The three failures share something. None of them is catchable by reading more of the subject’s material. Each requires opening my own config file, or a post I wrote five days ago, or my own project documentation. But what you naturally open at the start of an investigation is the subject — it is new and interesting, and you believe you already know your own drawer. “I already know” is exactly the layer-one illusion.
The lesson that cost something here is that findability does not produce reading. A collected, searchable library of every link exists, and all three times it went unopened. Making something findable and actually going to find it are different jobs, and doing the first one well does not bring the second along. So these three lines have to sit as the first three lines of the kickoff document, not as “something to remember” — decide where they get written down, or this section is just one more record nobody opens.
Layer three is counted by missing output, not by logs
Almost no tool records when it last ran. Instead, look for the files, folders and entries that tool was supposed to create. This is the cheapest method and usually decisive.
Two distinctions first. One: this is a post-hoc audit, not a pre-flight gate. The gate that catches an empty success finishing with exit code 0 inspects output right after a run to decide whether to pass that run; this section infers after the fact whether a run happened at all. The same observation — is the file there — is read for different things. Two: this is not a new axis but grade-two evidence applied to run-history auditing — sizes, counts and hashes confirmed by deterministic code. Which also answers “what grade is layer one, then”. Much lower.
Here is how it actually caught things. One cleanup skill declared in its own documentation that it creates a particular file and a particular folder, and had created neither — zero runs confirmed without reading a single log line. The output a tool promises itself becomes the ruler that audits it. The observation pipeline had 127 files and 20.6MB but zero instincts produced, and this method even separated which end of the channel was dead: the front was alive, the back was not. The last cell of the four-stage ladder gave items 0 · files 0 · hooks 0. The measurement skill had zero output, and the absence of any command to call it was confirmed against all 37 commands.
This beats logs for three reasons. It is cheap — checking whether a file exists takes seconds and involves no judgement. False passes are rare — if the output a tool was supposed to produce is absent, there is no way to claim it ran. It localises the break — knowing the pipeline’s front held 20.6MB while its output was zero told me exactly which joint had failed. Hunting through logs would only have surfaced logs confirming the wrong diagnosis.
The limit is equally clear. It does not work on a tool that produces no output at all. For those, this method cannot separate “never ran” from “ran and left nothing”.
The lesson that cost something here is where the design responsibility sits. A tool that leaves no output cannot prove it ran. That is not a difficulty of auditing but a design defect in the tool — the tool is what made layer three uncountable. So anything new gets “does it leave a trace when it runs” as a requirement. A one-line run log is enough, and without it nobody can measure that tool’s value six months later.
Copy in someone else’s preset and layer three is empty from the start
Anything that arrives by copying is not “something I put in to use” but “something that came with the bundle”, so its run history is zero from day one. It sits on layer one perfectly well, becomes grounds for the next verdict, and bills its cost on every single turn.
My own setup was built that way. 55 of 98 skills, 14 rules and 10 of 12 hooks came from someone else’s preset bundle, and only two hooks are my own. More than half is someone else’s, and nowhere is it recorded when that half last ran. Run history does not come along with a copy, which means layer three for this bundle is not an empty column — there is no column at all.
What state the arrivals were in became clear only by holding someone else’s ruler against them, twice — once with a course’s six-item self-assessment, once with another repo’s doctrine document. Both pointed the same way. A skill turned up that could never be auto-invoked at all, and exactly one of 96 had an invocation axis. The total always injected into a session’s first request is 69,871 characters — global rules 24,407 · automatic memory index 19,124 · working instructions 2,669 · skill descriptions 23,671.
The direction of the cost matters. Layer one grows for free and the bill arrives every turn. Copying in a skill takes seconds; its description then rides on the first request of every session thereafter. So copy-adoption looks like a gain at the start and turns into a loss over time — and because that loss never surfaces at any single moment, nobody reverses it.
The 16 over 500 lines I decided to merely know about. They came from upstream, so splitting them would collide at the next sync. What can be fixed and what has to be known are different things, and dressing the second up as the first only lengthens the list. The reason skill denominators differ between posts is worth recording too — folders 102 · documents 96 · registered 98–99, because the measurements were taken at different times. That the denominator wobbles is itself evidence that even layer one does not count cleanly in one pass.
The lesson that cost something here is the difference between a ledger and an asset. Copy-adoption fills layer one very fast and fills layer three not at all. So the larger the drawer, the more grounds there are for writing “we already have it” and the less actual value there is — the assets did not grow, only the ledger did. What you actually buy when you take on someone else’s preset wholesale is not tools but ledger entries, and those entries come back as “already have it” at the next verdict.
Turn an unused tool on before fixing it — there are three dispositions and deleting is not the default
When you find a defect in a tool whose layer three is zero, the real damage from that defect is zero. Damage is defect size × number of runs, and the second factor is zero. So the order is not fixing but turning on.
These are the ones actually handled that way. ① A measurement skill’s formula overcounted overhead by 37.6×. A ruler whose markings are off by a factor of 37. And it misled exactly zero decisions, because it had never run — so the correction was downgraded from “repair” to “absorb”. How wrong the markings are matters less than how many times you laid the ruler down. ② Two tools I was going to attach a control arm to had zero actual runs — instrumenting a tool that has never run is a liability, so it went on hold with the reopen condition “at the same time as the first run”. That is where a harvest of three became one. ③ Two dead hooks were switched off and 20.6MB plus 127 empty shells were deleted. Even deleting, I recorded what was removed along with a secret scan showing zero — though that zero was luck, not design.
The same judgement had already been made once. A sorting tool was rejected on the grounds that “there are already two tools doing this job and neither has ever been run. A third is not a solution, it is a deferral”. Bringing in a fourth while layer three sits at zero is avoidance, not adoption. And yet, having written that down, two more investigations into the same gap were opened — reaching a judgement and living by it are different jobs.
Disposition ③ carries one condition. When you delete, record what you deleted and why. Deleting erases layer one as well as layer three, so without a record the slot looks six months later like something that was never there. Then the same thing gets bought from outside again, which is exactly the path this piece set out to block.
The lesson that cost something here is two distinctions: order, and verdict. Not “no need to fix it” but “turn it on before fixing it” — improving a tool that never runs adds maintenance cost against zero output, and before turning it on you cannot even tell what needs fixing. And rejecting is not deleting. Deleting erases layer one as well, so without a reopen condition you buy the same thing from outside next time.
What I could not do
- There is still no mechanism that records “the date it last ran”. I used presence of output as a proxy, and that answers “did it ever run”, not a date. A tool that ran once six months ago and never again passes this method. And a tool that leaves no output cannot be counted at all — I have not counted how many of those are in my drawer.
- I did not audit layer three across all 99 skills. Only four, caught in three pieces, plus one of the four “equivalent” cells on the four-stage ladder; the rest are still paper. This piece talks about how to count and did not do a full count — the sample is “whatever happened to get caught”, so I cannot even give a rate.
- I could not measure the gain from filling layer three. With no hook to remove the 69,871 characters of always-injected context, building a control is impossible — that is not measurable, not “no problem”. There is no sentence in this piece saying “fill layer three and you gain N%”, and there should not be. Until a control exists, the grounds for this piece’s advice end at “I observed the repetition three times”.
- Every measurement comes from one personal setup. One Windows machine’s agent configuration. Whether the same three layers hold at team or organisation scale is untested — a shared tool might satisfy “somebody ran it” and hide the gap far longer, or might die faster because nobody treats it as theirs. I do not know which.
- I do not track the after-the-fact loss from disposition ③. I do not know whether anything I deleted turned out to be needed later. I recommend deleting without being able to measure its cost, and deletion is the hardest of these recommendations to reverse.
- The audit this piece recommends answers only “how many are not running”. “And what was that costing me” has no answer without a control. I found four tools that never ran and cannot say what would have been different had they run — and I publish in that state.