Teardownspublic

mattpocock/skills

I didn't take the 35 skills. I took the one document that defines how to use them — and measured against it, 1 of my 96 skills passed.

I didn’t install it — because of the repository’s own argument. Installing the plugin puts 25 skill descriptions resident in context every turn, which is precisely what this repository tells you not to do. Instead I used its doctrine document as a ruler and measured my own 96 skills. That measurement is this page: one had the invocation axis set, and the other 95 descriptions — 20,672 characters — load on every single turn.

35 skills — foldered by maturity engineering 18 promoted · ships productivity 7 promoted · ships misc 4 kept, not advertised in-progress 6 beta · public on purpose deprecated 0 the retirement shelf only the promoted 25 ship — and a doc enforces that invariant two distribution paths — different philosophies subscription — read-only bundle, auto-updates when the author ships a subscription, not a fork. You can't edit it ownership — editable files copied into your project you own it and change it. Updating is manual ⚠ Install both and you get every skill twice — a warning repeated in three places, itself copied from one source file
Matt Pocock publishing his own agent directory wholesale. It also ships as a plugin on the official marketplace.

What the repository sells

A bundle of slash commands. But the selling point is that it is not a framework that owns your process. The README names its competitors directly.

Those approaches help by owning the process for you. In doing so they take away control and make bugs in the process itself hard to fix. These skills are designed to be small, easy to fix, and composable.

And it doesn’t enumerate skills. It organises itself around four ways agents commonly fail, with a prescription attached to each.

Failure mode Diagnosis Prescription
It builds the wrong thing “Nobody knows exactly what they want” — the communication gap between a person and an agent An interrogation session
It’s far too verbose With no project jargon, it writes twenty words where one would do A shared vocabulary doc + domain modelling
The code doesn’t run No feedback loop. Without types, a browser or automated tests, it’s flying blind TDD · bug diagnosis
It built a ball of mud An agent accelerating coding accelerates software entropy with it Codebase design

The example under the second one is blunt. BEFORE: “the problem that happens when a lesson inside a section inside a course gets ‘materialised’ — given a place on the filesystem” → AFTER: “there’s a problem in the materialisation cascade.” And it justifies this on token logic rather than taste: a shared language means fewer tokens spent on thinking.

It has 207,586 stars and three contributors. Two hundred thousand at six months old is the author’s newsletter audience, not code scrutiny. That said 17,928 forks puts it at 11.6 to 1 — far healthier than Odysseus at 450 to 1, meaning people genuinely use it.

What actually came across — the two loads

The value isn’t the 35 skills. It’s the one doctrine document that defines how to use them. Its opening sentence states the target precisely.

The same levers make each of them predictable — getting the agent to walk the same process on every run, not to produce the same output.

Its central claim is that cost comes in two kinds.

Load Who bears it Treatment
Context load The machine — every resident line. It spends tokens and attention every turn whether it fires or not Minimise
Cognitive load The person — knowing what exists and when to reach for it. The human acts as the index Do not minimise

The second is the counterintuitive one. A person’s cognitive load isn’t a cost to be eliminated, it’s the price of their agency. Keep it where judgement matters; remove it only where it doesn’t.

Three working rules follow.

Progressive disclosure is a variance lever, not a token optimisation. Leave reference material inline that should have been disclosed and it buries the steps, turning whether attention lands on a given step into a coin flip. That’s a variance problem, not a readability one.

Steering by prohibition pulls the thing in. The forbidden behaviour enters the context and becomes more accessible — “don’t think of an elephant,” and now there’s nothing but elephant. Negation is a weak modifier, so a strongly activated concept tramples it and the prohibition reads half as an instruction. State the positive goal instead, and never name the forbidden side.

Hunt for no-ops. An instruction the model already follows by default costs load and says nothing. The test is “does behaviour change against the default” — and it’s a model question, not a reader question. So when two people argue about whether something is a no-op, they’re disagreeing about the default, and that gets settled by running the document, not by debating it.

A completion criterion has two properties

If an agent can’t tell “done” from “not done,” you get premature completion — it finishes before it’s finished, and attention slides toward “done.”

Property What
Clarity A fuzzy boundary (“reached an understanding”) invites premature completion. Sharpen the boundary first
Demand “Every modified model will be explained” forces thoroughness; “make a list of the changes” doesn’t

The strongest criteria are checkable and exhaustive at once. Demand rarely appears as its own step — it hides inside the phrasing.

Measured — 1 of 96

my 96 skills — one square each the one blue square = the only skill with an invocation axis The other 95 descriptions — 20,672 characters — load every turn On top of that sit tool schemas, nine rule drawers and a memory index
Having read the doctrine, I held the ruler up. As measured: 102 folders, 96 skill documents.

The repository’s only classification axis is user-invoked versus model-invoked. One question decides it — “is it useful for the model to reach for this on its own?”

Model-invoked (default) User-invoked
Description For the model — rich trigger phrasing For a person — one line read off a slash menu
Load Permanent context load ↔ discoverability Zero context load ↔ cognitive load
Reach Model or person — model-invoked includes human access People only

The sharpest example turned up in my own set. One skill’s description says, of itself: “explicit invocation only — do not stop the user with comprehension questions during ordinary work.”

The policy is declared in prose, and the flag the harness actually enforces isn’t set. Read on Matt’s axis, that’s a user-invoked skill shipped as model-invoked — left reachable by the model, with “don’t reach for this” written in the body. It is exactly the negation failure mode described above. One flag replaces that whole paragraph.

⚠ Applying it to all 96 would be wrong. Many of mine must stay model-invoked — the pipeline skills an orchestrator calls, for instance. The value of the axis isn’t conversion; it’s passing all 96 through that question once, and checking first whether an agent or hook calls the skill before setting the flag.

What I didn’t take, and why

Installing the plugin: rejected. All four reasons are measured.

  • It adds 25 descriptions to resident context. Stacking that onto an environment already loading 95 unbounded is a direct violation of the author’s own doctrine
  • Duplication — TDD, code review, handoff and research all already exist on my side in the same role
  • Mismatched premises — most of the engineering skills require wiring to a particular issue tracker, and my actual tools have no adapter
  • Language — the descriptions are in English, so model-invocation hit rate on Korean phrasing is lower than my existing skills’

Three disciplines came across instead. Keeping a folder that pins down “we don’t do this” with the reason (my own rejection list currently survives as abbreviations, with the why missing). Splitting review into two axes while forbidding merge and re-ranking (so a formatting note can’t mask a coverage gap). And asking every question you can ask right now in one round with a recommendation attached to each, rather than one at a time.

One line attached to that last one is good: “finding facts is my job and never the user’s. Don’t ask the user something I could look up. The decisions are theirs.”

There are things not to copy, too

This repository has no tests. Automated verification amounts to plugin-manifest validation; the quality of the skill bodies rests on the author checking by hand. There’s a doctrine that says settle it by running the document — and that run isn’t in CI. On this axis my side is ahead: two linters that compare documents against reality run mechanically.

The claim not to own your process has counterexamples. Several of the top engineering skills prescribe issue-tracker schemas, label vocabulary and document placement in some detail. Against the competitors the README attacks, that’s a difference of degree, not kind. What’s genuinely “small and composable” is the productivity bucket and the interrogation family.

And the examples are all web and TypeScript. What travels is the vocabulary and the discipline, not the examples.