Teardownspublic

Buzz group chat

Slack where the channel members are people and local CLIs. Two demos, one worked — and the presenter admitting which one didn't is what this video is worth.

Two demos with identical structure, one worked, and the presenter says so out loud — that single sentence is the most valuable thing in these fifteen minutes. Buzz is a Slack-shaped desktop app where the channel members are not only people but the coding CLIs installed on my machine. Claude Code, Codex and Kimi are each registered as a separate “staff member”; call @lead and the lead splits the work across the rest. The stock analysis came out properly after 34 turns. The web page build ran 24 turns and landed at single-model quality. The difference was not the tool — it was whether there was anything real to divide.

one agent = one identity + one harness, bound together. The two are completely separate axis A — identity · who it is people and agents share one address space · keypair + handle · pinned to the device so you can @ them, they sit beside people in the member list, and they have profile cards axis B — harness · what it runs on no model of its own. It drives CLIs already installed and already signed in on my machine Claude Code · Codex · Kimi Code · Goose … billed to my subscription, not metered API ⇒ this is why every agent gets its own private key — secrets grow to 1 + N
Think of these two as one thing and the per-agent keys make no sense. Split them and everything explains itself.

What it actually is

A Slack UI with no LLM of its own. You make a channel, invite people and agents, and call them with @. An invited agent first appears as waking — its harness process is starting.

The real difference from existing orchestration is the direction of the relationship. Subagents, or my own tc-team, have a parent calling children inside one process. Buzz inverts that — an agent has an address, not a process. They are peers rather than parent and child, so agents naming and questioning each other is natural, and the conversation log is the work record. Who asked whom for what, and how many times it was sent back, is simply the reply count on a thread.

And models mix inside one channel. The Warren Buffett seat runs on Kimi, Peter Lynch on Claude, Benjamin Graham on GPT. As we’ll see, that is the most important design decision in the video.

There is an economic axis too. It bills against a subscription you already pay for rather than metering an API key. Running five agents costs subscription usage, not extra API spend.

The technique — the lead’s instruction is one sentence

The orchestration logic isn’t in code. This is the lead agent’s entire system prompt:

You will not do the work yourself. You will use the specialists in this channel and manage the task to completion.

No routing logic, no tool definitions. The lead queries the channel membership at runtime with a CLI command and divides accordingly — the observation panel literally shows that command running. Team composition lives in channel membership, not in code.

The observation window being open is good too. Clicking an agent shows thinking blocks, the commands it ran, context burn as a real figure (26,537 / 258,400), the number of commands available, and inline permission requests.

How they unstuck the install was the fun part. When an installed CLI didn’t appear in the list, they pasted the harness configuration block into that CLI itself and said “this isn’t connecting — fix the path if you can and apply it.” Having an agent repair its own wiring generalises nicely as an environment-troubleshooting move.

What the demos showed

two demos — same structure, opposite outcomes stock analysis · 34 turns — it worked the three specialists ran on three models each viewpoint looks at genuinely different data = there was something real to divide web page build · 24 turns — it didn't designer · front end · back end, three roles the request was shallow, so there was nothing to split = only coordination cost went up ⇒ the presenter's own admission — "just using a single model might well be better" Multi-agent pays off only where viewpoints genuinely diverge. Splitting role names is not division of labour And with one model wearing personas, agreement isn't verification — it's an echo
Turn count is not a quality metric. 34 beat 24 because there was something to divide, not because there were more turns.

The best moment in the stock case is the lead waiting. One specialist reported first, and rather than finalising, the lead explicitly declared the result incomplete and held. Then it sent the gap back to the same specialist. A completeness gate that emerged from one line of prose. It was induced by writing, not enforced by code, so reproducibility isn’t guaranteed — but for something obtained at zero cost, it’s worth a lot.

And “the three specialists did not contradict each other” only means something under one condition. Because the three ran on different models, their agreement carries information. Three instances of one model wearing different personas would agree almost automatically, and that is not verification — it is an echo.

The web case was the inverse. The instruction was one line — “a modern web page referencing Apple and Tesla design” — and with no detail, the designer settled the concept itself. The output was fine, but there was little for three roles to divide. The presenter says it directly: “when you actually build a web page, just using a single model might well be better.”

Held against my own setup — and I didn’t install it

The only build confirmed is macOS on Apple Silicon. Windows support is outside the video’s scope and unverified. That one line made measurement impossible.

Even if I could install it, three things needed answers first.

  • Execution permission scope. The observation panel showed a permission-bypass mode, and the approval options include “allow all commands starting with this prefix.” Prefix approval widens the blast radius.
  • A third-party relay. Communities are relayed by an external service. Work content cannot go in until I know how far a channel conversation travels.
  • Secret count. One identity key plus one per agent, and each is revealed exactly once at creation. Five agents means six secrets. That is an operational burden to be counted, not a convenience feature.

Verdict — skipped the tool, took three principles

What Call
Heterogeneous model panel — split models, not personas, for verification passes Adopt. Applies directly to adversarial review design
A conductor that declares incompleteness and waits — a gate made from one line of prose Adopt. Zero cost on ad-hoc work with no code gate
Have an agent repair its own wiring Adopt. Generalises as environment troubleshooting
Adopt the tool On hold. Platform unsupported, permissions and relay unresolved

The first row is the big one. I have been running verification passes by giving one model several roles, and this video showed me why that is weak. For agreement to carry information, the ways of being wrong must differ. Swapping personas leaves the way of being wrong identical.

In fairness, the video admitting its own limit is what raised my confidence in it. Had it shown only the demo that worked, I would have read “34 turns, must be good.” Turn count is not a quality metric, and nothing in the video argues that more turns are better. Cost savings and accuracy gains over a single agent are not measured anywhere — any quotation stops there.