精選文章

MyR2D2 Update: I Switched to Opus 5.5 and Rewrote My Subagent Dispatch Rules

When I posted about the Opus 5.5 release, I said those of us who keep hitting Fable's weekly limit finally had somewhere to go. A week or so after switching, I also updated the dispatch rules in the public version of MyR2D2. If you're already using it, remember to update to the latest version.

As for Opus 5.5 itself, Anthropic's model guide says it plainly: start with it for most workloads, and switch to Fable 5.1 for deep reasoning or long-running agentic work. Pricing is $4 per million input tokens and $20 per million output tokens; the previous Opus 5 was $5 and $25. Default effort is medium.

But here's the new pitfall I ran into this month: dispatch the wrong way, and no model is cheap enough to save you.

A while back, I let four multi-agent workflows run on Opus end to end, with no model tiering. In one day they burned about 4.6 million subagent tokens, and every conversation stopped at once. After the tokens ran out, I burned through another $50 in credits in a single night. After that I rewrote my dispatch rules and shipped them in token-optimizer with v0.7.5. Four hard rules:

  1. Specify the model on every single dispatch; don't rely on a global switch.
    I originally thought the subagent-model environment variable in the settings file was a hard cap: set it to the mid tier and you're safe. I even wrote that into the skill. Two probes later, I found it's only the fallback for calls that don't specify a model. A model specified on the call really takes effect: specify the flagship, and you really pay flagship rates. So it only saves you when you forget to specify one, not when you specify the wrong one. Now every agent call specifies its own model, and the environment variable is just the fallback.

  2. A subagent is not a free clone.
    Every subagent reloads the rules, memory, and skill list at startup. In my setup, that fixed overhead comes to about 50,000 to 60,000 tokens, and subagents don't share the main conversation's context. Dispatch five, and the same batch of files gets read six times, and you pay six times.

  3. No dispatching for everyday work.
    I now dispatch in only three cases: the material doesn't fit in one conversation; the job needs an "attacker" view, like permission gates or credentials; or I've explicitly said to go all out. "This is a bit complicated" doesn't count. A counterexample: for one 621-line script, I dispatched five agents and spent 640,000 tokens, and two of them crashed outright with zero output. A same-size task I did myself that day cost about a tenth as much.

  4. Do the math before dispatching: if dispatching is estimated to cost at most 20% more than doing it yourself, you can dispatch.
    Convert both routes to API cost using the same assumptions. Doing it yourself = estimated number of calls × new tokens per call. Dispatching = each agent's fixed startup cost + number of calls × the subagent's tokens per call. If the estimated dispatch cost is within 1.2× of doing it yourself, you can dispatch; above that, do it yourself. Why is 20% more still OK? First, the estimate itself is off by about 20% either way, so within 20% is a tie. Second, on a tie, dispatching brings two more benefits: it keeps the main conversation's context from growing, and you save flagship-tier quota (Fable 5.1's input and output prices are 2.5× Opus 5.5's, and once the quota is gone, every conversation stops). This morning I estimated one verification agent at 200,000 tokens, and it actually ran 189,515. The estimate was close because I counted the fixed overhead. v0.7.5 also builds this step into the skill: before dispatching, state up front how many agents, which tier, and roughly how many tokens; at wrap-up, write back how far off the estimate was from what actually ran.

Three findings from testing are in this release too. The Agent tool has no effort parameter, so subagents run at the same effort as the main conversation, while Workflow dispatches need effort written explicitly, with medium as the floor. Don't use fork as an executor: it carries the whole context along with the parent model. And to see which version an alias like sonnet or opus actually ran, check the model field in the subagent transcript, not what you wrote.

I now log my token burn every day. As it turns out, you really do have to look at the numbers. Comparing the two weeks before and after the rule change, new tokens per day fell from about 57.4 million on average to about 25.5 million. Subagents dropped the most: in the first two weeks they ate about 300 million tokens, nearly 40% of the total; in the last two weeks, about 28.3 million, under 10%. The main conversation itself fell by only a little over 30%.

Daily tokens before and after the dispatch rule change

I scratch my own itch.

MyR2D2 v0.7.5 release page

How to install and upgrade:

New install (installs at the project level by default; add -g for a global install):

npx skills add tingyulu/MyR2D2

Already on an older version (one line updates every installed skill):

npx skills update

Installed as a plugin: update the marketplace from the /plugin menu.

To get notified of new versions, Watch the repo on GitHub and subscribe to Releases.

Repo: https://github.com/tingyulu/MyR2D2
This release: https://github.com/tingyulu/MyR2D2/releases/tag/v0.7.5
What MyR2D2 is, from the start: https://www.uncleric.com/2026/09/myr2d2-open-source-claude-skills.html

When you dispatch subagents, what pitfalls have you hit, and how many tokens did they burn?

I'm Eric Lu ("Uncle Eric"), a product consultant, headhunter, and career coach. These skills are my actual daily workflow.

留言