Matching is a retrieval problem, not a routing table

Once a runtime has more than a handful of skills installed, the question stops being "which skill should run" and becomes "how does the runtime decide, cheaply, for every single request, without a human writing a rule for each new skill." That's a retrieval problem: given a request and a catalog of short description strings, find the closest match. Get the retrieval step wrong and the rest of the skills architecture -- however cleanly decoupled -- never actually fires correctly.

This is the mechanism half of the skills architecture piece: not why skills are decoupled from the runtime, but the concrete matcher that sits between a request and a skill's body being loaded into context.

Advertisement

Three ways to implement the match

Keyword and heuristic matching. The cheapest option: tokenize the request, tokenize each skill's name and description, score overlap. It's fast and needs no model call, but it's brittle -- a request phrased as "can you look over this diff before I merge" won't token-match a description that only says "review a PR," even though a human reads the two as identical intent.

Embedding similarity. Encode the request and every skill description into the same vector space, rank by cosine distance. This closes most of the keyword-matching gap (paraphrases land close together) at the cost of maintaining an embedding index and a similarity threshold that needs tuning per catalog size. It scales well to large catalogs because the expensive part -- encoding the descriptions -- happens once, offline, not per request.

LLM-as-router. Hand the model the request plus the full list of name/description pairs and ask it to pick the best match, or none. This is the most accurate of the three because the router shares the same language understanding as the model that will eventually execute the skill, and it's what Claude Code and comparable tools actually do at moderate catalog sizes. The cost is a model call before the model call, which is real but usually small relative to running the matched skill itself, since the router only ever sees metadata, never full skill bodies.

In practice the three aren't exclusive: a cheap keyword or embedding pass narrows a large catalog to a handful of plausible candidates, and an LLM call makes the final call among those few. That two-stage design is the same coarse-then-precise pattern that shows up in agent request routing generally.

Advertisement

Why the description matters more than the algorithm

No matching algorithm rescues a bad description. A description that says only "helps with code" will collide with half the catalog under any of the three approaches above, because there's no signal in it to discriminate on. The two properties that actually move match quality are independent of which matcher you pick:

Specific triggers. "Review a diff, PR, or set of changed files for correctness bugs, security issues, and unnecessary complexity" gives every matcher concrete nouns to latch onto -- diff, PR, changed files, bugs -- that a vague "helps with code" doesn't.

Explicit negative scope. The single highest-leverage addition to a description is a clause stating what the skill is not for. This site's own code-review skill description ends with "not for general code questions" specifically because without that clause, any message that merely mentions code risked a false match. A negative clause does something a matching algorithm structurally cannot do on its own: it tells the retrieval step to actively rule a candidate out, not just rank it lower.

# weak -- true of nearly every request in a coding tool
description: Helps with code.

# strong -- specific triggers, explicit negative scope
description: Review a diff, PR, or set of changed files for correctness
  bugs, security issues, and unnecessary complexity. Use when asked to
  review code or check a PR before it merges -- not for general code
  questions or explaining how existing code works.