Three ways to change what an agent does, at three different layers

"Just have the agent do X differently" is not one decision, it's a choice between three genuinely different mechanisms that happen to produce similar-looking results from the outside: write a skill (change what's in context for this request), add a tool (give the model something new it can actually call), or fine-tune (change the model's weights). They fail differently, cost differently, and iterate at different speeds, and picking the wrong one for a given problem is a common, avoidable source of wasted effort.

This follows directly from the boundary discussion in the skills architecture piece: that article establishes that a skill is not a tool and not a fine-tune; this one is the decision framework for choosing correctly among the three in the first place.

Advertisement

What each mechanism actually changes

A skill changes what text is in the model's context window for a given request. It's instructions -- methodology, style, a checklist, a format to follow -- and the model still has to reason its way through applying them fresh, every time, using whatever general capability it already has.

A tool gives the model a function it can call: a schema describing inputs and outputs, and real code behind it that executes, reads state, or reaches an external system. A tool doesn't tell the model how to think; it gives the model something to do that the model itself cannot do by generating text -- query a database, run a test suite, fetch a live price.

Fine-tuning changes the model's weights via a training run, so the behavior is baked in before any request-time context exists at all. Nothing needs to be in the prompt for a fine-tuned behavior to show up; that's both the appeal and the cost.

Advertisement

Comparison

SkillToolFine-tune
Iteration speedEdit a file, next request picks it upShip code + schema, redeployTraining run + eval, hours to days
Cost to changeNear zeroEngineering timeCompute + data curation
Reviewable as a diffYes, plain textYes, code reviewOnly indirectly, via eval deltas
Can access external stateNo -- only via tools it invokesYes, that's its purposeNo -- weights have no live state
Typical failure modeModel doesn't follow instruction under load / distractionSchema mismatch, execution error, timeoutSilent regression on cases outside training distribution
Where it winsBehavior, methodology, output format, judgment callsAnything requiring real execution or fresh external dataNarrow output format at very high volume / low latency

The failure-mode row is the one worth sitting with. A skill fails softly -- the model drifts from the instruction, and the fix is a clearer instruction, reviewable the same day. A tool fails loudly and locally -- a bad schema or a timeout is a normal engineering bug with a normal engineering fix. A fine-tune fails quietly and broadly -- a regression shows up as a shift in behavior on inputs nobody thought to eval, discovered well after the training run that caused it, which is exactly why it should be the last resort, not the first idea.

Worked example: teaching an agent to write better commit messages

This is a skill, not a tool or a fine-tune. There's no external state to fetch and no execution to perform -- the entire job is judgment about phrasing, tense, and what belongs in a commit body, applied to a diff the model can already see. A tool would be the wrong layer because there's nothing to call; a fine-tune would be enormous overkill for a behavior that a well-written instruction handles on the very next request, reviewable as a plain-text diff before it ships. This site's own commit-message-writer skill is exactly this case.