Three ways to change what an agent does, at three different layers
"Just have the agent do X differently" is not one decision, it's a choice between three genuinely different mechanisms that happen to produce similar-looking results from the outside: write a skill (change what's in context for this request), add a tool (give the model something new it can actually call), or fine-tune (change the model's weights). They fail differently, cost differently, and iterate at different speeds, and picking the wrong one for a given problem is a common, avoidable source of wasted effort.
This follows directly from the boundary discussion in the skills architecture piece: that article establishes that a skill is not a tool and not a fine-tune; this one is the decision framework for choosing correctly among the three in the first place.
What each mechanism actually changes
A skill changes what text is in the model's context window for a given request. It's instructions -- methodology, style, a checklist, a format to follow -- and the model still has to reason its way through applying them fresh, every time, using whatever general capability it already has.
A tool gives the model a function it can call: a schema describing inputs and outputs, and real code behind it that executes, reads state, or reaches an external system. A tool doesn't tell the model how to think; it gives the model something to do that the model itself cannot do by generating text -- query a database, run a test suite, fetch a live price.
Fine-tuning changes the model's weights via a training run, so the behavior is baked in before any request-time context exists at all. Nothing needs to be in the prompt for a fine-tuned behavior to show up; that's both the appeal and the cost.
Comparison
| Skill | Tool | Fine-tune | |
|---|---|---|---|
| Iteration speed | Edit a file, next request picks it up | Ship code + schema, redeploy | Training run + eval, hours to days |
| Cost to change | Near zero | Engineering time | Compute + data curation |
| Reviewable as a diff | Yes, plain text | Yes, code review | Only indirectly, via eval deltas |
| Can access external state | No -- only via tools it invokes | Yes, that's its purpose | No -- weights have no live state |
| Typical failure mode | Model doesn't follow instruction under load / distraction | Schema mismatch, execution error, timeout | Silent regression on cases outside training distribution |
| Where it wins | Behavior, methodology, output format, judgment calls | Anything requiring real execution or fresh external data | Narrow output format at very high volume / low latency |
The failure-mode row is the one worth sitting with. A skill fails softly -- the model drifts from the instruction, and the fix is a clearer instruction, reviewable the same day. A tool fails loudly and locally -- a bad schema or a timeout is a normal engineering bug with a normal engineering fix. A fine-tune fails quietly and broadly -- a regression shows up as a shift in behavior on inputs nobody thought to eval, discovered well after the training run that caused it, which is exactly why it should be the last resort, not the first idea.
Worked example: teaching an agent to write better commit messages
This is a skill, not a tool or a fine-tune. There's no external state to fetch and no execution to perform -- the entire job is judgment about phrasing, tense, and what belongs in a commit body, applied to a diff the model can already see. A tool would be the wrong layer because there's nothing to call; a fine-tune would be enormous overkill for a behavior that a well-written instruction handles on the very next request, reviewable as a plain-text diff before it ships. This site's own commit-message-writer skill is exactly this case.