A prompt in production is code that is interpreted by a model you do not control. Change one sentence and accuracy can move by several points; keep the sentence and let the provider update the model behind an alias, and it can move again. Prompt versioning is the discipline that lets you answer three questions at any time: exactly what produced this output, what changed between the version that worked and the one that does not, and how fast can we go back.
This article covers the version itself: what it must contain, how to identify it, how changes are reviewed and gated, and how versions are resolved, logged, rolled out and rolled back. The service that stores and serves versions is described in Prompt registry architecture, and designing the evaluations that gate changes is covered in Prompt Evaluation Architecture.
The prompt is the wrong unit
Teams usually start by versioning the template text, and then discover that outputs changed when nobody touched it. Model behaviour is a function of everything sent and everything that configures the call. A version must therefore capture the whole bundle: the system and user templates; the few-shot examples, which are often loaded from a file or a database and change independently; the tool definitions and output schema, because a renamed field or reworded tool description changes behaviour as much as the template does; the model identifier; and the sampling parameters such as temperature and the token limit.
Two further inputs are often relevant and frequently forgotten: retrieval configuration for prompts fed by search, such as the index, the number of chunks and the ranking model, and any preprocessing applied to user input before it is substituted. If an input can change the output and is not in the bundle, it is an unversioned dependency, and an incident investigation will eventually hit it.
Identity: content hashes underneath, labels on top
A version needs an identity that cannot lie. Hand-maintained numbers drift: someone edits a file and forgets to bump the number, and two different bundles share one label. The robust approach is content addressing. Serialize the bundle canonically, with sorted keys, fixed separators and referenced files inlined, then hash it. The hash is the version identifier; any change to any input yields a new one, and identical bundles get identical ids wherever they are built.
Humans still need names, so layer two kinds of label on top. A semantic version such as 1.5.0 communicates intent in review. Aliases such as prod and canary are mutable pointers to immutable versions, and releasing or rolling back means moving a pointer. Never edit a released version in place; if it is wrong, release a new one.
# prompts/ticket_classifier/1.5.0.yaml -- immutable once released
name: ticket_classifier
version: 1.5.0
model: vendor-model-2026-06-15 # a dated snapshot id, never a moving alias
params: {temperature: 0, max_tokens: 200}
template:
system: |
You classify customer support tickets into exactly one category.
Categories: billing, outage, account_access, feature_request, other.
user: |
Ticket:
{ticket_text}
few_shot: examples/ticket_classifier/v3.jsonl
output_schema: schemas/ticket_category.v2.jsonimport hashlib, json, pathlib, yaml
ROOT = pathlib.Path("prompts")
def load_bundle(name: str, version: str) -> dict:
spec = yaml.safe_load((ROOT / name / f"{version}.yaml").read_text())
# Inline every referenced file so the hash covers its contents, not its path.
spec["few_shot"] = (ROOT.parent / spec["few_shot"]).read_text()
spec["output_schema"] = json.loads((ROOT.parent / spec["output_schema"]).read_text())
return spec
def version_id(bundle: dict) -> str:
identity = {k: v for k, v in bundle.items() if k != "version"} # labels are not identity
canonical = json.dumps(identity, sort_keys=True, separators=(",", ":"), ensure_ascii=False)
return "pv_" + hashlib.sha256(canonical.encode("utf-8")).hexdigest()[:16]
def resolve(name: str, alias: str, aliases: dict) -> tuple[dict, str]:
version = aliases[name][alias] # e.g. {"ticket_classifier": {"prod": "1.4.2"}}
bundle = load_bundle(name, version)
return bundle, version_id(bundle)Note that the hash excludes the version label itself, so relabelling does not create a new identity, and that it inlines the few-shot file and schema so editing an example changes the hash even though the manifest did not change. The templates are kept as separate fields rather than concatenated, which keeps diffs readable; reusable template patterns are covered in Prompt templates.
Change classes and what each requires
Semantic versioning maps onto prompts imperfectly, because even a typo fix can shift behaviour. It is still worth defining change classes by their effect on consumers. A major change alters the output contract: a new category, a renamed field, a different format. Downstream parsers and stored data must change with it, so it ships with a schema version bump and a migration plan. A minor change alters behaviour within the contract: new instructions, new examples, a different model snapshot. A patch is intended to be behaviour-neutral: wording, formatting, comments.
The rule that matters is that every class runs the evaluation suite, because intent is not evidence. What differs is the bar and the rollout. Patches need no regression beyond the noise floor; minor changes need an improvement or a justification; major changes need consumer sign-off and a coordinated release, because a parser expecting the old schema will fail on the new output immediately. Structured output contracts are discussed in Structured output.
The git and CI workflow
Keep prompt bundles in the same repository as the code that calls them, in plain text files, so a pull request shows the diff of the template, the examples and the schema together with the code that consumes them. Reviewers should be able to read the change as prose. On every pull request that touches a bundle, CI computes the new version id, runs the evaluation set against both the production version and the candidate, and fails the build if quality drops beyond agreed limits.
# ci/check_prompt_change.py -- runs on every pull request that touches prompts/
import json, sys
base = json.load(open("eval/base_report.json")) # scores for the version in prod
cand = json.load(open("eval/candidate_report.json")) # scores for the changed bundle
assert cand["version_id"] != base["version_id"], "bundle unchanged but files differ?"
failures = []
for metric, floor in {"accuracy": -0.01, "schema_valid": 0.0}.items():
delta = cand["overall"][metric] - base["overall"][metric]
if delta < floor:
failures.append(f"{metric} moved {delta:+.3f} (allowed {floor:+.3f})")
for slice_name, s in cand["slices"].items(): # regressions hide in slices
if s["accuracy"] < base["slices"][slice_name]["accuracy"] - 0.03:
failures.append(f"slice {slice_name} regressed")
if failures:
sys.exit("prompt change blocked:\n" + "\n".join(failures))
print("prompt change ok:", cand["version_id"])Gate on slices, not only on the overall score. A change that improves average accuracy by two points while dropping a 5% slice, say non-English tickets, by ten points will pass an average-only gate and fail real users. Store each evaluation report keyed by version id, so the evidence for every release is retrievable later. When the suite is expensive, run a fast subset on every commit and the full set before merge.
Some teams keep prompts in a registry service with a web editor so non-engineers can edit them. That is a valid choice, but the same properties must hold: immutable versions, content identity, mandatory evaluation before an alias moves, and an audit log of who moved which alias. A registry without those is a shared text box.
Runtime: resolve, call, log the hash
At runtime the application asks for a prompt by name and alias, resolves it to a bundle and version id, and calls the model. Resolution can happen at deploy time, where the bundle is baked into the build, or at request time, where the service reads the alias from a registry with a short cache. Deploy-time resolution is simpler and reproducible; request-time resolution allows rollback without a deploy. Many teams resolve at request time with a cache of a minute or so, and fall back to the last known good bundle if the registry is unavailable.
Whatever the mechanism, log the version id with every model call, alongside the model identifier, the alias, latency, token counts and whether the output parsed. Without it, a complaint about a bad answer cannot be traced to the bundle that produced it, and A/B comparisons between versions are impossible. Store the rendered inputs for a sample of calls as well, subject to your data-retention rules, so failures can be replayed against a candidate version. The debugging loop that uses these traces is covered in Prompt debugging workflow.
Pinning the model
Providers typically offer two kinds of model identifier: a moving alias that is updated to newer models over time, and a dated or versioned snapshot that is not supposed to change. A bundle that names an alias is not a version, because its behaviour can change without any change on your side. Pin the snapshot in the bundle, and treat moving to a newer model as a minor change that goes through the full evaluation gate like any other.
Snapshots are retired on a published schedule, so model migration is routine work rather than an emergency. Track the deprecation dates of every snapshot your bundles pin, start the migration well before them, and expect to retune: prompts tuned for one model often need adjustment for its successor, particularly few-shot examples and formatting instructions. Running the evaluation suite against the new snapshot with the old prompt first tells you how much retuning is needed.
Rollout and rollback
Moving an alias directly from one version to another exposes every user at once. A canary alias receives a small, deterministic share of traffic, chosen by hashing a stable key such as the user id so a user does not flip between versions mid-conversation. Compare the canary's online metrics with production: parse failure rate, refusals, latency, cost per call, and any downstream signal such as escalation rate. Widen the share in steps, and promote by pointing prod at the canary's version.
import hashlib
def pick_alias(user_id: str, canary_percent: int) -> str:
# Deterministic: the same user always sees the same version during a rollout.
bucket = int(hashlib.sha256(user_id.encode()).hexdigest(), 16) % 100
return "canary" if bucket < canary_percent else "prod"
alias = pick_alias(request.user_id, canary_percent=5)
bundle, vid = resolve("ticket_classifier", alias, ALIASES)
result = call_model(bundle, ticket_text=request.text)
log.info("llm_call", prompt="ticket_classifier", prompt_version=vid, alias=alias,
model=bundle["model"], latency_ms=result.latency_ms, parsed_ok=result.parsed_ok)Rollback is moving the alias back, which takes effect within the resolution cache's lifetime. Two things can make it harder than that. Outputs stored under a new schema may not be readable by consumers of the old one, so major changes need consumers that accept both formats during the rollout. And caches keyed only by input will serve outputs from the bad version after rollback; include the version id in any cache key for model outputs.
Worked example: a ticket classifier release
A support team's classifier at version 1.4.2 routes tickets into five categories. Agents report that login problems are often classified as billing. An engineer adds two instructions and four new examples, producing 1.5.0. The figures that follow are illustrative. On the 1,200-ticket evaluation set, overall accuracy rises from 0.91 to 0.93 and the account-access slice from 0.82 to 0.90, but the Spanish-language slice falls from 0.88 to 0.81, because all the new examples are English and the model now over-weights their phrasing. The slice gate blocks the merge.
The engineer adds two Spanish examples, producing a new bundle with a new hash. All slices are now within tolerance and the change merges. In canary at 5% for a day, parse failures and latency match production, and agent reclassification of canary tickets falls. The alias moves to 25% and then 100%. A week later the provider announces a retirement date for the pinned snapshot; the team runs 1.5.1, identical except for the newer snapshot, through the same gate, finds formatting drift in the output, adjusts one instruction in 1.5.2 and ships it through the same canary process. Every trace from the period identifies which of the four versions handled the ticket.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
| Outputs changed, nobody edited the prompt | model alias, examples or schema outside the bundle | content-hash the whole bundle; pin snapshots |
| Cannot say which prompt produced a bad answer | version id not logged | log the hash on every call |
| Average improved, complaints rose | slice regression | gate on slices |
| Rollback did not fix it | output cache keyed without version | include the version id in cache keys |
| Parser errors after release | output contract changed as a minor | classify contract changes as major; dual-read |
| Two environments disagree | label reused for different content | identity by hash, labels advisory |
Trade-offs
Git-based versioning gives you review, history and atomic changes with code for free, but makes every prompt change a deploy unless aliases are resolved at runtime. A registry service lets product owners iterate without engineers and allows instant rollback, but needs its own access control, audit and availability. Content hashing adds a little machinery and guarantees honest identity. Strict evaluation gates slow iteration and catch the regressions that cost the most. For most teams the right shape is bundles in git, hashes as identity, aliases resolved at runtime with a short cache, and evaluation as a required check.
What to do next
- List every input to each production prompt call and move anything outside the bundle, especially examples and schemas, into it.
- Compute a canonical content hash per bundle and log it, with the model id, on every model call.
- Replace moving model aliases in bundles with dated snapshots, and record each snapshot's retirement date.
- Add a CI gate that compares candidate and production evaluation reports overall and per slice.
- Introduce prod and canary aliases with deterministic user bucketing, and rehearse a rollback.
- Add the version id to every cache key for model outputs.