Once a company ships more than one LLM feature, the hard problems stop being about wording. Someone has to own the evaluation set that tells you whether a change helped. Someone has to decide whether a new model version can replace the old one across a dozen prompts. Someone has to answer the page when the extraction feature starts returning malformed JSON at two in the morning. If nobody owns these things, prompts drift, evaluations rot, and every model upgrade becomes a frightening big-bang change.
This article is about structuring the people and responsibilities around prompts, not about prompting techniques. It inventories the work, compares three operating models, defines roles, shows how to encode ownership in a repository with a CI check, walks a growing company through the transitions, and ends with metrics and a checklist. It deliberately quotes no industry survey numbers; the structures here follow from where the work and the risk actually sit.
The work that actually exists
Before choosing a structure, list the work. Most of it is not writing prompts.
- Authoring and iteration: writing prompts, templates and tool descriptions, and fixing regressions. Typed templates are covered in prompt templates in depth.
- Evaluation datasets: collecting representative and adversarial cases, labelling expected outputs, and keeping the set current as the product changes. This needs domain knowledge more than prompt skill.
- Evaluation infrastructure: the harness, graders, dashboards and CI integration, described in prompt evaluation architecture.
- Release mechanics: versioning, staged rollout and attribution of outputs to prompt versions, which is the job of a prompt registry.
- Model selection and migration: testing new model versions, managing deprecations, and balancing cost and latency.
- Safety and security review: prompt injection exposure, tool permissions, data handling.
- Production operations: monitoring, incident response and debugging live failures.
The first two items belong close to the product because they need product context. The middle three are shared infrastructure that every feature repeats. The last two need specialist judgement and consistent standards. A good structure puts each item where that property is satisfied.
Three operating models
Centralised prompt team. One team writes and maintains every prompt and product teams file requests. This gives consistency and concentrates expertise, and it works when there are one or two features. It fails as features multiply: the central team lacks product context, becomes a queue, and product engineers learn to route around it by embedding prompt strings in code.
Fully embedded. Every product team writes its own prompts with no shared function. Context and speed are excellent. The cost appears later: five evaluation harnesses, five ways of calling models, no one who can test a model upgrade across features, and safety review that varies by team.
Hub and spoke. Product teams own their prompts, evaluation sets and on-call; a small platform hub owns the registry, the evaluation harness, the model gateway, cost tracking and migration tooling; a guild of practitioners across teams shares patterns and reviews; security and safety review high-risk changes. This is the model most organisations converge on once they have several features, because it puts each item from the inventory where it belongs.
Roles and responsibilities
| Role | Owns | Typical home |
|---|---|---|
| Feature owner | Outcome metrics, priorities, final ship decision | Product team |
| Prompt author | Prompt and template changes, regression fixes | Product team engineer |
| Evaluation owner | Eval set coverage, labelling quality, thresholds | Product team, with domain experts |
| Domain reviewer | Judging correctness where graders cannot | Support, legal, clinical or other experts |
| Platform engineer | Registry, harness, gateway, migration tooling | Platform hub |
| Safety and security reviewer | Injection exposure, tool scope, data use | Security team |
| On-call engineer | Live failures, rollback, incident review | Product team rotation |
Two rules make the table work. First, every prompt has exactly one owning team, recorded in the repository, not in a wiki. Second, prompt author is a responsibility, not a job title. In most organisations the people best placed to write a prompt are the engineers who own the feature, supported by the guild; a dedicated prompt engineer title is useful mainly in the hub or as a senior guild lead who reviews and teaches.
Encoding ownership in the repository
Structure is only real when tooling enforces it. Keep each prompt in its own directory with a manifest, use the code host's code-owners file so the owning team must approve changes, and add a CI check that refuses prompts without an owner, an evaluation suite and thresholds, and that requires a safety reviewer for prompts with tool access.
# prompts/support_reply/manifest.yaml
# id: support_reply
# owner: team-support-assistant
# eval_suite: evals/support_reply.jsonl
# thresholds: {task_success: 0.92, policy_violation_max: 0.0}
# tools: [refund_lookup]
# risk: high # tool access => security review required
import pathlib, sys, yaml
REQUIRED = {"id", "owner", "eval_suite", "thresholds", "risk"}
errors = []
for manifest in pathlib.Path("prompts").glob("*/manifest.yaml"):
m = yaml.safe_load(manifest.read_text())
missing = REQUIRED - m.keys()
if missing:
errors.append(f"{manifest}: missing {sorted(missing)}")
continue
if not pathlib.Path(m["eval_suite"]).is_file():
errors.append(f"{manifest}: eval suite {m['eval_suite']} not found")
if m.get("tools") and m["risk"] != "high":
errors.append(f"{manifest}: prompts with tools must be risk: high")
if errors:
print("\n".join(errors))
sys.exit(1)
# CODEOWNERS (excerpt; the LAST matching pattern wins, so generic lines go first)
# /evals/ @org/llm-platform # harness and shared graders
# /prompts/support_reply/ @org/team-support-assistant # prompt and manifest
# /evals/support_reply.jsonl @org/team-support-assistant # the team owns its eval set
# /tools/manifest_check.py @org/llm-platform # this check belongs to the hubA second CI job runs the prompt's evaluation suite on every change and fails the build if a threshold drops, posting a diff of failed cases to the pull request. The reviewer then reads the failures, not just the prompt diff, which is where most regressions become visible.
The change workflow
- Propose: the author opens a pull request with the prompt change and the reason, ideally linked to failing cases from production.
- Evaluate offline: CI runs the suite; new failure cases found while debugging are added to the eval set in the same change.
- Review: a code owner from the product team reviews outputs and failures; high-risk prompts also need the safety reviewer.
- Roll out: the registry ships the new version to a small share of traffic, attributing every output to its version.
- Watch and promote: the feature owner checks online metrics, then promotes or rolls back.
When a change fails in production anyway, follow the capture, reproduce, minimise loop from debugging broken prompts and add the case to the suite so the same failure cannot ship twice.
Worked example: a company growing from one feature to six
Consider an illustrative company of about forty engineers. It launches a support reply assistant built by one product team, with prompts in code and a spreadsheet of fifty test cases. That is fine: fully embedded is the right model for one feature.
A year later there are four features. Symptoms appear: two teams wrote separate evaluation scripts, nobody knows the total model spend per feature, and a deprecation notice for the model version all four use causes a scramble. The company forms a guild with one member from each team, standardises on one evaluation harness, and moves prompts into manifest directories with owners. One engineer from the support team, the most experienced with evaluation, rotates into a half-time platform role.
With six features, including one that can issue refunds through a tool, it creates a platform hub of three engineers who own the gateway, registry, harness and a migration playbook, and it requires security review for prompts with tool access, following the patterns in prompt-injection defence.
Then a new model version arrives. The hub runs every feature's suite against it in one batch and publishes a table: four features meet thresholds unchanged, one needs prompt edits, and one regresses on a policy check. Each owning team fixes its own prompt against its own suite; the hub coordinates the cut-over date and keeps the old version routable for rollback. A migration that used to take weeks of uncoordinated effort becomes a scheduled change, because ownership, suites and tooling were in place before it was needed.
Rituals and interfaces between teams
Ownership boundaries only hold if the teams have a few lightweight, recurring touch points. Four are enough for most organisations, and each has a clear input and output.
- Weekly evaluation review, run inside each product team: the evaluation owner brings the cases that failed in the past week, from CI and from production sampling, and the team decides which become permanent suite entries, which reveal a labelling error, and which need a prompt change. The output is a changed suite, not a meeting note.
- Guild session, every two weeks across teams: one team presents a regression and its fix, or a technique that worked, and the guild decides whether it becomes a shared template, grader or lint rule. This is how a lesson learned by the search team reaches the extraction team without a central owner.
- Model and cost review, monthly, led by the hub: spend and latency per feature, upcoming deprecations, and candidate model versions with their batch evaluation results. Product teams leave with dates, not surprises.
- Prompt incident review, after any user-visible failure: the same blameless format used for service incidents, with one extra question, which evaluation case would have caught this, and an action to add it.
The interface between hub and product teams should be written down as a short contract: the hub guarantees the harness, registry and gateway with an on-call of its own, publishes migration results with a fixed notice period, and never edits another team's prompt content; product teams keep their manifests valid, their suites current and their on-call staffed. Disagreements about thresholds go to the feature owner, because the threshold encodes a product decision.
Measuring whether the structure works
- Eval coverage: share of production prompts with an owner, a suite and thresholds. The target is all of them.
- Lead time for a prompt change: from pull request to full rollout. Long times usually mean review or tooling bottlenecks in the hub.
- Regression escape rate: incidents caused by prompt or model changes that passed CI. Each one should add cases to a suite.
- Migration duration: time from a new model version being available to all features deciding to adopt or skip it.
- Cost and latency per feature: visible to the owning team, not only to the platform.
Failure modes
- The central queue. A central team becomes a bottleneck and product teams bypass it with hard-coded prompts.
- Orphaned prompts. A team is reorganised and its prompts lose an owner; nobody updates their suites. The CI ownership check catches this if owners map to live teams.
- Rotting evaluation sets. Suites written at launch no longer reflect real traffic, so passing CI means little. Budget recurring time for the evaluation owner.
- Hub overreach. The platform team starts approving every prompt change, recreating the central queue. The hub owns tools and standards; product teams own content.
- Unreviewed tool access. A prompt gains a tool in a routine change without security review. Enforce the risk rule in CI, not by convention.
Trade-offs
| Model | Strengths | Weaknesses | Fits |
|---|---|---|---|
| Centralised | Consistency, concentrated expertise | Queueing, weak product context | One or two features |
| Embedded | Speed, product context | Duplicated tooling, uneven safety | Early stage, one team |
| Hub and spoke | Context plus shared tooling and standards | Needs a funded hub and clear boundaries | Several features and teams |
The evaluation-first discipline that makes any of these work is described in eval-driven development; structure only decides who does it.
What to do next
- List every prompt in production and record its owning team; anything without one is your first problem.
- Write down the inventory of work from this article and mark who does each item today.
- Move prompts into manifest directories and add the CI ownership and evaluation checks.
- Pick one person per product team as the evaluation owner and give them recurring time.
- Write a model migration playbook before the next deprecation notice, and run it once on a non-critical feature.
- Start tracking eval coverage, prompt change lead time and regression escapes monthly.