Most organisations that build with large language models have an AI ethics statement: fairness, transparency, accountability, privacy, safety, human oversight. Almost none of those statements has ever stopped a release, changed a feature or paid for a fix. The words are right; what is missing is the machinery that turns them into decisions an engineer can act on and evidence a reviewer can check.
This article treats AI ethics as that machinery. It assumes you are building or operating an LLM product and want a repeatable process: identify who can be hurt and how, decide in writing what you will do when values conflict, convert decisions into controls and measurable thresholds, gate releases on evidence, and give people who are harmed a way to be heard. The running example is a homework tutor for 13 to 17 year olds, chosen because it puts real pressure on every principle at once.
Why principles alone do not ship
Principles fail in production for three structural reasons. First, they are unranked: transparency and privacy, autonomy and protection, helpfulness and caution regularly pull in opposite directions, and a list gives no rule for which wins. Second, they are unowned: "we value fairness" has no team, no budget and no on-call rotation. Third, they are unmeasured: without a metric and a threshold there is no moment at which a release is visibly blocked.
The fix is not better principles but a pipeline that forces each of those gaps closed. Every principle must become, for a specific product, a list of concrete harms; every harm must get an owner, a control, a piece of evidence and a threshold; and every conflict between principles must be resolved by a named person in a dated, written decision that the next team can read and challenge.
Step 1: map stakeholders and harms
Start with people, not model behaviour. For the tutor the stakeholders are students, parents, teachers, schools, and people mentioned in student work. Then enumerate harms per stakeholder across five families: quality harms (wrong answers presented confidently), allocation harms (worse service for some groups, for example non-native English speakers), dignity and safety harms (demeaning output, mishandled disclosures of abuse or self-harm), autonomy harms (doing the homework instead of teaching, building dependence), and privacy harms (retaining or training on minors' writing without meaningful consent).
Score each harm on severity, likelihood and reversibility, and record who bears it. The last field matters more than it looks: a harm borne by the person least able to notice it, such as a student who cannot tell a correct derivation from a fluent wrong one, deserves more weight than the raw likelihood suggests. A harm that cannot be undone, such as a missed crisis disclosure, should be treated as blocking regardless of how rare it is.
Keep the register as a versioned file next to the code, so changes go through review like any other change:
# harms.yaml -- one entry per harm, owned and versioned with the product
- id: H3
harm: "Tutor gives a confidently wrong worked solution"
who: [student, parent, teacher]
severity: medium # wrong grade, lost trust; reversible
likelihood: high
reversibility: high
bearer: student # the person least able to detect it
controls: [show-reasoning-steps, cite-textbook-section, 'flag "check this" on low confidence']
evidence: eval/math_accuracy.jsonl
threshold: {metric: step_accuracy, min: 0.92, slice_min: 0.88}
owner: learning-quality
- id: H7
harm: "Self-harm disclosure handled as a homework question"
who: [student]
severity: critical # irreversible
likelihood: low
reversibility: none
bearer: student
controls: [crisis-classifier, fixed-safe-response, human-escalation-path]
evidence: eval/crisis_recall.jsonl
threshold: {metric: recall, min: 0.98}
owner: trust-safety
Step 2: turn value conflicts into written decisions
Every non-trivial product hits conflicts the principles do not settle. The tutor hits at least four. Autonomy versus learning: students want answers; the product's purpose is understanding. Privacy versus safety: detecting self-harm disclosures requires reading content you would otherwise minimise. Transparency versus misuse: explaining exactly how the integrity filter works teaches students to evade it. Helpfulness versus equity: a feature that works well only in English widens a gap if it launches only in English.
Do not resolve these in meetings that leave no trace. Write a short decision record per conflict: the options considered, who is affected by each, the choice, the reasoning, the residual risk accepted, the owner, and the condition that would reopen the decision. For the tutor, a defensible set is:
| Conflict | Decision | Reopen if |
|---|---|---|
| Autonomy vs learning | Hints and worked steps first; full solutions only after an attempt, with the attempt shown | Teachers report it blocks legitimate revision |
| Privacy vs safety | Run a crisis classifier on every message; store only the classifier outcome, not the text, unless escalated | Recall below threshold, or escalation volume exceeds staff capacity |
| Transparency vs misuse | Disclose that integrity checks exist and what they are for; do not publish their mechanics | Schools require auditability |
| Helpfulness vs equity | Launch only languages whose step accuracy meets the same threshold as English | Slice gap narrows below 2 points |
A written decision is what makes accountability real: it identifies who chose, on what evidence, and gives a later reviewer something specific to disagree with. The AI Bill of Rights article covers per-decision records for individual automated outcomes; this is the same idea applied to design choices.
Step 3: scope the use and restrict misuse
Ethical scope is a technical artefact, not a sentence in the terms of service. Write an intended-use statement ("homework help for students aged 13 to 17 in maths and science, supervised by a school") and an explicit prohibited-use list ("not for grading, not for disciplinary decisions, not for counselling"). Then enforce the boundary in the product: topic classifiers that redirect out-of-scope requests, an API contract that refuses to return per-student risk scores to schools, and rate limits that prevent the tutor being used as a general essay generator.
The prohibited list matters most where a downstream customer could repurpose your output into a consequential decision. If a school could use conversation logs to flag students for discipline, you have built a surveillance system whatever your intent. Removing the capability, by not exposing the data at all, is stronger than a contractual promise not to use it.
Step 4: data provenance and consent
Ask three questions of every dataset that touches the system. Where did it come from, and does its licence permit this use? Who is in it, and did they, or for minors their guardians, have a meaningful choice? What happens to it next: is user content retained, for how long, and does it feed training?
For the tutor, a defensible default is that student conversations are not used for training, are retained for a short fixed period for safety review, and are deletable on request, with a separate opt-in for research use that is off by default. Evaluation sets should be built from consented or synthetic material and checked for the same demographic and language slices as production traffic, or the evidence in Step 5 will describe a population you do not serve. Privacy law, such as the obligations in GDPR for LLM systems, sets a floor; the ethical question is whether the people affected would be surprised by what you do.
Step 5: release gates built on evidence
Each harm in the register points to an evaluation and a threshold. A release candidate runs every evaluation, and a gate compares results against thresholds, including the worst slice, not only the average. The fairness article explains how to compute per-group metrics with uncertainty; the safety evaluation article explains what a jailbreak or refusal number does and does not license. The gate itself is simple:
import json, yaml
def load(path):
return [json.loads(l) for l in open(path, encoding="utf-8")]
def release_gate(register_path, results):
# results: {evidence_path: {"metric": value, "slices": {name: value}}}
blocking, waivable = [], []
for h in yaml.safe_load(open(register_path)):
failures = blocking if h["severity"] == "critical" else waivable
r = results.get(h["evidence"])
if r is None:
failures.append((h["id"], "no evidence for this build"))
continue
t = h["threshold"]
if r["metric"] < t["min"]:
failures.append((h["id"], f"{t['metric']}={r['metric']:.3f} < {t['min']}"))
worst = min(r["slices"].values(), default=r["metric"])
if worst < t.get("slice_min", t["min"]):
failures.append((h["id"], f"worst slice {worst:.3f}"))
# critical harms always block; the rest ship only with a named, dated waiver
return blocking, waivableTwo rules make the gate trustworthy. Missing evidence is a failure, never a pass, so an evaluation that silently stopped running blocks the release. And failures on non-critical harms can be waived only by a named owner with an expiry date, recorded beside the build, so waivers become visible debt rather than quiet precedent. For the tutor, H7 (crisis recall) is never waivable.
Step 6: disclosure, redress and human oversight
Transparency is useful only if it is aimed at a decision someone can make. Students need to know they are talking to an AI, that it can be wrong, and what to do when it is. Parents and schools need a plain-language system card: intended use, prohibited uses, languages supported, known weaknesses with measured rates, data handling, and a change log. Regulators may need more; the EU AI Act compliance guide maps transparency duties to engineering artefacts.
Redress means a person who is harmed can get the harm noticed and fixed. Give every response a one-click "this is wrong or harmful" control, route reports into a triaged queue with a response time, and close the loop: each confirmed report either adds a test case, updates the harm register, or both. Human oversight for the crisis path means a trained person reachable within a defined time, not a reviewer reading logs a week later.
Worked example: the tutor&amp;amp;amp;amp;amp;#x27;s first release review
Suppose the release candidate produces these results. Step accuracy is 0.94 overall but 0.86 for Spanish-language maths. Crisis recall is 0.985 on a 400-item set. Full-solution leakage, where the tutor gives complete answers without an attempt, is 7 percent against a 5 percent threshold.
The gate output is three findings. Crisis recall passes, but with 400 items the confidence interval is wide, so the owner commissions a larger set before the next release. Spanish fails the slice threshold of 0.88; per the written decision, Spanish launches later, and the team records the gap as a dated equity debt. Solution leakage fails a non-critical threshold; the product lead signs a two-week waiver because the fix, a stricter attempt check, is already in review. The release ships in English only, with three tracked items, and every one of those choices can be reconstructed months later from the register, the decision records and the waiver log.
Monitoring and the review cadence
Offline evaluation describes the past; production describes the present. Monitor the same harms in production with cheap proxies: report rates per thousand sessions, escalation counts, out-of-scope redirect rates, language mix drift, and a weekly sample of conversations reviewed by trained staff under the retention policy. Treat a breached threshold as an incident with an owner, a timeline and a post-incident update to the register.
Re-run the full harm mapping on a fixed cadence, quarterly for most products, and immediately when scope changes: a new age group, a new language, a new tool such as web search, or a new customer type. Scope changes are where most ethical failures enter, because the original analysis silently stops applying. The AI safety frameworks article shows how this cadence lines up with NIST AI RMF and ISO/IEC 42001 if you need to report against them.
Failure modes
| Failure | What it looks like | Countermeasure |
|---|---|---|
| Ethics-washing | Principles page, no owners, no blocked releases | Track how many releases the gate changed; zero is a warning sign |
| Metric fixation | Average passes, worst slice ignored | Gate on slice minimums and show them first |
| Diffused responsibility | Committee approves; nobody owns the harm | One named owner per harm and per waiver |
| Stale analysis | Harm map written for v1, product is now v4 | Scope-change trigger plus fixed cadence |
| Redress theatre | Report button feeds an unread inbox | Response-time target and a test case per confirmed report |
| Evidence drift | Evals stop running and nobody notices | Missing evidence fails the gate |
Trade-offs
Every control has a cost, and pretending otherwise produces controls that get quietly removed. Stricter gates slow releases; a crisis classifier with high recall produces false alarms that consume human time; refusing full solutions frustrates students with legitimate revision needs; delaying a language launch for quality withholds a useful product from people who might prefer an imperfect one. The goal is not to minimise risk at any price but to make the price explicit, decided by someone accountable, and revisited when evidence changes.
What to do next
- Write an intended-use statement and a prohibited-use list for one product, and enforce at least one prohibition in code.
- Create a versioned harm register with severity, likelihood, reversibility, bearer, owner, control, evidence and threshold for each harm.
- List the value conflicts you have already resolved informally and write a decision record for each, with a reopen condition.
- Wire a release gate that fails on missing evidence and on worst-slice thresholds, with named, expiring waivers.
- Publish a plain-language system card and add a report control to every response, with a response-time target.
- Monitor harm proxies in production and treat breaches as incidents that update the register.
- Schedule a quarterly harm-mapping review and trigger one on every scope change.