Persuasion capability evals try to answer a deceptively simple question with real rigor: how much can a model move what a person believes? This is not a capability you read off a benchmark leaderboard the way you read off code-completion accuracy, because the ground truth lives inside human minds and only reveals itself through a controlled comparison. A persuasion eval is therefore a small social-science experiment wearing an AI-safety hat: you randomize people into arms, expose one arm to a model’s argument, measure the attitude they held before and after, and estimate the difference the model made against a fair control. This piece is about the measurement — the estimand, the experimental design, the effect-size math, durability, and the ethics constraints — not about how to write a persuasive message. Getting the measurement right is what separates a defensible claim about disinformation risk from a scary anecdote.

What the eval actually measures

The quantity of interest — the estimand — is a change in a person’s stated attitude or belief that is caused by exposure to the model’s output. Concretely, you pick a claim (‘a carbon tax would help the economy’), measure a participant’s agreement on a scale, expose them to a model-generated argument, measure agreement again, and ask how much it moved.

The word caused is load-bearing. A raw before/after difference is not the estimand, because attitudes drift on their own, people anchor to the act of being asked twice, and simply reading any text on a topic nudges opinion. The persuasion effect is the shift attributable to this model’s argument over and above those nuisance effects. That is why a persuasion eval is never a single-arm measurement: the number only becomes meaningful as a contrast between arms, and everything in the design exists to make that contrast clean.

Advertisement

The randomized human-subject design

The workhorse instrument is a randomized controlled trial. Recruit a sample of human participants, randomly assign each to a condition, deliver the treatment, and measure the outcome. Randomization is what buys causal interpretation: if assignment is truly random, the arms are balanced in expectation on age, prior beliefs, education, and every unmeasured confounder, so a post-treatment difference can be attributed to the treatment rather than to who ended up where.

A typical flow looks like this:

recruit N participants
measure baseline attitude a_pre  (on topic set T)
randomize -> {control arm, AI-persuader arm, [human arm]}
deliver arm-specific message
measure attitude a_post
shift = a_post - a_pre         (per participant)
effect = mean(shift | AI) - mean(shift | control)

The per-participant shift controls for that person’s starting point; the across-arm difference of mean shifts controls for the drift that hits everyone. Only the final line is the causal persuasion estimate.

Control arms and comparison conditions

The single most important design decision is what the AI arm is compared against, because the honest question is rarely ‘does reading an argument change minds’ — it obviously does — but ‘how does the model’s persuasiveness compare to a meaningful reference.’ Several controls answer different questions.

A no-message (or neutral-text) control isolates the raw effect of the AI argument against doing nothing; it tends to overstate danger because any argument beats silence. A static human-written control asks the sharper policy question: is the model super-human, or merely as good as a competent human writer? A live-human-persuader arm is the most demanding and expensive comparison. Choosing the control is choosing the claim you are entitled to make, so pre-register it: a result stated against a no-message control must never be reported as if it beat human persuaders.

Measuring attitude and the shift outcome

The outcome variable is usually a self-reported attitude on an ordered scale — a 0–100 slider or a 1–7 Likert item — sometimes supplemented with a behavioral proxy such as willingness to sign, donate, or share. The per-participant outcome is the change Δ = a_post − a_pre, which removes stable individual differences in where people sit on the scale.

Two measurement hazards recur. Ceiling and floor effects: a participant already at 100 cannot move up, so topics are chosen where baseline attitudes leave room to shift. Demand characteristics: people guess the study wants them to change and oblige. Good designs blur the purpose, interleave filler items, and separate the pre- and post-measures so the participant is not simply reporting ‘did the essay convince me,’ which inflates the apparent effect.

Effect size, not just significance

A p-value tells you an effect is unlikely to be zero; it says nothing about whether it matters. Persuasion evals therefore report effect sizes. The two common currencies are the raw shift in the outcome’s own units (‘the AI arm moved agreement 8.3 points more than control on a 0–100 scale’) and a standardized effect such as Cohen’s d = (μ_AI − μ_ctrl) / σ_pooled, which expresses the gap in standard deviations so results are comparable across studies and scales.

A worked feel for the numbers: suppose the AI arm shifts a mean of 9.0 points and the control shifts 3.0, with a pooled standard deviation of 20. The persuasion effect is 6.0 points, and d = 6.0 / 20 = 0.30 — a small-to-moderate effect by convention. Reporting both keeps the finding interpretable: the point estimate says how big in real terms, the standardized value says how big relative to natural variation, and a confidence interval says how much the sample lets you trust it.

Durability: does the shift survive the week?

An attitude change measured thirty seconds after reading an argument may be a genuine belief update or a transient echo that evaporates by dinner. Because the real-world concern — radicalization, sustained disinformation — depends on lasting change, serious evals add delayed follow-up measurements, re-contacting participants after a day, a week, or a month to re-measure the same attitude.

Persuasion research consistently finds decay: immediate effects shrink over time, sometimes toward zero. So durability is part of the estimand you should specify up front — an eval that reports only the immediate shift systematically overstates the standing risk. The complication is attrition: people drop out of follow-ups non-randomly (the unconvinced may not bother returning), which can bias the surviving sample, so analysts pre-plan how to handle missing follow-ups rather than quietly dropping non-responders — that choice alone can flip a decay result into an apparent persistence result.

Advertisement

Static messages versus interactive dialogue

The task format shapes the ceiling of what you can detect. A static task shows every participant one fixed model-generated message; it is clean, reproducible, and easy to compare across arms, but it under-uses what makes a model distinctive.

An interactive (multi-turn) task lets the model converse: it can ask what the participant believes, rebut their specific objection, and tailor the next turn to the person in front of it. This is closer to the deployed threat model and typically yields larger effects, because personalization is where AI persuasion may exceed a one-size-fits-all human pamphlet. It is also far harder to measure: every participant sees a different conversation, so the ‘treatment’ is no longer a fixed stimulus but a distribution of adaptive dialogues. Analysts trade the tidy identical-message comparison for ecological realism, and must describe the treatment statistically rather than exhibit a single artifact.

Statistical power and sample size

Persuasion effects are often small, and small effects demand large samples. Power is the probability your study finds a real effect of a given size; under-powered evals are doubly dangerous because they both miss real capability and, when they hit significance by luck, exaggerate the magnitude (the ‘winner’s curse’). Sample size scales roughly with the inverse square of the effect you want to catch: reliably detecting d = 0.2 needs on the order of hundreds of participants per arm, and each arm multiplies recruitment. This is why persuasion evals are expensive and why pre-registration matters — fixing the sample size, primary outcome, and analysis before seeing data prevents the garden of forking paths that would otherwise manufacture false positives from noise.

Ethics, consent, and the IRB

Deliberately trying to change real people’s real beliefs is human-subjects research, and it is fenced by ethics review for good reason. Studies pass through an institutional review board (or equivalent), obtain informed consent, and are bound to minimize harm. That creates a genuine tension: fully disclosing ‘we are testing whether an AI can change your mind on X’ contaminates the measurement via demand characteristics, yet concealment must never cross into harmful deception.

The standard resolutions are careful topic selection — favor low-stakes or counter-balanced positions over content that could entrench genuinely harmful views — a layer of authorized deception about the specific hypothesis where the board permits it, and thorough debriefing afterward that explains the true purpose and, when warranted, actively works to undo any induced shift. Consent, harm minimization, and debriefing are not bureaucratic friction; a persuasion result obtained without them is not usable evidence regardless of how clean the statistics look.

Reading the number responsibly

A persuasion eval outputs an effect size with a confidence interval, tied to a specific control, a specific population, a specific topic set, and a specific measurement horizon. The discipline is to carry all of that context with the number. ‘The model shifted attitudes by d = 0.3’ is meaningless alone; ‘d = 0.3 versus human-written messages, on US crowd-workers, on low-salience policy topics, measured immediately’ is a claim you can act on.

Tracked across model generations with a fixed methodology, these numbers become a capability trend line — the actual product of the exercise. The point is not one headline effect size but a reproducible instrument sensitive enough to say whether the next model is meaningfully more persuasive than the last, and trustworthy enough that a policy decision can rest on the comparison rather than a demo.

A persuasion capability eval is a controlled experiment, not a benchmark you read off a model: you randomize people into arms, measure attitude before and after, and estimate the causal shift the model’s argument produced over a fair control. Every design choice is a claim commitment — the control arm decides whether you can say ‘better than a human writer’ or only ‘better than silence’; the effect size and its confidence interval say how big and how trustworthy; the delayed follow-up says whether the shift lasts or decays. Interactive dialogue raises both the effect and the measurement difficulty, small effects demand large pre-registered samples, and consent, harm minimization, and debriefing are what make the evidence usable at all. The deliverable is a reproducible instrument that tracks persuasiveness across model generations — so a number always travels with its control, population, topics, and time horizon.