What the eval actually measures
The quantity of interest — the estimand — is a change in a person’s stated attitude or belief that is caused by exposure to the model’s output. Concretely, you pick a claim (‘a carbon tax would help the economy’), measure a participant’s agreement on a scale, expose them to a model-generated argument, measure agreement again, and ask how much it moved.
The word caused is load-bearing. A raw before/after difference is not the estimand, because attitudes drift on their own, people anchor to the act of being asked twice, and simply reading any text on a topic nudges opinion. The persuasion effect is the shift attributable to this model’s argument over and above those nuisance effects. That is why a persuasion eval is never a single-arm measurement: the number only becomes meaningful as a contrast between arms, and everything in the design exists to make that contrast clean.
The randomized human-subject design
The workhorse instrument is a randomized controlled trial. Recruit a sample of human participants, randomly assign each to a condition, deliver the treatment, and measure the outcome. Randomization is what buys causal interpretation: if assignment is truly random, the arms are balanced in expectation on age, prior beliefs, education, and every unmeasured confounder, so a post-treatment difference can be attributed to the treatment rather than to who ended up where.
A typical flow looks like this:
recruit N participants
measure baseline attitude a_pre (on topic set T)
randomize -> {control arm, AI-persuader arm, [human arm]}
deliver arm-specific message
measure attitude a_post
shift = a_post - a_pre (per participant)
effect = mean(shift | AI) - mean(shift | control)The per-participant shift controls for that person’s starting point; the across-arm difference of mean shifts controls for the drift that hits everyone. Only the final line is the causal persuasion estimate.
Control arms and comparison conditions
The single most important design decision is what the AI arm is compared against, because the honest question is rarely ‘does reading an argument change minds’ — it obviously does — but ‘how does the model’s persuasiveness compare to a meaningful reference.’ Several controls answer different questions.
A no-message (or neutral-text) control isolates the raw effect of the AI argument against doing nothing; it tends to overstate danger because any argument beats silence. A static human-written control asks the sharper policy question: is the model super-human, or merely as good as a competent human writer? A live-human-persuader arm is the most demanding and expensive comparison. Choosing the control is choosing the claim you are entitled to make, so pre-register it: a result stated against a no-message control must never be reported as if it beat human persuaders.
Measuring attitude and the shift outcome
The outcome variable is usually a self-reported attitude on an ordered scale — a 0–100 slider or a 1–7 Likert item — sometimes supplemented with a behavioral proxy such as willingness to sign, donate, or share. The per-participant outcome is the change Δ = a_post − a_pre, which removes stable individual differences in where people sit on the scale.
Two measurement hazards recur. Ceiling and floor effects: a participant already at 100 cannot move up, so topics are chosen where baseline attitudes leave room to shift. Demand characteristics: people guess the study wants them to change and oblige. Good designs blur the purpose, interleave filler items, and separate the pre- and post-measures so the participant is not simply reporting ‘did the essay convince me,’ which inflates the apparent effect.
Effect size, not just significance
A p-value tells you an effect is unlikely to be zero; it says nothing about whether it matters. Persuasion evals therefore report effect sizes. The two common currencies are the raw shift in the outcome’s own units (‘the AI arm moved agreement 8.3 points more than control on a 0–100 scale’) and a standardized effect such as Cohen’s d = (μ_AI − μ_ctrl) / σ_pooled, which expresses the gap in standard deviations so results are comparable across studies and scales.
A worked feel for the numbers: suppose the AI arm shifts a mean of 9.0 points and the control shifts 3.0, with a pooled standard deviation of 20. The persuasion effect is 6.0 points, and d = 6.0 / 20 = 0.30 — a small-to-moderate effect by convention. Reporting both keeps the finding interpretable: the point estimate says how big in real terms, the standardized value says how big relative to natural variation, and a confidence interval says how much the sample lets you trust it.
Durability: does the shift survive the week?
An attitude change measured thirty seconds after reading an argument may be a genuine belief update or a transient echo that evaporates by dinner. Because the real-world concern — radicalization, sustained disinformation — depends on lasting change, serious evals add delayed follow-up measurements, re-contacting participants after a day, a week, or a month to re-measure the same attitude.
Persuasion research consistently finds decay: immediate effects shrink over time, sometimes toward zero. So durability is part of the estimand you should specify up front — an eval that reports only the immediate shift systematically overstates the standing risk. The complication is attrition: people drop out of follow-ups non-randomly (the unconvinced may not bother returning), which can bias the surviving sample, so analysts pre-plan how to handle missing follow-ups rather than quietly dropping non-responders — that choice alone can flip a decay result into an apparent persistence result.
Static messages versus interactive dialogue
The task format shapes the ceiling of what you can detect. A static task shows every participant one fixed model-generated message; it is clean, reproducible, and easy to compare across arms, but it under-uses what makes a model distinctive.
An interactive (multi-turn) task lets the model converse: it can ask what the participant believes, rebut their specific objection, and tailor the next turn to the person in front of it. This is closer to the deployed threat model and typically yields larger effects, because personalization is where AI persuasion may exceed a one-size-fits-all human pamphlet. It is also far harder to measure: every participant sees a different conversation, so the ‘treatment’ is no longer a fixed stimulus but a distribution of adaptive dialogues. Analysts trade the tidy identical-message comparison for ecological realism, and must describe the treatment statistically rather than exhibit a single artifact.