Know Your Customer is the part of a payments business that decides who you are willing to hold an account for — and it is the part that breaks first when the entity pressing ‘buy’ is software. In a classic flow a human onboards, proves who they are, gets a risk rating, and is monitored from then on. In an AP2 flow the same human onboards, but the thing that later shows up at the merchant is an agent carrying a signed mandate. Nothing about that removes the KYC obligation; it relocates it. This piece walks the onboarding programme itself: how identity evidence is gathered and graded, how risk tiers gate what an account can do, who the customer actually is when an agent transacts, how refresh keeps the file honest, why sanctions screening is a different control, why one global policy cannot work, and what over-tight screening quietly costs you.
KYC is a programme, not a gate
The most common architectural mistake is treating KYC as a one-time gate at signup. It is a continuing programme with three broad limbs: identifying and verifying the customer at onboarding, forming and maintaining a risk-based understanding of what that customer is likely to do, and monitoring activity against that understanding for as long as the relationship lasts.
That framing matters for agentic payments because an agent’s behaviour is where the second and third limbs get tested. An account whose expected profile was ‘occasional consumer purchases’ and which suddenly emits hundreds of machine-paced transactions has not necessarily done anything wrong — but its profile is stale, and a programme that looked at the customer once has no mechanism to notice. Designing KYC as a living record with an update path, rather than a signup form with a pass/fail outcome, determines whether the rest of this architecture is possible.
The onboarding flow and what each evidence type proves
Onboarding runs in three beats: collect claimed identity attributes, verify them against independent evidence, and decide — accept, reject, or escalate to manual review. The interesting engineering is in the middle beat, because different evidence types prove genuinely different things and are not interchangeable.
Document checks (a passport or licence image, with security-feature and tamper analysis) establish that a credential exists and appears genuine. Biometric checks — a selfie matched to the document photo, plus liveness detection — establish that the person presenting it is the person it depicts and is present rather than a replayed image. Database and bureau matches confirm the claimed identity has a history: electoral rolls, credit files, telco or utility records, government registries where available. Instrument-based checks, such as verifying control of a named bank account, bind the identity to a funding source. A robust programme layers these; any one alone has a well-known bypass.
Risk-based tiering — what each tier unlocks
Verification is not binary, and treating it that way is expensive in both directions. Mature programmes assign a risk tier at onboarding from the strength of evidence gathered plus inherent risk factors — geography, funding instrument, expected activity, whether the relationship is face-to-face or entirely remote — and then let the tier decide what the account can actually do.
The tier is a capability grant. A lightly verified account might be limited to small, domestic, low-velocity purchases; a fully verified one unlocks higher ceilings, cross-border activity, irreversible rails, and delegation to agents. Because the tier is data rather than code, it also gives you a clean upgrade path: a customer who wants a capability they lack is asked for the evidence that would justify it, rather than refused. For AP2 specifically, the tier is the natural place to express whether an account may grant mandates to an agent at all, and to bound the value those mandates may carry.
Who is the customer when an agent transacts?
This is the question AP2 forces, and the answer is less exotic than it looks: the customer is the human or legal entity on whose behalf the agent acts. An agent is not a legal person, cannot hold obligations, and cannot be the subject of customer due diligence. It is a channel — closer in kind to a mobile app or a standing instruction than to a new account holder.
What genuinely changes is the evidentiary chain. When an agent transacts, the principal is absent by design, so authority has to be carried explicitly. That is precisely the role AP2’s mandates play: signed, verifiable statements of the user’s intent that travel with the transaction. The KYC programme’s job is not to onboard the agent but to make sure every mandate resolves to a known, verified, currently-in-good-standing customer, and that the link between the two is recorded rather than inferred.
Recording delegated authority in the customer file
If the mandate is the evidence of authority at transaction time, the customer file has to be where that delegation is registered. Practically, the onboarding and ongoing-diligence record should capture which agents a customer has authorised, the keys or credentials those agents sign with, the scope and limits attached, when the authorisation was granted, and — critically — when and how it can be revoked.
Two properties make this worth building carefully. First, revocation must be immediate and centrally observable: a customer who withdraws an agent’s authority should not have to trust that every downstream party notices. Second, the record must support attribution after the fact. When a disputed transaction is examined months later, the question is not only whether the mandate signature was valid but whether this agent was authorised by this customer, within these limits, on that date — a due-diligence record, not a runtime check.
Refresh: periodic cycles and trigger-based re-verification
Customer information decays. Addresses change, companies restructure, beneficial ownership shifts, documents expire, and a risk rating set at onboarding slowly stops describing reality. Programmes handle this two ways, and both are needed.
Periodic refresh re-examines the file on a cadence driven by risk tier — higher-risk relationships reviewed far more often than low-risk ones. (Any specific interval is illustrative; the cadence is a policy choice, not a universal constant.) Trigger-based re-verification is the more valuable half: an event makes the existing file untrustworthy, so you re-verify immediately rather than waiting for the cycle. Typical triggers include a document reaching expiry, a sanctions or adverse-media hit, a step change in transaction volume or geography, a tier upgrade request, a change in ownership or control, or a customer authorising an agent for the first time. Agentic activity deserves explicit trigger design, because its behavioural signature legitimately differs from a human’s.
Sanctions and PEP screening is a separate control
Screening is routinely bundled with identity verification and should not be. Verification asks are you who you claim to be; screening asks are you someone we are prohibited or restricted from dealing with. A customer can pass one perfectly and fail the other, and the two controls have completely different failure modes.
Screening matches customer data against sanctions designations, politically-exposed-person lists, and adverse-media sources. It is fuzzy by necessity — transliteration, name ordering, partial dates of birth, and common names all defeat exact matching — so it produces alerts requiring disposition, not verdicts. Three design points follow. Screening must run at onboarding and continuously, because lists change while your customers do not. List ingestion latency is itself a control: a designation you have not loaded is one you cannot enforce. And you screen the customer, not the agent — the agent has no identity for a list to designate.
Jurisdictional divergence — why one global policy fails
KYC rules are set nationally, and they diverge on nearly every axis that matters: which identity attributes must be collected, which documents are acceptable, whether remote onboarding is permitted, how long records must be kept, what triggers enhanced due diligence, and how beneficial ownership is defined. Data-protection law adds a second axis, since biometric and identity data are among the most tightly regulated categories and may not be free to leave a region.
The engineering consequence is that jurisdiction cannot be a configuration afterthought. A workable architecture separates a common pipeline — collect, verify, score, screen, decide, record — from a per-jurisdiction policy layer that supplies required attributes, acceptable evidence, thresholds, and retention rules as versioned data. Hard-coding one market’s rules into the pipeline guarantees a rewrite at the first expansion; treating them as policy makes a new market a data change and a review, not a re-architecture.
Audit trail, retention, and reconstructability
The obligation is not merely to make correct decisions but to be able to demonstrate that you made them, sometimes years later. That is a stronger requirement than logging. It means capturing not just the outcome but the inputs, the evidence, the policy version in force, the vendor or model that produced each score, the alert and its disposition, and whoever made a manual call — enough to reconstruct the decision as it stood at the time.
Two forces pull against each other. Retention rules require you to keep records for defined periods after a relationship ends; data-protection rules require you to minimise what you hold and delete what you no longer need. The usual reconciliation is to keep the decision record for the retention period while limiting raw sensitive artefacts — storing a verified result and an evidence reference rather than an indefinite copy of every document image and biometric template.
The cost of screening too tightly
It is tempting to treat KYC as a control where more is always safer. It is not. Screening thresholds trade false negatives against false positives, and false positives are not free — they are simply paid by someone other than the compliance team.
Tighten too far and three things happen. Legitimate customers abandon onboarding, disproportionately those with thin files, recent moves, or names that match list entries poorly — a fairness problem, not only a revenue one. Review queues fill with alerts that are almost all noise, and analyst attention, the scarcest resource in the programme, gets spent clearing them; genuine hits are likelier to be missed in the flood. And agent-initiated flows break badly, because an agent cannot answer a step-up prompt on the principal’s behalf. Calibration is the real work: measure alert precision, queue latency, and abandonment alongside detection, and treat all of them as programme metrics.