← All articles
Research note

financial compliance · regulation as code · continuous audit

Regulation Is Code. You're Running It on Humans.

Working note No. 14 — on point-in-time compliance and the systems that drift out from under it.

Abstract. Financial compliance is still assembled from spreadsheets, siloed point solutions, and headcount — and its central ritual is the point-in-time attestation: a consultant or an internal review examines the firm as it stands, and issues a green light. This note applies the vocabulary of the preceding thirteen to that ritual and finds it familiar. A compliance sign-off is an in-sample measurement of a continuously drifting system; a regulator's examination is the out-of-sample test; and the gap discovered between them is the same phenomenon this series has documented in every other domain — a snapshot mistaken for a property. We argue that the terminal architecture is the one regulators themselves have begun prototyping: obligations expressed as machine-readable rules, checked continuously against operations, with agents executing the checks and humans keeping the judgment calls. The green light should be a monitor, not a memory.

1. The ritual

Every regulated firm knows the choreography. An assessment is commissioned — external counsel, a Big Four team, or an internal review burning a quarter of the compliance calendar. Documents are gathered from the tools that hold fragments of the truth: the KYC platform, the transaction-monitoring vendor, the spreadsheet where the exceptions live, the other spreadsheet that reconciles the first one.1 Findings are remediated, or scheduled to be. And then the moment the entire exercise exists to produce: someone states, in writing, that the firm is compliant.

The statement is true the way a photograph is true. It describes the instant of exposure. The filing cabinet where it lands does not know that the firm kept moving.

2. Drift

The firm's compliance state is a function of at least three arguments — its obligations, its products, and its jurisdictions — and all three move continuously. Regulators publish; product teams ship; expansion adds a jurisdiction whose rulebook interacts with the existing ones in ways nobody priced. The arguments do not move independently, which is the expensive part: obligations multiply combinatorially across products and jurisdictions, while the machinery for tracking them — reviews, spreadsheets, headcount — scales linearly at best.2 A cost curve that compounds against one that accrues is not a management problem; it is an arithmetic verdict, delayed.

Fig. 1 draws the consequence. In the point-in-time regime, the gap between what the firm does and what its obligations require is reset to approximately zero at each assessment and then compounds unobserved — each launch, each rule change, each new market a step. The regulator does not arrive at t₀. The regulator arrives later, and examines the firm that exists then, against the obligations that exist then.

Compliance drift between point-in-time assessments

Fig. 1. This figure shows a firm's gap to its current obligations under two regimes. In the point-in-time regime (gray), a sign-off at t₀ measures the gap as near zero, after which drift events — product launches, rule changes, new jurisdictions, vendor migrations — compound unobserved until an examination at t₁ finds the firm deep in territory it last verified from a distance. In the continuous regime (red), every drift event triggers an automated re-check, and the gap saws back toward zero at the cadence of change rather than the cadence of the assessment calendar. The trajectory is stylized.

3. The audit is the out-of-sample test

Readers of this series will recognize the shape before the vocabulary arrives. Self-assessment is in-sample evaluation: the firm checks its controls against its own documentation of its own processes, and — like every backtest scored on its training data — the result is biased toward reassurance by construction. The controls were written by the people being measured; the evidence is curated by the process being evidenced. Ten notes ago we would have called the resulting green light what it is: an in-sample fit, reported without a holdout.

The examination is the holdout. A regulator samples transactions the firm did not select, reads the rulebook as written rather than as remembered, and tests the firm that exists at t₁ — not the one photographed at t₀. Firms describe the resulting findings as surprises. They are not surprises; they are the out-of-sample result, arriving on the only schedule out-of-sample results ever arrive on: after the position has been held for a while.3

The remedy is the same one this series reached for backtests, because it is the only remedy there is: move the validation to where the drift happens. A control that is checked when the process changes cannot silently decay for eighteen months, for the same reason a strategy monitored by degradation diagnostics cannot — the check rides the change, not the calendar.

4. The port

Continuous checking at combinatorial scale is not a staffing plan; no headcount curve reaches it. It is a computing plan, and its premise is older than the tooling: regulation is already code. A rulebook is conditions, thresholds, obligations, and exceptions — pseudo-code that happens to execute today on the most expensive and least deterministic runtime available: professionals reading PDFs and transcribing judgments into spreadsheets. Expressing obligations in machine-readable form is not an invention. It is a port.

The regulators know. NIST's OSCAL project exists precisely to render control catalogs and assessments in machine-readable form; the FCA and Bank of England ran multi-phase Digital Regulatory Reporting pilots on the thesis that reporting obligations could be expressed as executable rules against firms' own data.4 The supervisory side of the market is, quietly, further along in this thinking than most of the supervised side.

What the port makes possible is the division of labor this series has already field-tested in a smaller lab: agents execute the checks — continuously, against operations as they run — and escalate to humans exactly the residue that deserves humans: ambiguity, novelty, judgment. The escalation boundary is the design surface that matters, and it is the same boundary note No. 11 drew for research: the agent writes pages, not verdicts. An agent that flags a gap in a licensing obligation is a clerk with perfect stamina. An agent that decides the gap is immaterial has been promoted past the constitution.5

Under that architecture the point-in-time professional does not disappear; the work changes altitude. The consultant's product stops being a photograph of the firm and becomes an audit of the pipeline that watches the firm — which is, not coincidentally, exactly what happened to the analyst when the backtest replaced the anecdote.

5. Conclusion

Every domain this series has touched has produced the same lesson in different clothes: a measurement taken once, on a system that keeps moving, is not knowledge — it is a souvenir. Compliance is simply the domain where the souvenir is framed, signed, and billed by the hour, and where the out-of-sample test is conducted by an examiner with enforcement powers.

The firms that internalize this will stop buying green lights and start building the monitor: obligations as code, checks as infrastructure, agents as clerks, humans as judges. The rest will keep the photograph on the wall — and discover, on the regulator's schedule rather than their own, the difference between a light that is green and a light that was.


Notes

  1. The tooling market concentrated where the workflows were standardizable — identity verification and transaction screening — because standardizable workflows are what software vendors could productize. The result is the industry's current shape: automated islands of KYC and AML in a sea of manual review. The economics of that long tail, and why agentic development changes them, deserve a note of their own.
  2. Industry surveys consistently price the machinery in the tens of billions annually for financial-crime compliance alone (LexisNexis Risk Solutions, 2023), with the majority of cost in labor rather than technology — a ratio that is itself the diagnostic: labor is what you buy when your obligations are not machine-readable.
  3. The parallel extends to the failure taxonomy. An examination finding can mean the firm decayed (controls eroded), or that the environment moved (new obligations), or that the original sign-off overfit — measured the documentation rather than the operation. Note No. 10's dictum applies verbatim: one statistic, several meanings, and the remedy is instrumentation that distinguishes them before the examiner has to.
  4. NIST (2021); Financial Conduct Authority (2020). The DRR pilots are candid about the hard parts — ambiguity in rule text does not compile — which is precisely why the human escalation path is a design requirement rather than a transitional apology.
  5. The boundary also disciplines the vendor conversation. "AI compliance" that cannot state, in writing, which decisions its agents may take and which they must escalate has not designed a compliance system; it has designed a liability with a dashboard.

References

Financial Conduct Authority, 2020. Digital Regulatory Reporting: Phase 2 Viability Assessment. FCA, London.

LexisNexis Risk Solutions, 2023. True Cost of Financial Crime Compliance Study — Global Report.

NIST, 2021. Open Security Controls Assessment Language (OSCAL). National Institute of Standards and Technology, csrc.nist.gov/projects/oscal.

Keywords: financial compliance, regulation as code, continuous audit, drift, RegTech, agents.