Anti-money laundering (AML) is not a singular workload. It is seven data jobs wearing one regulatory label, and almost nobody building or buying AML technology treats them separately. If you work in financial services infrastructure, you’ve likely seen “AML” on a requirements document, a vendor slide, or a budget line. Compliance officers quote regulation at you. Vendors quote their own slide deck. Platform architects know it is expensive but not why.
This post is for the platform architects. It takes “AML” apart and shows what it actually is. Once you see the individual jobs, and the workload characteristics that fall out of them, vendor claims and architecture decisions stop being mysterious. They become consequences of the work.
Side note: If your mental model of money laundering comes from Ozark, you are closer than you think. Marty Byrde’s problem is exactly what the regulation exists to catch. The difference is scale, automation, and having to prove you caught it.
Defining AML
AML is not a singular system. It is a legal obligation. If money moves through your organisation, you must: know who is moving it, watch for patterns that suggest it is dirty, investigate when something looks wrong, report when it probably is, and prove (later) that you did all of this.
That obligation applies to banks, payment processors, insurance companies, casinos, crypto exchanges, money service businesses, and broker-dealers. The specifics vary by jurisdiction. The US has the Bank Secrecy Act. The EU has its Anti-Money Laundering Directives. The global standard-setter is FATF. The shape ends up being the same everywhere: know, watch, investigate, report, prove.

Figure 1: AML Lifecycle
There Are Seven Different Jobs Hiding Under “AML”
The legal obligations for AML produce at least seven specific data centric jobs. Each answers a different question, each comes with a different computational shape. The decomposition below is not the only way to draw the boundaries (you could split some, merge others, add a few), but it is a useful one.
- KYC / CDD. Who is this, and how risky? A data integration problem before it is an analytics one. Assembling a coherent customer picture from fragmented sources across jurisdictions is the hard part, not the scoring model.
- Sanctions screening. Are they on a prohibited list? A distinctive asymmetric shape: a tiny reference set (OFAC lists thousands of entries) scanned against every payment and customer event, with spikes on list-update days when the entire customer base may need rescreening.
- Transaction monitoring. Does this pattern look suspicious? Scenario logic run over months of transaction history to flag suspicious activity. The heaviest job in the list, and the one hiding the most complexity inside it.
- Alert triage. Which alerts deserve a human? Transaction monitoring produces far more alerts than turn out to be real (false positives). Triage scores and prioritises the queue so that the highest-risk alerts reach investigators first and the rest can be deprioritised.
- Investigation. Did suspicious activity actually occur? Interactive, bursty, and graph-shaped. Analysts follow counterparty chains several degrees out from the original alert. The access pattern is unpredictable because the investigation follows the evidence.
- Reporting and retention. Can we prove what we knew and did? Five-year retention (US) that is queryable and reconstructable, not archived. An examiner can ask you to reproduce how a decision was made.
- Model validation. Do the detection models actually work? Being reshaped by OCC Bulletin 2026-13, which narrowed the definition of “model” to exclude deterministic rule-based tools. Many transaction monitoring systems are exactly that.
That is already seven different workload shapes. But transaction monitoring hides even more variety inside it.
Defining Transaction Monitoring
Transaction monitoring is what most people picture when they hear “AML.” Unfortunately for anyone building the infrastructure, it is also not one job.
What banks call “transaction monitoring” is a bundle of technical operations running on different clocks. Understanding AML at the job level is useful. Understanding it at the component level is where you start seeing the system requirements.
| Component | Execution mode | Access pattern |
| Event intake | Continuous, event-driven | Append, sequential, write-heavy |
| Context enrichment | Continuous + micro-batch | Mixed scan and random reads |
| Detection / scoring | Real-time + micro-batch | Read-heavy, feature and history lookup |
| Historical correlation | Scheduled batch + ad hoc | Scan-heavy, long historical windows |
| Investigation / disposition | Interactive, analyst-triggered | Graph traversal, unpredictable |
| Model and rule development | Periodic batch | Mixed scan and random, compute-heavy |
| Evidence and reporting retention | Archival, append + selective retrieval | Write-once, read episodically |
There is a two-tempo problem inside transaction monitoring that has created architectural drift in the systems that support this workload. The continuous components (ingestion, enrichment, detection) are being pushed toward payment-time by real-time rails (FedNow in the US, Faster Payments in the UK, SEPA Instant in Europe). If a payment settles in seconds and cannot be recalled, overnight batch monitoring means the money is gone before the alert fires. But the batch components (historical correlation, model training, backtesting) are not moving. They cannot. They need the full lookback window. The result is a system with two operational tempos that have to share data and stay consistent.

Figure 2: Decomposing transaction monitoring
History is the working set. A detection scenario evaluating whether a sudden spike in international wire activity is suspicious cannot answer that question without seeing the previous six, twelve, or twenty-four months of customer behaviour. The data a monitoring engine needs is not today’s transactions. It is today’s transactions plus the full behavioural history needed to contextualise them. This single characteristic (history as active computation, not cold archive) shapes more AML infrastructure decisions than any other.
The false positive problem from alert triage lives here too. Alert volumes are high and false positives dominate analyst capacity. This is structural, not a tuning failure: detection scenarios are deliberately calibrated to over-alert rather than under-alert, because a missed SAR is a regulatory finding and a false positive is just analyst time.
Side note: You will see a “95% false positive rate” cited everywhere for AML transaction monitoring. The number traces back to vendor white papers and industry folklore, not to a primary regulatory or academic source. The directional claim (false positives dominate) is universally agreed. The specific percentage is not something I can stand behind with a citation.
Each job is a separate workload
Step back and look at the seven jobs. They share source data (customers, transactions, counterparties) but diverge on every operational dimension. Transaction monitoring alone, as the previous section showed, contains seven components with different execution modes. The other jobs decompose similarly.
The divergence spans every axis. Sanctions screening operates at sub-second latency against a small reference set. Transaction monitoring scans months of history in batch windows that are compressing toward real-time. Investigation is interactive and graph-shaped, running at human speed across years of customer relationships. Model validation replays the full historical record episodically. Evidence retention persists for five years and must stay queryable. No two jobs share the same profile.

Figure 3: Workload Map
A system optimised for sanctions screening (sub-second fuzzy matching at payment volume) looks nothing like a system optimised for investigation (interactive graph traversal across years of history). A platform built around a scheduled batch operates on a different cycle from one built around event-driven matching or episodic replay.
When a vendor says “AML platform,” ask which of these jobs it actually covers. You will usually find it covers three or four and punts on the rest.
Fraud and AML Convergence
Fraud and AML use the same source data. They both score risk. They both generate alerts. In many banks the same data warehouse feeds both programs.
They historically produced different systems because the obligations differ in timing, evidence, and audience.
| Fraud | AML | |
| Decision point | Inline, during the transaction | Post-event, over a lookback window |
| Latency | Milliseconds (approve/decline) | Hours to days (batch monitoring) |
| Optimises for | Speed and accuracy | Completeness and reproducibility |
| Audience | The customer and the loss book | The regulator |
| Evidence obligation | Resolve the dispute | Prove the investigation, for years |
Fraud asks: should we allow this transaction? AML asks: should we investigate this customer? Those two questions produced two separate technology stacks, two separate vendor markets, and two separate teams in most banks.
Real-time payment rails are compressing AML’s decision window toward fraud’s. If a payment settles in seconds and cannot be recalled, the post-event monitoring model that justified overnight batch starts to fail. But only the detection and screening components move toward transaction-time. Investigation, evidence retention, and model validation stay where they are. The convergence is partial. Whether the two stacks merge or stay parallel with shared infrastructure is an open question.
AML Data Architecture
If you hold the decomposition and the workload map, the infrastructure consequences follow from what the jobs require.
- History is the working set. This showed up first in transaction monitoring, but it generalises. Investigation pulls the full customer relationship. Model validation replays the full historical record. Reporting must keep years of context reconstructable. In most analytics workloads, history is a cold tier you archive and rarely touch. In AML, history is part of the computation. Three of the seven jobs cannot function without deep historical access, and two more exist specifically to prove what the historical record contains.
- What gets retained is not dead storage. It is live evidence. SAR supporting documentation carries a five-year retention obligation (US). KYC/CDD records likewise. “Retained” does not mean archived. It means queryable and reconstructable. An examiner can ask to see how a specific alert was generated, what evidence the investigator reviewed, and why the disposition was what it was. The system must be able to answer.
- Reproducibility is mandatory. A regulator can ask you to reproduce how a decision was made. If the monitoring logic has changed since the alert was generated, you still need the original version. This is not an audit trail. It is an examinable obligation.
- Copy sprawl is structural, not negligent. Most banks do not run a single integrated AML platform. They run a KYC system from one vendor, a transaction monitoring engine from another, a case management tool from a third, and a data warehouse underneath. The same customer and transaction data gets replicated into each. Every platform review calls this bad architecture. It is not. These jobs were built by different teams, bought from different vendors, regulated by requirements that arrived at different times, and run by compliance functions that could not wait for consolidation to finish before satisfying an examiner. The physical data estate for AML is a multiple of the logical data set. That multiple is the accumulated cost of building each job independently over two decades of evolving regulation. It is not a cleanup project. It is structural. Any new platform that enters this space either accepts the copy sprawl and integrates with the existing estate, or it asks the bank to rip out and replace systems that are currently under regulatory examination.
In Conclusion!
Architecture does not care what the regulator calls the program. It cares what the data actually has to do.
Next time someone puts “AML” on a requirements document, a vendor slide, or a budget line, the useful question is not “does it support AML?” It is: which job, which component, what history, what latency, what evidence obligation? The workload map is how you get there. You will be surprised how few people have an answer.
