Methodology

A screening result is only worth what you know about how it was produced. This page is the method, written so a compliance officer can check it and a developer can reproduce it. The numbers it quotes are read from the code that runs, not typed in. Where the method has limits, they are stated.

1. Sources and refresh

Five government lists, each fetched from the publisher’s own endpoint, never from an aggregator:

ListPublisherLicencePublisher cadence
OFAC Specially Designated Nationals (SDN) ListUS Department of the Treasury, Office of Foreign Assets ControlUS Government work, public domainOn designation days; typically several times a week
OFAC Consolidated Non-SDN Sanctions ListUS Department of the Treasury, Office of Foreign Assets ControlUS Government work, public domainOn designation days; less frequent than the SDN
EU Consolidated Financial Sanctions ListEuropean Commission, DG FISMAEU open data, attributionOn each Council regulation; typically weekly
UK Consolidated List of Financial Sanctions TargetsHM Treasury, Office of Financial Sanctions ImplementationOpen Government Licence v3On each designation; typically weekly
UN Security Council Consolidated ListUnited Nations Security CouncilUN public dataOn each committee decision; irregular

Every day the refresh downloads each file, records its SHA-256 and size, parses it, and swaps it in as a new batch inside one database transaction. If the parse yields fewer than half the previous batch’s entries (or under 100), the new file is rejected and the previous batch stays live: a publisher outage or a truncated download can never empty a list. Old batch rows are kept because decisions cite them. A list is reported as fresh on the status page when it was fetched successfully within 36 hours, regardless of whether the publisher changed anything.

Public chain labels (mixers, hacks, scams, darknet markets, ransomware) come from feeds whose licences allow commercial use, each row carrying its source URL. No OpenSanctions dumps, no Chainalysis or TRM data, no scraped explorer labels: the status page lists exactly what is loaded.

2. Address matching

An address is normalised by family: EVM hex and bech32 (bc1…, ltc1…) are lowercased; base58 forms (Bitcoin legacy, Solana, Tron, Monero) are kept as published because their case is significant. The chain is detected from the shape when not supplied; every EVM chain is one family, so an address listed under ETH matches on Base, Arbitrum or any other EVM network, and shapes that are ambiguous (Dogecoin, Dash, Zcash, XRP) match across chains rather than being guessed.

Matching is exact against the set of crypto addresses published on the lists above, held in process and refreshed every 60 seconds, so a direct check takes about a millisecond and does not depend on a database round-trip. A listed address is blocked under the default policy. A labelled address is handled by category: by default mixer → review, hack → review, darknet → review, ransomware → block, scam → review, sanctioned service → block. Both are policy settings you can change.

3. Name matching

Names are the hard part, and where vendors differ most. The pipeline has four stages.

Normalisation. Unicode NFKD, combining marks removed (É → E), letters NFKD cannot decompose mapped (Ø → O, ł → l, ß → ss, æ → ae), lowercased, punctuation and symbols turned into spaces, whitespace collapsed. The same function is applied to every listed name and alias at import and to every query at screen time, so “PUTIN, Vladimir Vladimirovich” and “vladimir vladimirovich putin” become the same token set.

Candidate retrieval. Every primary name and alias on every list is indexed with PostgreSQL trigrams. A query retrieves aliases whose trigram similarity to it exceeds 0.3 in either natural or token-sorted order, plus aliases that largely contain the query’s trigrams (word similarity ≥ 0.6), which is what finds a two-token query inside a four-token listed name. Up to 100 candidates per query, ordered by similarity.

Scoring. Each candidate alias is scored against the query in [0, 1]:

scoreNames(query, alias) in [0, 1]
particles are folded first: el/ul → al, "abd al" → abdal, abdul/abdel → abdal
tokens are aligned greedily, best pair first, no token reused
pair similarity: identical → 1.00
                 initial vs the word it abbreviates ("v" ↔ "vladimir") → 0.95
                 same phonetic key ("mohammed" ↔ "muhammad", "sergei" ↔ "sergey") → 0.94
                 otherwise Jaro-Winkler of the two tokens
a pair "counts" at ≥ 0.93 (one typo in a long token still counts; "vlad" vs "vladimir" does not)

if every token of the shorter name counts, and there are at least two of them:
    score = mean pair similarity − 0.08 per token the longer name has on its own
            (a dropped patronymic or an added honorific lands at about 0.92: review, not block)
    capped at 0.92 when an initial was used (an initial is never enough to block)
otherwise:
    score = 0.5 × Jaro-Winkler(token-sorted strings) + 0.5 × (Σ counting pair similarity ÷ larger token count)
            (near-misses stay visible in `matches` without reaching the review band)

a lone surname never satisfies the first branch; the best alias per entity is kept, one score per listed entity

The phonetic key folds the common Latin spellings of one sound (ph/f, kh/h, zh/j, ts/c, ck/k, q/k, x/ks, w/v, y/i, ou/u), folds a leading vowel, and drops interior vowels. It is deliberately narrow: it targets the variation between OFAC, EU, UK and UN spellings of the same person, not general phonetics. The score is the same function in the API, the dashboard and the public checker.

4. Corroboration and thresholds

Three thresholds live in your policy; the defaults are:

  • below 0.80: not reported at all;
  • 0.80 to 0.85: listed in matches for transparency, no reason, no effect on the outcome;
  • at or above 0.85: a possible match reason → review;
  • at or above 0.93: a strong match. Corroborated by a matching date of birth (year), nationality or identifier → block. Not corroborated → still review, unless your policy turns on blocking for strong uncorroborated matches.

The rule that matters: a name alone never blocks under the default policy. Thousands of people share a name with someone on a list. Blocking needs a second fact that the list also publishes about that person. Corroboration uses the identifiers the lists carry: dates of birth as published (a year match counts, since lists often give only a year or “circa”), nationalities as ISO codes, and passport, national-ID and tax numbers normalised to alphanumerics. What you send is compared against every identifier of the candidate entity, not just the alias that matched.

5. Countries

A country subject is an ISO alpha-2 code, or an IP address geolocated with DB-IP Lite (attribution required and given) with Tor exit nodes flagged. Under the default policy CU, IR, KP, SY are blocked and RU, BY, VE, MM, AF, YE, LY, SD, SO, IQ are sent to review. Both lists are yours to edit; they are a policy, not a legal determination of which jurisdictions are sanctioned.

6. Counterparty exposure

For an address or a transaction, exposure looks at the subject’s most recent 200 transactions on its chain (public RPC, or Helius and Etherscan where keys are configured), collects the unique counterparties with the value moved, and checks each against the sanctioned and labelled address sets. A direct (one-hop) counterparty on a government list is blocked by default; a labelled direct counterparty is sent to review; two-hop exposure is reported as information and allowed. Exchanges, bridges and protocols are ignored as intermediaries because everything touches them. Exposure runs inside a budget (2.5 seconds at depth 1, 8 at depth 2); if the chain data cannot be fetched in time the screen completes and the exposure check is recorded as unavailable rather than guessed. Exposure is counterparty analysis, not attribution: it does not claim who controls an address.

7. Decisions and evidence

Every screen produces one decision: allow, review or block, the list of reasons with their evidence (which list, which entry, which alias, the score, the thresholds in force, the batch id), the echoed normalised subject, every match above 0.80, and the exact batch version of every list it was evaluated against. Decisions are append-only; a review resolution is a separate record and the original is never altered. An allowlisted subject that matches is still reported with the match attached and the outcome downgraded, so the evidence shows the false positive and who cleared it. On paid plans decisions are kept for the life of the account; on Free, 30 days.

8. What it does not do

  • No politically exposed person (PEP) lists, no adverse media. Sanctions and public chain risk only.
  • No identity verification. A name screen checks a name you already trust; it does not establish who the person is.
  • No wallet attribution or clustering. Exposure checks counterparties; it does not claim that two addresses belong to the same owner.
  • No legal opinion. The attestation describes the controls in operation and never states that an organisation is compliant.
  • No guarantee the publishers’ lists are complete or correct; the crosscheck below covers our parsing of them, not their content.

9. Verification

Two checks run daily after the list refresh and their latest results are on the status page, with the date they ran.

OFAC crypto-address crosscheck. Our parse of the SDN crypto addresses is compared, asset by asset, with an independently maintained extract of the same file (0xB10C, MIT). Any address present on one side and not the other is reported as drift. Nothing is ingested from the extract; it is a second parser looking at the same source.

Name-matching benchmark. A deterministic sample of 120 listed individuals and organisations (ordered by a hash of the entity id, so the same batch gives the same sample) is screened under the default policy, with each name perturbed the ways real customer data is messy:

  • Exact name as listed
  • Different casing and punctuation
  • Tokens in a different order
  • A middle token dropped
  • Two letters transposed
  • One letter missing
  • A common transliteration variant
  • First name reduced to an initial
  • An extra token added

Recall is the share of queries for which the listed entity comes back at or above the review threshold. Alongside it, 60 ordinary personal and business names are screened and the share that would reach a review queue is published as the noise figure. Both numbers are reported as they come out; the benchmark is in the repository and runs with npm run verify.