At first glance, The Flower Shop appears to contain thousands of routine money transfers. The investigation begins by asking what normal activity looks like, then follows the transactions that repeatedly break from that pattern. Amount, frequency, timing, identification and geography are examined first. The analysis then connects people and physical agents through a focused network, allowing individual warning signs to be viewed as part of a larger movement of funds.
Transactions reviewed 11,399
Observed period Jan 2001 to Nov 2020
Leading person suspect name_10 (6/8)
Leading agent suspect agent_39 (6/6)
The analysis produced a shortlist of three people and three physical agents for enhanced investigation. The leading people are name_10, name_1119, name_143. The leading physical agents are agent_39, agent_73, agent_246. They were not selected because of one large payment or one fast transfer. They rose to the top because several independent parts of the story pointed in the same direction, including repeated activity, unusual values, immediate onward movement and concentrated network relationships.
The strongest person-level agent-concentration signal appears for name_10, while agent_39 receives the highest agent Evidence Score. The person combines unusually high activity with repeated large or recordkeeping-adjacent transactions, concentrated use of one physical agent and substantial operator diversity. The leading agent combines missing identification, upper-tail transaction values, immediate-relay connections and repeated transactions where several indicators overlap.
| Rank | Person | Score | Transactions | Total value | Immediate relays | Large amounts | Recordkeeping adjacent | Primary agent |
|---|---|---|---|---|---|---|---|---|
| 1 | name_10 | 6/8 | 301 | $370,671 | 0 | 72 | 20 | agent_39 |
| 2 | name_1119 | 6/8 | 37 | $50,557 | 2 | 12 | 5 | agent_11 |
| 3 | name_143 | 6/8 | 28 | $28,545 | 2 | 4 | 0 | agent_11 |
| Rank | Agent | Location | Score | Transactions | Total value | Missing ID | Large amount rate | Immediate relays |
|---|---|---|---|---|---|---|---|---|
| 1 | agent_39 | SAINT LOUIS PARK, MN | 6/6 | 491 | $603,306 | 87.0% | 24.0% | 67 |
| 2 | agent_73 | MONTERREY, NL | 4/6 | 182 | $225,429 | 100.0% | 30.8% | 1 |
| 3 | agent_246 | SAINT LOUIS PARK, MN | 4/6 | 18 | $10,090 | 88.9% | 0.0% | 8 |
Decision boundary: The shortlist is a prioritisation result. It does not establish criminal intent, beneficial ownership, a predicate offence, or control of the funds. The next decision should be whether transaction records, customer relationships, agent controls, and source-of-funds evidence provide a credible lawful explanation.
What is Missing_ID?: Missing ID refers to transactions where the customer identification field required for the analysis (for example, a government-issued identification number or equivalent customer identifier) is absent or unavailable. Missing identification is not inherently suspicious and may arise for legitimate operational reasons. However, consistently high rates of missing identification reduce traceability and may weaken customer due diligence, particularly when observed alongside other indicators such as unusually large values, immediate movement of funds, recordkeeping-adjacent transfers or concentrated transaction patterns. Throughout this report, missing ID is treated as a supporting risk indicator rather than evidence of money laundering on its own.
What is Recordkeeping adjacent?: Recordkeeping-adjacent refers to transactions valued from $2,400 to below $3,000. This range was selected because it sits close to the $3,000 external recordkeeping benchmark used in the analysis. A transaction in this range is not automatically suspicious and may reflect an ordinary customer payment. The concern begins when the same people, agents or connected groups repeatedly operate just below the benchmark, particularly when the activity also involves rapid onward movement, missing identification or unusual network relationships. Repetition can suggest that transaction amounts are being deliberately managed to remain below a known control point. Throughout this report, recordkeeping-adjacent activity is treated as a supporting indicator rather than proof of threshold avoidance or money laundering. Its value comes from the wider pattern around the transaction, not from the amount alone.
At first glance, the transactions passing through The Flower Shop look like ordinary money transfers. People send funds, recipients collect them, and agents process each payment. The concern begins when the same behaviour appears again and again. Several people may send money to one recipient, who then forwards most of it shortly afterwards. Certain agents may repeatedly handle transactions with missing identification, similar amounts, unusual locations, or rapid onward movement. No single transaction proves that money laundering has occurred. The stronger story emerges when separate warning signs begin to connect and point towards the same people and locations. This analysis follows those connections, tests whether the patterns are unusual, and narrows thousands of transactions into a focused shortlist of individuals and agents whose activity deserves closer investigation.
The guiding question is:
Which people and physical agents demonstrate the strongest combination of unusual transaction volume, rapid movement, managed amounts, weak identification, or network coordination?
Extract and validate
↓
Understand ordinary activity
↓
State hypotheses and decision rules
↓
Test transaction and network patterns
↓
Score evidence without false precision
↓
Present the investigative shortlist
The objective of this investigation is not simply to identify unusual transactions, but to determine which people, locations or transaction patterns warrant further investigation. The analysis therefore follows the same process an investigator would typically use when reviewing a large transaction dataset.
The investigation begins by understanding what normal activity looks like inside The Flower Shop. Transaction amounts, frequency, timing, locations and agent usage are reviewed to establish a clear baseline. Once that picture is in place, the analysis looks for behaviour that breaks from the norm. Each suspicious pattern is treated as a separate hypothesis, such as rapid movement of funds, repeated near-threshold transfers, concentrated agent activity or money flowing through connected individuals. These hypotheses are tested one at a time before the results are brought together. The aim is not to label anyone as guilty. It is to identify the people, agents and transactions where several warning signs overlap, creating a focused shortlist for deeper investigation.
The cleaned file is treated as the analytical starting point. The workflow removes any residual row-count field, standardises column names, parses dates, converts amount fields, and constructs one transaction identifier per row. Routine preparation code is intentionally hidden because it does not contribute to the investigative story.
| Measure | Value |
|---|---|
| Transactions | 11399 |
| Exact duplicate rows | 0 |
| Missing send timestamps | 0 |
| Missing pay timestamps | 11 |
| Non-positive amounts | 0 |
| Pay before send | 422 |
| Unique senders | 3019 |
| Unique payees | 4658 |
Before looking for suspicious behaviour, the file is checked to make sure the investigation starts on reliable ground. Duplicate records, zero or negative amounts, missing dates and transactions where the payment appears to occur before the send time are all reviewed. These records are not quietly deleted, because each one may reveal either a data-quality issue or an unusual transaction that deserves attention. Instead, they are kept and clearly flagged. This creates a transparent record of what was found, what was retained and how the data was prepared before any conclusions were drawn.
Every investigation begins with a few necessary assumptions about what the data represents. A name is assumed to refer to the same person across transactions, timestamps are assumed to record the true sequence of events, and transaction amounts are assumed to be complete and comparable. The analysis also assumes that the dataset captures enough of the activity passing through The Flower Shop to reveal meaningful connections. These assumptions are made visible because they shape every finding that follows. When an assumption is uncertain, the related result is treated with greater caution. The final shortlist is therefore only as reliable as the identities, dates, transaction values and network coverage available in the file.
| Assumption | Why was the assumption made | Potential pitfalls |
|---|---|---|
| An exact anonymised name consistently represents the same person. | The dataset does not contain a usable identification number, so names are the provisional entity key. | One person may be split across spellings, or different people may be incorrectly combined. |
| Addresses support context but do not automatically merge people. | Households, businesses, apartment buildings, and data-entry defaults can legitimately share addresses. | Some hidden relationships may remain unidentified, but false entity merges are reduced. |
| Each row represents one completed transfer. | The file does not include a status field for cancelled, declined, refunded, or reversed transfers. | Counts and value totals would be overstated if unsuccessful transfers are present. |
| The amount field is comparable across the dataset. | No currency field or exchange rate is provided at the transaction level. | Cross-country value comparisons may be distorted if amounts use different currencies. |
| Send and pay timestamps use a common time standard. | Timing tests require direct comparison of the two timestamps. | Settlement and relay intervals may be incorrect if time zones differ. |
| Pay DateTime represents actual in-person receipt. | The assignment describes the receiver as collecting the transfer in person. | A processing timestamp would weaken the interpretation of corridor-adjusted settlement and receive-to-send relay. |
| The file captures only The Flower Shop channel. | The dataset cannot observe cash, bank accounts, goods, digital assets, or other operators outside the file. | Acyclic or incomplete transaction paths may still be part of a broader laundering network. |
| Agent names and cities identify stable physical locations. | Agent analysis depends on grouping repeated activity at the same location. | Spelling variants or moved locations may split one agent into several records. |
| Binary ID fields mean identification was or was not indicated. | The fields do not contain document numbers or validation quality. | Missing ID may reflect collection policy, system migration, or incomplete records rather than avoidance. |
| An evidence score is a screening tool, not a finding of guilt. | The dataset lacks occupation, source of funds, relationship purpose, and verified beneficial ownership. | Every shortlisted entity requires enhanced review and a lawful-explanation test. |
The analysis does not rely on hidden assumptions or a complex model that cannot be explained. Every screening decision is stated clearly, including the thresholds used, the time windows selected, the behaviours treated as unusual and the reasons behind each choice. This makes the process easy to follow and allows another analyst to repeat the same steps using the same data. It also makes the findings open to challenge. A different threshold or time window may produce a different shortlist, and that is important to acknowledge. The goal is not to present the score as unquestionable, but to show exactly how each conclusion was reached.
| Decision | Rule used | Why was the decision made? |
|---|---|---|
| Large transaction | At least $1,400 | The cutoff is the dataset’s 90% percentile. It identifies transactions that are unusual for The Flower Shop without claiming that the value is suspicious by itself. |
| Very large transaction | At least $2,000 | The cutoff is the dataset’s 95% percentile. It separates the most extreme values for additional context while leaving the broader large-transaction screen intact. |
| Recordkeeping reference | $3,000 | The value is retained as a specific external reference for funds-transfer recordkeeping. It is not used as the definition of a large Flower Shop transaction. |
| Recordkeeping-adjacent range | $2,400 to below $3,000 | The lower bound is 80% of the recordkeeping reference. This range tests whether values repeatedly cluster immediately below the reference. |
| Fast settlement | Fastest 10% within the same send-country and pay-country corridor | Settlement speed differs materially by corridor. A peer-relative rule identifies unusually fast collection without treating every transfer as though it follows the same operating pattern. |
| Timing peer minimum | At least 50 corridor transactions | Smaller corridors can produce unstable percentiles. When the peer group is smaller, the overall 10% cutoff of 2 minutes is used. |
| Immediate relay | At most 120 minutes | The receive-to-send distribution contains a distinct immediate cluster. The two-hour rule captures that cluster while excluding overnight activity. |
| Short-horizon relay | At most 24 hours | The one-day measure is retained as supporting context and sensitivity analysis. It does not drive the primary relay point in the Evidence Score. |
| Relay search horizon | At most 7 days | A seven-day search captures delayed forwarding without linking transactions separated by long periods. |
| Person outlier cutoff | Top 5% | The 95th percentile focuses on the extreme tail while retaining enough people for comparison and investigation. |
| Agent eligibility | At least 10 transactions | A minimum denominator reduces misleading 100% rates produced by one or two transactions. |
| Agent outlier cutoff | Top 10% | The broader agent cutoff reflects a larger and more fragmented physical-agent population than the person shortlist. |
| Repeated evidence minimum | At least two events | One event can be accidental or operational. Repetition is required for immediate-relay and multi-signal points. |
Important interpretation: Two different questions require two different amount rules. A transaction is considered unusually large for The Flower Shop when it falls at or above the dataset’s 90% percentile, equal to $1,400. The separate $3,000 value is retained only as an external recordkeeping reference and for the recordkeeping-adjacent structuring test. Neither rule is evidence of wrongdoing on its own.
Before asking who appears suspicious, the investigation first needs to understand what The Flower Shop considers routine. This section establishes the normal range of transaction values, the people who appear most often, the speed at which transfers are collected and the physical agents through which the activity moves. Each chart narrows the field, but no chart is allowed to decide the case by itself. The story becomes useful only when the same names and locations continue to reappear across different views of the data.
The amount distribution is right-skewed. The median is $833, the mean is $909.94, the standard deviation is $515.88, and the 95th percentile is $2,000. The mean sits above the median because a smaller upper tail pulls the average upward. That upper tail matters, but the dataset also makes clear that a single universal dollar rule would miss much of the relevant behaviour.
Why use thresholds at all?
Thresholds turn broad phrases such as “unusually large” into decisions that another analyst can see, reproduce and challenge. They serve four purposes. They reduce thousands of records to a manageable review population. They make the reason for selection visible. They allow the same rule to be repeated. They also create a specific test for values that repeatedly sit near an external control point.
The weakness is equally important. A transaction immediately below a cutoff is not fundamentally different from one immediately above it. For that reason, this report does not use one number for every purpose.
Transaction size is evaluated through separate lenses. The dataset’s 90th percentile, equal to $1,400, identifies transactions that are unusually large for The Flower Shop. The 95th percentile, equal to $2,000, identifies the most extreme values. The separate $3,000 benchmark is retained only as a recordkeeping reference. Transactions from $2,400 to below that reference are tested for repeated threshold-adjacent behaviour. These rules do not establish suspicious conduct. They become useful when unusual values overlap with repeated activity, immediate onward movement, missing identification or network relationships.
| Statistic | Amount |
|---|---|
| Minimum | $30.00 |
| 25th percentile | $589.77 |
| Median | $833.00 |
| 75th percentile | $1,000.00 |
| 90th percentile | $1,400.00 |
| 95th percentile | $2,000.00 |
| 99th percentile | $2,900.00 |
| Maximum | $9,000.00 |
The comparison shows that a threshold is only useful when it captures a meaningful pattern, not simply a large number of transactions.
Using $2,000 as the reference identifies 190 transactions between $1,600 and $2,000. This is a relatively narrow group, but it does not address the specific recordkeeping question the analysis is trying to test.
The $2,500 reference captures 362 transactions, the largest group in the comparison. However, 70.7% of them are exactly $2,000. This suggests that the window is being driven by a common transaction amount rather than people deliberately operating just below $2,500.
The $3,000 reference captures 167 transactions, or 1.5% of the dataset. Exact $2,500 transfers still appear frequently, but they account for a smaller 43.7% of the window. This leaves a more varied and selective set of transactions to examine for repeated threshold-adjacent behaviour.
The $5,000 reference captures only 21 transactions. That is too small a group to support a reliable recurring-pattern test and risks placing too much weight on isolated high-value transfers.
| Candidate_threshold | Near_window | Transactions | Share_of_dataset | Most_common_amount | Dominant_amount_share |
|---|---|---|---|---|---|
| $2,000 | $1,600 to below $2,000 | 190 | 1.7% | $1,800 | 17.4% |
| $2,500 | $2,000 to below $2,500 | 362 | 3.2% | $2,000 | 70.7% |
| $3,000 | $2,400 to below $3,000 | 167 | 1.5% | $2,500 | 43.7% |
| $5,000 | $4,000 to below $5,000 | 21 | 0.2% | $4,000 | 47.6% |
The conclusion is not that $3,000 defines a large Flower Shop transaction. Large transactions are identified separately using the dataset’s 90th and 95th percentiles. The $3,000 reference is retained for a narrower purpose: to test whether the same people, agents or connected groups repeatedly send amounts just below an external recordkeeping benchmark, and the $3,000 reference creates a more selective range from $2,400 to below $3,000. The important reason it remains in the analysis because it answers a narrower question: do people repeatedly select values immediately below an external recordkeeping point?
The selected threshold code is available below for review.
threshold_options <- sort(
unique(c(2000, 2500, params$recordkeeping_threshold, 5000))
)
threshold_comparison <- map_dfr(
threshold_options,
function(candidate_threshold) {
lower_bound <-
candidate_threshold *
params$recordkeeping_lower_ratio
values <- tx$amount[
!is.na(tx$amount) &
tx$amount >= lower_bound &
tx$amount < candidate_threshold
]
dominant_amount <- if (length(values) == 0) {
NA_real_
} else {
as.numeric(names(sort(table(values), decreasing = TRUE))[1])
}
dominant_share <- if (length(values) == 0) {
NA_real_
} else {
max(table(values)) / length(values)
}
tibble(
Candidate_threshold = candidate_threshold,
Near_window = glue(
"{dollar(lower_bound)} to below {dollar(candidate_threshold)}"
),
Transactions = length(values),
Share_of_dataset = length(values) / nrow(tx),
Most_common_amount = dominant_amount,
Dominant_amount_share = dominant_share
)
}
)
amount_quantiles <- quantile(
tx$amount,
probs = c(0, 0.25, 0.50, 0.75, 0.90, 0.95, 0.99, 1),
na.rm = TRUE
)
amount_summary <- tibble(
Statistic = c(
"Minimum", "25th percentile", "Median", "75th percentile",
"90th percentile", "95th percentile", "99th percentile", "Maximum"
),
Amount = as.numeric(amount_quantiles)
)Amount narrows the field, but it cannot distinguish a legitimate high-value customer from a person repeatedly moving funds through the system. The next question is therefore not simply who moved the most money, but who kept returning.
A single transaction rarely tells the whole story. Some people appear only once, while others repeatedly move money through The Flower Shop as both senders and recipients. The analysis therefore measures frequency across both roles to identify the individuals who return to the network again and again. This also prevents one unusually large transfer from pushing someone to the top of the shortlist without supporting behaviour. The stronger concern is not simply how much money a person moves, but how often they appear, which roles they play and whether those repeated transactions connect to other warning signs.
The x-axis shows how often each person appears in the dataset, counting activity as both a sender and a recipient. Moving further right means the person is involved in more transactions. The y-axis shows the total value of money they sent and received, so higher points represent greater financial activity. Both axes use a log scale, which keeps occasional users and extreme outliers visible on the same chart without allowing a few very large values to overwhelm the overall pattern.
Frequency shows who repeatedly uses the channel. It still does not explain whether the movement is operationally normal. To answer that, the investigation turns to time. Was the money collected unusually quickly for its corridor, and did the recipient then become a sender almost immediately?
Two timing measures are kept separate because they describe different parts of the story. Settlement time measures the interval from send to pay. It can reflect product speed, corridor operations and customer collection behaviour. Relay time begins only after the money is collected and measures how long it takes the same person to send another transfer. That second interval is more directly relevant to possible pass-through activity.
The overall settlement distribution is highly dispersed. The 10th percentile is only 2 minutes, while the median is more than 17 hours. The mean and standard deviation are pulled upward by an extreme upper tail, including a maximum interval of more than 988,000 minutes. This is why the median and corridor percentiles carry more weight than the overall mean.
| Measure | Settlement time |
|---|---|
| 5th percentile | 1 minutes |
| 10th percentile | 2 minutes |
| 25th percentile | 79 minutes |
| Median | 1,147 minutes, or 19.1 hours |
| 75th percentile | 40,315 minutes, or 671.9 hours |
| Mean | 214,565 minutes, or 3576.1 hours |
| Standard deviation | 1,105,173 minutes |
| Maximum | 9,545,245 minutes |
A fixed 60-minute rule would classify 23.0% of valid transactions as fast. The difficulty is that 60 minutes does not mean the same thing everywhere. A domestic United States transfer can settle in minutes, while the fastest transactions in another corridor may still take close to two hours.
| Corridor | 10th-percentile settlement |
|---|---|
| United States to United States | 0 minutes |
| United States to Guatemala | 2 minutes |
| United States to El Salvador | 3 minutes |
| United States to Mexico | 3 minutes |
| United States to Colombia | 48 minutes |
For that reason, a transaction is classified as fast settlement when it falls within the fastest 10% of its send-country and pay-country corridor. The corridor-specific rule is used only when at least 50 comparable transactions exist. Smaller corridors use the overall fallback of 2 minutes. This keeps the rule sensitive to the way the product actually operates rather than importing a universal one-hour standard.
1,077 transactions fall within the fastest corridor-adjusted settlement group. This remains a supporting indicator because wire-transfer products are designed for speed. The stronger timing signal is the identification of 77 receipts followed by a new send from the same person within 120 minutes. A broader 24-hour measure captures 83 transactions and is retained only as supporting context.
Timing reveals how quickly value moves. The next question is where that movement is being facilitated. If one location dominates the file, is it simply the main Flower Shop outlet, or does its activity also contain the strongest concentration of independent warning signs?
One physical send agent, agent_11, processed 91.8% of send transactions. This is an immediate operational finding because activity is concentrated in one location. It is not, by itself, an AML conclusion. A dominant agent may simply be The Flower Shop’s principal business channel.
The busiest agent becomes an investigative concern only when its volume overlaps with missing identification, an unusually high share of large transactions, immediate-relay connections or concentration around one person. This prevents the scoring process from treating business importance as evidence of wrongdoing. With the baseline now established, the investigation can move from observation to hypothesis: which repeated patterns would be difficult to explain as ordinary remittance activity?
The exploratory analysis identifies where activity is concentrated, but concentration alone does not explain intent. The next step is to translate recognised laundering patterns into rules that the dataset can either support or reject. Each hypothesis is tested independently across amount, timing, geography and network relationships. A zero result is retained because the purpose is to learn what the data shows, not to force every theory to produce a suspect.
| Hypothesis | Why investigate this pattern | Dataset Implementation |
|---|---|---|
| Agent facilitation | A particular location may repeatedly appear in transactions containing unrelated warning signs. When this occurs across several people, it may indicate weak controls or coordinated routing rather than one unusual customer. | Compare transaction volume, missing-ID rate, large-transaction rate, immediate-relay involvement and concentration around one person. |
| Many senders to one payee followed by relay | A recipient may act as a temporary collection point by receiving funds from several people and sending value onward before there is time for ordinary use or retention. | Require at least three unique senders within one month and at least one outgoing transfer within 120 minutes after receipt. |
| Near-threshold transfers to a common recipient | Repeated values immediately below an external recordkeeping point may indicate deliberate amount selection, particularly when several senders fund the same recipient. | Require at least two transfers from $2,400 to below $3,000 from at least two senders in one month. |
| Split and recombine | Dividing funds across several intermediaries and later directing them to one recipient can make the original path harder to follow. | Test the full three-intermediary motif and retain a two-intermediary path only as partial screening evidence. |
| Geographic funnel | Funds arriving from several locations and moving onward quickly may indicate collection activity across a wider network rather than an ordinary relationship between one sender and one recipient. | Require at least two origin states in one month and at least one onward transfer within 120 minutes. |
| Role switching and network prominence | People who repeatedly alternate between sending and receiving, while connecting many parts of the network, may be acting as intermediaries rather than final users. | Measure repeated two-sided activity and weighted degree in the person-to-person network. |
The hypotheses now define what the investigation is looking for. The next section asks a harder question: which patterns actually survive contact with the data?
The purpose of the tests is not to produce the largest possible alert count. It is to determine which proposed patterns are visible under rules that were stated before the shortlist was assembled. Some hypotheses are expected to produce candidates, while others may fail under the strict definition. Both outcomes help narrow the investigation.
| Hypothesis | Test | Result | Interpretation |
|---|---|---|---|
| Fast send-to-pay interval | Settlement in the fastest 10% of comparable corridors; 2 minutes used as the fallback | 1077 | Observed transactions |
| Immediate payee-to-next-send relay | Next send within 120 minutes | 77 | Across 74 people |
| Many senders to one payee followed by relay | At least three senders in one month plus at least one immediate relay | 0 | Not supported under the stated rule |
| Near-threshold transfers to a common recipient | At least two transfers from $2,400 to below $3,000 from at least two senders in one month | 1 | Supported for screening |
| Split-and-recombine network path | At least three intermediaries fund one later recipient within 7 days; two intermediaries retained as a partial screening pattern | 0 | 1 partial pattern(s) with at least two intermediaries |
| Multiple origin states followed by redistribution | At least two origin states in one month plus at least one immediate relay | 0 | Not supported under the stated rule |
The findings are intentionally mixed. Fast corridor-adjusted settlement and immediate relay are both present. Recordkeeping-adjacent convergence appears in a limited number of monthly windows. The strict funnel and geographic rules may return few or no candidates. The full split-and-recombine rule is also retained even when only a partial path is observed. These negative findings are not discarded. They show which theories were tested and where the available file did not provide enough support.
The summary identifies the patterns that deserve attention, but it does not yet show the people behind them. The investigation therefore moves from the hypothesis level to the transaction sequences that create each result.
| Person | Immediate relays | Median minutes | Minimum minutes | Amount received | Amount sent next (onward transaction) |
|---|---|---|---|---|---|
| name_1119 | 2 | 0.1 | 0.1 | $5,600.00 | $5,600 |
| name_143 | 2 | 0.1 | 0.1 | $3,099.74 | $1,980 |
| name_392 | 2 | 2.5 | 0.1 | $1,509.99 | $1,400 |
| name_134 | 1 | 0.0 | 0.0 | $629.99 | $620 |
| name_1471 | 1 | 0.0 | 0.0 | $509.99 | $500 |
| name_2095 | 1 | 0.0 | 0.0 | $809.99 | $800 |
| name_2314 | 1 | 0.0 | 0.0 | $1,021.00 | $1,000 |
| name_2346 | 1 | 0.0 | 0.0 | $500.00 | $500 |
An immediate relay does not require the incoming and outgoing amounts to match exactly. Transfer fees, partial forwarding, pooled funds and activity outside the dataset can create differences between the two values. The test asks a simpler and more defensible question: does the same person repeatedly collect funds and become a sender again within 120 minutes?
That timing pattern is more meaningful when it is joined by deliberate-looking amount selection. The next test therefore asks whether several senders repeatedly converge on the same recipient using values immediately below the recordkeeping reference.
| Person | Month | Near threshold transactions | Near threshold amount | Unique senders | Unique send agents | Missing sender ID |
|---|---|---|---|---|---|---|
| name_6330 | 2019-06 | 2 | $5,000 | 2 | 1 | 0.0% |
The structuring test remains intentionally separate from the general large-transaction screen. A transaction at or above $1,400 is unusual for The Flower Shop. A transaction between $2,400 and below $3,000 answers a different question: does the value repeatedly sit immediately below an external reference point? Keeping the tests separate prevents one threshold from being used to explain two different behaviours.
The full hypothesis requires at least three intermediaries. A two-intermediary path is retained separately as partial evidence because the dataset may not capture every operator or channel. The partial rule is visible to the investigator, but it does not receive the same narrative weight as a complete three-intermediary motif.
| Origin person | Month | Later recipient | Intermediaries | Split transactions | Recombine transactions | Split value | Recombined value | Path |
|---|---|---|---|---|---|---|---|---|
| name_10 | 2020-09 | name_2401 | 2 | 2 | 2 | $2,019.98 | $4,000 | name_2816; name_2911 |
The selected hypothesis logic is available below.
monthly_receipts <- tx |>
filter(!is.na(payee_name), !is.na(pay_month)) |>
group_by(person_id = payee_name, analysis_month = pay_month) |>
summarise(
incoming_transaction_count = n(),
incoming_amount = sum(amount, na.rm = TRUE),
unique_senders = n_distinct(sender_name, na.rm = TRUE),
unique_origin_states = n_distinct(sender_state, na.rm = TRUE),
unique_send_agents = n_distinct(send_agent_name, na.rm = TRUE),
rapid_relay_count = sum(rapid_relay, na.rm = TRUE),
near_threshold_count = sum(near_threshold, na.rm = TRUE),
.groups = "drop"
)
funnel_candidates <- monthly_receipts |>
filter(unique_senders >= 3, rapid_relay_count >= 1)
geographic_candidates <- monthly_receipts |>
filter(unique_origin_states >= 2, rapid_relay_count >= 1)
structured_candidates <- tx |>
filter(near_threshold, !is.na(payee_name), !is.na(pay_month)) |>
group_by(person_id = payee_name, analysis_month = pay_month) |>
summarise(
near_threshold_count = n(),
near_threshold_amount = sum(amount, na.rm = TRUE),
unique_senders = n_distinct(sender_name, na.rm = TRUE),
unique_send_agents = n_distinct(send_agent_name, na.rm = TRUE),
sender_id_missing_rate = mean(sender_id_missing, na.rm = TRUE),
.groups = "drop"
) |>
filter(near_threshold_count >= 2, unique_senders >= 2)
split_receipts <- tx |>
filter(!is.na(sender_name), !is.na(payee_name), !is.na(pay_dt)) |>
transmute(
origin_id = sender_name,
intermediary_id = payee_name,
split_transaction_id = transaction_id,
split_amount = amount,
intermediary_received_dt = pay_dt,
origin_month = floor_date(pay_dt, "month")
)
split_outgoing <- tx |>
filter(!is.na(sender_name), !is.na(payee_name), !is.na(send_dt)) |>
transmute(
intermediary_id = sender_name,
recombine_recipient_id = payee_name,
recombine_transaction_id = transaction_id,
recombine_amount = amount,
recombine_send_dt = send_dt
)
split_paths <- suppressWarnings(
inner_join(split_receipts, split_outgoing, by = "intermediary_id")
) |>
filter(
recombine_send_dt >= intermediary_received_dt,
recombine_send_dt <= intermediary_received_dt + days(params$relay_search_days)
)
split_motifs <- split_paths |>
group_by(origin_id, origin_month, recombine_recipient_id) |>
summarise(
intermediary_count = n_distinct(intermediary_id),
split_transaction_count = n_distinct(split_transaction_id),
recombine_transaction_count = n_distinct(recombine_transaction_id),
split_total = sum(split_amount, na.rm = TRUE),
recombined_total = sum(recombine_amount, na.rm = TRUE),
intermediaries = collapse_values(intermediary_id),
.groups = "drop"
) |>
mutate(
full_split_recombine = intermediary_count >= 3,
partial_split_recombine = intermediary_count >= 2
)
full_split_candidates <- split_motifs |>
filter(full_split_recombine)
partial_split_candidates <- split_motifs |>
filter(partial_split_recombine)The hypothesis tests produce several types of evidence, but they do not naturally rank one person against another. The next step is therefore to combine the findings without pretending that the dataset can estimate a probability of guilt.
The report uses an Evidence Score, not a model probability. Every condition contributes one point. Equal weights are deliberate because the dataset contains no confirmed laundering cases that could justify statistical weighting. A score of six does not mean a 60% probability of money laundering. It means that six independently stated screening conditions are present.
This simple structure also keeps the reasoning visible. A reader can see which point was earned, why the condition matters and which transaction behaviour produced it. The score is therefore a prioritisation tool. Its purpose is to identify where limited investigative time should be spent first.
A person can score from 0 to 8. An agent can score from 0 to 6. Scores are compared only within the same entity type.
| Evidence component | Point | Why does this matter? |
|---|---|---|
| transaction frequency in the top 5% | 1 | Repeated activity increases the opportunity for coordinated movement and distinguishes persistent behaviour from an isolated transfer. |
| combined sent and received value in the top 5% | 1 | Large aggregate value identifies people who control or touch a material share of the observed flow. |
| at least two receipts followed by a new send within 120 minutes | 1 | Repeated receive-to-send movement within 120 minutes is more consistent with immediate pass-through behaviour than ordinary retention. |
| at least two unusually large or recordkeeping-adjacent transactions | 1 | Repeated transactions at or above the dataset’s 90% percentile, or immediately below the $3,000 reference, show that the person repeatedly operates in an unusual part of the amount distribution. |
| repeated activity as both sender and payee | 1 | Repeated movement between sender and payee roles can identify intermediaries who collect and redistribute funds. |
| fan-in activity or a partial split-and-recombine path | 1 | Fan-in and split paths identify people positioned inside a coordinated flow rather than at a single end point. |
| high concentration at one agent with repeated missing identification and operator diversity | 1 | A concentrated relationship with one physical agent becomes more material when identification is repeatedly absent and many operators are used. |
| weighted network degree in the top 5% | 1 | Weighted degree identifies actors connected to an unusually large number of transactions in the person-to-person network. |
The top 5% cutoffs are dataset-relative. In this file, the person frequency cutoff is 11 transactions and the person value cutoff is $10,328.00. The amount-management point is also dataset-aware. It is earned when a person is linked to at least two transactions at or above $1,400, or at least two transactions between $2,400 and below $3,000. The first condition captures unusual value for The Flower Shop. The second captures repeated activity next to a specific recordkeeping reference.
The immediate-relay point requires repetition. One receive-to-send event can arise by chance or from an ordinary remittance purpose. Requiring at least two events within 120 minutes makes the point reflect a pattern rather than an isolated timestamp.
The score construction code is shown below.
sender_summary <- tx |>
group_by(person_id = sender_name) |>
summarise(
sent_count = n(),
total_sent = sum(amount, na.rm = TRUE),
unique_payees = n_distinct(payee_name, na.rm = TRUE),
large_sent_count = sum(large_amount, na.rm = TRUE),
near_sent_count = sum(near_threshold, na.rm = TRUE),
first_send = min(send_dt, na.rm = TRUE),
last_send = max(send_dt, na.rm = TRUE),
.groups = "drop"
)
receiver_summary <- tx |>
group_by(person_id = payee_name) |>
summarise(
received_count = n(),
total_received = sum(amount, na.rm = TRUE),
unique_senders = n_distinct(sender_name, na.rm = TRUE),
unique_origin_states = n_distinct(sender_state, na.rm = TRUE),
large_received_count = sum(large_amount, na.rm = TRUE),
near_received_count = sum(near_threshold, na.rm = TRUE),
first_receipt = min(pay_dt, na.rm = TRUE),
last_receipt = max(pay_dt, na.rm = TRUE),
.groups = "drop"
)
relay_person_summary <- relay_first |>
filter(rapid_relay) |>
group_by(person_id) |>
summarise(
rapid_relay_count = n(),
median_relay_minutes = median(relay_minutes, na.rm = TRUE),
minimum_relay_minutes = min(relay_minutes, na.rm = TRUE),
relayed_received_amount = sum(received_amount, na.rm = TRUE),
relayed_outgoing_amount = sum(next_send_amount, na.rm = TRUE),
.groups = "drop"
)
person_edges <- tx |>
filter(!is.na(sender_name), !is.na(payee_name)) |>
group_by(from = sender_name, to = payee_name) |>
summarise(
frequency = n(),
total_amount = sum(amount, na.rm = TRUE),
.groups = "drop"
)
person_graph <- igraph::graph_from_data_frame(
person_edges,
directed = TRUE
)
person_network <- tibble(
person_id = igraph::V(person_graph)$name,
weighted_degree = igraph::strength(
person_graph,
mode = "all",
weights = igraph::E(person_graph)$frequency
)
)
sender_agent_profile <- tx |>
filter(!is.na(sender_name), !is.na(send_agent_name)) |>
group_by(
person_id = sender_name,
dominant_agent = send_agent_name,
dominant_agent_city = send_agent_city,
dominant_agent_state = send_agent_state
) |>
summarise(
dominant_agent_transactions = n(),
dominant_agent_amount = sum(amount, na.rm = TRUE),
dominant_agent_missing_id_rate = mean(
sender_id_missing,
na.rm = TRUE
),
dominant_agent_operator_count = n_distinct(
send_operator_name,
na.rm = TRUE
),
dominant_agent_unique_payees = n_distinct(
payee_name,
na.rm = TRUE
),
.groups = "drop"
) |>
left_join(
tx |>
count(sender_name, name = "all_sent_transactions") |>
rename(person_id = sender_name),
by = "person_id"
) |>
mutate(
agent_share =
dominant_agent_transactions /
all_sent_transactions
) |>
group_by(person_id) |>
slice_max(
order_by = agent_share,
n = 1,
with_ties = FALSE
) |>
ungroup()
partial_split_origins <- unique(
partial_split_candidates$origin_id
)
person_summary <- full_join(
sender_summary,
receiver_summary,
by = "person_id"
) |>
full_join(relay_person_summary, by = "person_id") |>
full_join(person_network, by = "person_id") |>
left_join(sender_agent_profile, by = "person_id") |>
mutate(
across(
c(
sent_count, total_sent, unique_payees,
large_sent_count, near_sent_count,
received_count, total_received, unique_senders,
unique_origin_states, large_received_count,
near_received_count, rapid_relay_count,
relayed_received_amount, relayed_outgoing_amount,
weighted_degree, dominant_agent_transactions,
dominant_agent_amount,
dominant_agent_missing_id_rate,
dominant_agent_operator_count,
dominant_agent_unique_payees,
all_sent_transactions, agent_share
),
~ replace_na(.x, 0)
),
transaction_count = sent_count + received_count,
total_value = total_sent + total_received,
large_amount_count =
large_sent_count + large_received_count,
near_threshold_count =
near_sent_count + near_received_count
)
person_frequency_cutoff <- quantile(
person_summary$transaction_count,
probs = 0.95,
na.rm = TRUE
)
person_value_cutoff <- quantile(
person_summary$total_value,
probs = 0.95,
na.rm = TRUE
)
network_cutoff <- quantile(
person_summary$weighted_degree,
probs = 0.95,
na.rm = TRUE
)
person_flag_labels <- c(
high_frequency =
"transaction frequency in the top 5%",
high_value =
"combined sent and received value in the top 5%",
rapid_relay_flag = glue(
"at least two receipts followed by a new send within ",
"{params$immediate_relay_minutes} minutes"
),
managed_amount_flag =
"at least two unusually large or recordkeeping-adjacent transactions",
role_switching =
"repeated activity as both sender and payee",
network_flow_motif =
"fan-in activity or a partial split-and-recombine path",
agent_anomaly =
"high concentration at one agent with repeated missing identification and operator diversity",
network_prominent =
"weighted network degree in the top 5%"
)
person_scored <- person_summary |>
mutate(
high_frequency =
transaction_count >= person_frequency_cutoff,
high_value =
total_value >= person_value_cutoff,
rapid_relay_flag =
rapid_relay_count >= 2,
managed_amount_flag =
large_amount_count >= 2 |
near_threshold_count >= 2,
role_switching =
sent_count >= 2 &
received_count >= 2,
network_flow_motif =
(unique_senders >= 3 & received_count >= 3) |
person_id %in% partial_split_origins,
agent_anomaly =
dominant_agent_transactions >= 5 &
agent_share >= 0.80 &
dominant_agent_missing_id_rate >= 0.90 &
dominant_agent_operator_count >= 5,
network_prominent =
weighted_degree >= network_cutoff
)
person_flag_columns <- names(person_flag_labels)
person_scored$evidence_score <- rowSums(
person_scored[person_flag_columns]
)
person_scored <- person_scored |>
rowwise() |>
mutate(
evidence_reasons = paste(
unname(
person_flag_labels[
c_across(all_of(person_flag_columns))
]
),
collapse = "; "
)
) |>
ungroup() |>
arrange(
desc(evidence_score),
desc(agent_anomaly),
desc(rapid_relay_count),
desc(total_value),
desc(transaction_count)
)
top_people <- person_scored |>
slice_head(n = 3) |>
mutate(rank = row_number())| Evidence component | Point | Why does this matter? |
|---|---|---|
| transaction volume in the top 10% of established agents | 1 | Sustained volume gives an agent greater opportunity to facilitate repeated suspicious behaviour. |
| missing-identification rate in the top 10% of established agents | 1 | A high missing-identification rate can indicate weak controls or deliberate avoidance when compared with established agents. |
| large-transaction rate in the top 10% of established agents | 1 | A high share of transactions at or above $1,400 shows that the agent processes an unusually large proportion of the upper tail relative to peers. |
| at least two immediate-relay transactions and an immediate-relay rate of at least 5% | 1 | Repeated association with payees who become senders within 120 minutes connects the agent to immediate pass-through activity. |
| at least two transactions containing at least three of four indicators: missing identification, unusually large value, fast settlement, or immediate relay | 1 | A transaction containing at least three of four indicators is stronger than any single flag because identification, value, settlement, and relay behaviour overlap. |
| one person represents at least half of activity while several operators are used | 1 | A physical location dominated by one person but using several operators may indicate coordinated routing through the location. |
Rate comparisons exclude agents with fewer than 10 transactions. This prevents a location with one missing-ID event or one large transfer from receiving a 100% rate and outranking an agent with sustained activity. Among eligible agents, the score compares volume, missing identification, upper-tail value, immediate-relay involvement, overlapping transaction indicators and concentration around one person.
Recordkeeping-adjacent activity remains visible in the agent tables, but it no longer drives the general amount point. This avoids treating a narrow external reference as the only definition of unusual value.
The score construction code is shown below.
agent_send_long <- tx |>
transmute(
transaction_id,
agent_name = send_agent_name,
agent_city = send_agent_city,
agent_state = send_agent_state,
person_id = sender_name,
operator_name = send_operator_name,
role = "send",
amount,
id_missing = sender_id_missing,
large_amount,
near_threshold,
rapid_relay,
short_settlement
)
agent_pay_long <- tx |>
transmute(
transaction_id,
agent_name = pay_agent_name,
agent_city = pay_agent_city,
agent_state = pay_agent_state,
person_id = payee_name,
operator_name = pay_operator_name,
role = "pay",
amount,
id_missing = payee_id_missing,
large_amount,
near_threshold,
rapid_relay,
short_settlement
)
agent_long <- bind_rows(
agent_send_long,
agent_pay_long
) |>
filter(!is.na(agent_name)) |>
mutate(
agent_city = replace_na(agent_city, "UNKNOWN")
)
agent_transaction_flags <- agent_long |>
group_by(agent_name, agent_city, transaction_id) |>
summarise(
agent_state = mode_value(agent_state),
amount = first(amount),
id_missing = any(id_missing, na.rm = TRUE),
large_amount = any(large_amount, na.rm = TRUE),
near_threshold = any(near_threshold, na.rm = TRUE),
short_settlement = any(short_settlement, na.rm = TRUE),
rapid_relay = any(rapid_relay, na.rm = TRUE),
.groups = "drop"
) |>
mutate(
indicator_count =
as.integer(id_missing) +
as.integer(large_amount) +
as.integer(short_settlement) +
as.integer(rapid_relay),
multi_signal = indicator_count >= 3
)
agent_transaction_summary <- agent_transaction_flags |>
group_by(agent_name, agent_city) |>
summarise(
agent_state = mode_value(agent_state),
transaction_count = n_distinct(transaction_id),
total_amount = sum(amount, na.rm = TRUE),
missing_id_rate = mean(id_missing, na.rm = TRUE),
large_amount_rate = mean(large_amount, na.rm = TRUE),
near_threshold_rate = mean(
near_threshold,
na.rm = TRUE
),
rapid_relay_rate = mean(rapid_relay, na.rm = TRUE),
rapid_relay_count = sum(rapid_relay, na.rm = TRUE),
multi_signal_count = sum(multi_signal, na.rm = TRUE),
.groups = "drop"
)
agent_role_summary <- agent_long |>
group_by(agent_name, agent_city) |>
summarise(
role_events = n(),
unique_people = n_distinct(person_id, na.rm = TRUE),
unique_operators = n_distinct(
operator_name,
na.rm = TRUE
),
send_role_count = sum(role == "send"),
pay_role_count = sum(role == "pay"),
.groups = "drop"
)
agent_person_concentration <- agent_long |>
count(
agent_name,
agent_city,
person_id,
name = "person_role_events"
) |>
group_by(agent_name, agent_city) |>
summarise(
dominant_person_share =
max(person_role_events) /
sum(person_role_events),
.groups = "drop"
)
agent_summary <- agent_transaction_summary |>
left_join(
agent_role_summary,
by = c("agent_name", "agent_city")
) |>
left_join(
agent_person_concentration,
by = c("agent_name", "agent_city")
) |>
mutate(
eligible_for_rate_comparison =
transaction_count >=
params$minimum_agent_transactions
)
eligible_agents <- agent_summary |>
filter(eligible_for_rate_comparison)
agent_volume_cutoff <- quantile(
eligible_agents$transaction_count,
probs = 0.90,
na.rm = TRUE
)
agent_missing_id_cutoff <- quantile(
eligible_agents$missing_id_rate,
probs = 0.90,
na.rm = TRUE
)
agent_large_amount_cutoff <- quantile(
eligible_agents$large_amount_rate,
probs = 0.90,
na.rm = TRUE
)
agent_flag_labels <- c(
high_volume =
"transaction volume in the top 10% of established agents",
high_missing_id =
"missing-identification rate in the top 10% of established agents",
high_large_amount =
"large-transaction rate in the top 10% of established agents",
rapid_relay_flag = glue(
"at least two immediate-relay transactions and an immediate-relay rate of at least 5%"
),
multi_signal_flag = glue(
"at least two transactions containing at least three of four indicators: missing identification, unusually large value, fast settlement, or immediate relay"
),
concentration_anomaly =
"one person represents at least half of activity while several operators are used"
)
agent_scored <- agent_summary |>
mutate(
high_volume =
eligible_for_rate_comparison &
transaction_count >= agent_volume_cutoff,
high_missing_id =
eligible_for_rate_comparison &
missing_id_rate >= agent_missing_id_cutoff,
high_large_amount =
eligible_for_rate_comparison &
large_amount_rate >= agent_large_amount_cutoff,
rapid_relay_flag =
rapid_relay_count >= 2 &
rapid_relay_rate >= 0.05,
multi_signal_flag =
multi_signal_count >= 2,
concentration_anomaly =
eligible_for_rate_comparison &
dominant_person_share >= 0.50 &
unique_operators >= 5
)
agent_flag_columns <- names(agent_flag_labels)
agent_scored$evidence_score <- rowSums(
agent_scored[agent_flag_columns]
)
agent_scored <- agent_scored |>
rowwise() |>
mutate(
evidence_reasons = paste(
unname(
agent_flag_labels[
c_across(all_of(agent_flag_columns))
]
),
collapse = "; "
),
location = paste(
na.omit(c(agent_city, agent_state)),
collapse = ", "
)
) |>
ungroup() |>
arrange(
desc(evidence_score),
desc(multi_signal_count),
desc(rapid_relay_count),
desc(transaction_count),
desc(total_amount)
)
top_agents <- agent_scored |>
slice_head(n = 3) |>
mutate(rank = row_number())When person scores are equal, the shortlist is ordered by agent-concentration anomaly, immediate-relay count, total value and transaction count. When agent scores are equal, the order is determined by multi-signal transaction count, immediate-relay count, transaction count and total amount. The tie-breakers give priority to repeated combined evidence before raw value.
The scoring process now answers how many concerns are present. The shortlist must answer the more important question: which people and physical agents carry the clearest, most explainable combination of those concerns?
The objective is not to label every unusual transaction. It is to reduce a large file to a small group of people and locations that can be reviewed in detail. The final shortlist therefore includes only the highest-ranked entities after the amount, timing, agent and network evidence has been combined.
| Rank | Person | Evidence score | Transactions | Sent | Received | Total value | Immediate relays | Large amounts | Recordkeeping adjacent | Unique senders | Unique payees | Dominant agent |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | name_10 | 6/8 | 301 | 301 | 0 | $370,671 | 0 | 72 | 20 | 0 | 246 | agent_39 |
| 2 | name_1119 | 6/8 | 37 | 30 | 7 | $50,557 | 2 | 12 | 5 | 2 | 1 | agent_11 |
| 3 | name_143 | 6/8 | 28 | 26 | 2 | $28,545 | 2 | 4 | 0 | 1 | 9 | agent_11 |
Observed behaviour: 301 transactions with a combined value of $370,671. The person sent 301 transfers and received 0 transfers. The file links this person to 0 immediate relays, 72 unusually large transactions and 20 recordkeeping-adjacent transactions.
Why the person is shortlisted: transaction frequency in the top 5%; combined sent and received value in the top 5%; at least two unusually large or recordkeeping-adjacent transactions; fan-in activity or a partial split-and-recombine path; high concentration at one agent with repeated missing identification and operator diversity; weighted network degree in the top 5%.
Investigation focus: Review the relationship with agent_39, the stated purpose of repeat counterparties, source of funds and the chronology of linked receipts and sends.
Observed behaviour: 37 transactions with a combined value of $50,557.40. The person sent 30 transfers and received 7 transfers. The file links this person to 2 immediate relays, 12 unusually large transactions and 5 recordkeeping-adjacent transactions.
Why the person is shortlisted: transaction frequency in the top 5%; combined sent and received value in the top 5%; at least two receipts followed by a new send within 120 minutes; at least two unusually large or recordkeeping-adjacent transactions; repeated activity as both sender and payee; weighted network degree in the top 5%.
Investigation focus: Review the relationship with agent_11, the stated purpose of repeat counterparties, source of funds and the chronology of linked receipts and sends.
Observed behaviour: 28 transactions with a combined value of $28,545.28. The person sent 26 transfers and received 2 transfers. The file links this person to 2 immediate relays, 4 unusually large transactions and 0 recordkeeping-adjacent transactions.
Why the person is shortlisted: transaction frequency in the top 5%; combined sent and received value in the top 5%; at least two receipts followed by a new send within 120 minutes; at least two unusually large or recordkeeping-adjacent transactions; repeated activity as both sender and payee; weighted network degree in the top 5%.
Investigation focus: Review the relationship with agent_11, the stated purpose of repeat counterparties, source of funds and the chronology of linked receipts and sends.
The person table explains who accumulated the strongest evidence. It does not yet show whether those people relied on the same physical locations. The agent shortlist provides that operational view.
| Rank | Agent | Location | Evidence score | Transactions | Total value | Missing ID | Large amount rate | Recordkeeping adjacent rate | Immediate relays | Multi signal | Dominant person share | Operators |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | agent_39 | SAINT LOUIS PARK, MN | 6/6 | 491 | $603,306 | 87.0% | 24.0% | 5.9% | 67 | 62 | 59.1% | 234 |
| 2 | agent_73 | MONTERREY, NL | 4/6 | 182 | $225,429 | 100.0% | 30.8% | 8.2% | 1 | 19 | 16.5% | 2 |
| 3 | agent_246 | SAINT LOUIS PARK, MN | 4/6 | 18 | $10,090 | 88.9% | 0.0% | 0.0% | 8 | 8 | 61.1% | 11 |
Observed behaviour: 491 distinct transactions with $603,306 in facilitated value. Missing identification appears in 87.0% of transactions. Unusually large transactions represent 24.0%, recordkeeping-adjacent activity represents 5.9%, and 67 transactions connect to an immediate relay.
Why the agent is shortlisted: transaction volume in the top 10% of established agents; missing-identification rate in the top 10% of established agents; large-transaction rate in the top 10% of established agents; at least two immediate-relay transactions and an immediate-relay rate of at least 5%; at least two transactions containing at least three of four indicators: missing identification, unusually large value, fast settlement, or immediate relay; one person represents at least half of activity while several operators are used.
Investigation focus: Review agent onboarding, operator relationships, identification procedures, transaction logs, employee activity and the concentration of business around the dominant person.
Observed behaviour: 182 distinct transactions with $225,429 in facilitated value. Missing identification appears in 100.0% of transactions. Unusually large transactions represent 30.8%, recordkeeping-adjacent activity represents 8.2%, and 1 transactions connect to an immediate relay.
Why the agent is shortlisted: transaction volume in the top 10% of established agents; missing-identification rate in the top 10% of established agents; large-transaction rate in the top 10% of established agents; at least two transactions containing at least three of four indicators: missing identification, unusually large value, fast settlement, or immediate relay.
Investigation focus: Review agent onboarding, operator relationships, identification procedures, transaction logs, employee activity and the concentration of business around the dominant person.
Observed behaviour: 18 distinct transactions with $10,089.89 in facilitated value. Missing identification appears in 88.9% of transactions. Unusually large transactions represent 0.0%, recordkeeping-adjacent activity represents 0.0%, and 8 transactions connect to an immediate relay.
Why the agent is shortlisted: missing-identification rate in the top 10% of established agents; at least two immediate-relay transactions and an immediate-relay rate of at least 5%; at least two transactions containing at least three of four indicators: missing identification, unusually large value, fast settlement, or immediate relay; one person represents at least half of activity while several operators are used.
Investigation focus: Review agent onboarding, operator relationships, identification procedures, transaction logs, employee activity and the concentration of business around the dominant person.
The shortlist identifies the leading actors, but a list alone can hide the structure that made them important. The next section places the people, agents and transaction paths on one network so that the relationships can be examined as a connected story.
Transaction-level tests identify unusual records, but laundering patterns often become clearer only when the records are connected. A person who appears ordinary in isolation may sit between several senders and recipients. An agent with high volume may be routine until the same location repeatedly appears on paths containing missing identification, unusual values or immediate relays.
The investigative network therefore represents people and physical agents as nodes. Transaction frequency determines edge width. Larger nodes represent greater weighted activity. Red edges identify paths containing at least two of four transaction indicators: missing identification, an unusually large or recordkeeping-adjacent value, fast corridor-adjusted settlement and immediate relay.
The graph is deliberately restricted to the top people and agents, together with their most relevant connected nodes. Plotting every transaction would turn the network into a dense picture without an investigative focus. The purpose of the graph is not to show everything. It is to show why the shortlisted entities deserve attention.
# Create one physical key for each agent location.
tx_network <- tx |>
mutate(
send_agent_key = paste(
"AGENT",
send_agent_name,
send_agent_city,
sep = "::"
),
pay_agent_key = paste(
"AGENT",
pay_agent_name,
pay_agent_city,
sep = "::"
),
sender_key = paste(
"PERSON",
sender_name,
sep = "::"
),
payee_key = paste(
"PERSON",
payee_name,
sep = "::"
),
amount_indicator =
large_amount |
near_threshold,
transaction_indicator_count =
as.integer(any_id_missing) +
as.integer(amount_indicator) +
as.integer(short_settlement) +
as.integer(rapid_relay),
suspicious_transaction =
transaction_indicator_count >= 2
)
selected_person_keys <- paste(
"PERSON",
top_people$person_id,
sep = "::"
)
selected_agent_keys <- paste(
"AGENT",
top_agents$agent_name,
top_agents$agent_city,
sep = "::"
)
relevant_transactions <- tx_network |>
filter(
sender_key %in% selected_person_keys |
payee_key %in% selected_person_keys |
send_agent_key %in% selected_agent_keys |
pay_agent_key %in% selected_agent_keys
)
node_relevance <- bind_rows(
relevant_transactions |>
transmute(
node = sender_key,
suspicious_transaction,
amount
),
relevant_transactions |>
transmute(
node = payee_key,
suspicious_transaction,
amount
),
relevant_transactions |>
transmute(
node = send_agent_key,
suspicious_transaction,
amount
),
relevant_transactions |>
transmute(
node = pay_agent_key,
suspicious_transaction,
amount
)
) |>
group_by(node) |>
summarise(
activity = n(),
suspicious_activity =
sum(suspicious_transaction),
amount = sum(amount, na.rm = TRUE),
.groups = "drop"
) |>
mutate(
selected_suspect =
node %in%
c(
selected_person_keys,
selected_agent_keys
)
) |>
arrange(
desc(selected_suspect),
desc(suspicious_activity),
desc(activity),
desc(amount)
)
retained_nodes <- node_relevance |>
filter(selected_suspect) |>
bind_rows(
node_relevance |>
filter(!selected_suspect) |>
slice_head(n = 24)
) |>
distinct(node) |>
pull(node)
network_edges <- bind_rows(
relevant_transactions |>
transmute(
from = sender_key,
to = send_agent_key,
suspicious_transaction,
amount
),
relevant_transactions |>
transmute(
from = send_agent_key,
to = pay_agent_key,
suspicious_transaction,
amount
),
relevant_transactions |>
transmute(
from = pay_agent_key,
to = payee_key,
suspicious_transaction,
amount
)
) |>
filter(
from %in% retained_nodes,
to %in% retained_nodes
) |>
group_by(from, to) |>
summarise(
frequency = n(),
total_amount = sum(amount, na.rm = TRUE),
suspicious_edge =
any(suspicious_transaction),
.groups = "drop"
)
network_nodes <- tibble(
name = unique(
c(network_edges$from, network_edges$to)
)
) |>
mutate(
entity_type = if_else(
str_starts(name, "PERSON::"),
"Person",
"Agent"
),
display_name = str_remove(
name,
"^(PERSON|AGENT)::"
),
suspect_type = case_when(
name %in% selected_person_keys ~
"Top person suspect",
name %in% selected_agent_keys ~
"Top agent suspect",
entity_type == "Person" ~
"Connected person",
TRUE ~
"Connected agent"
)
) |>
left_join(
node_relevance,
by = c("name" = "node")
) |>
mutate(
activity = replace_na(activity, 1),
label = if_else(
suspect_type %in%
c(
"Top person suspect",
"Top agent suspect"
),
display_name,
NA_character_
)
)
investigative_graph <- tbl_graph(
nodes = network_nodes,
edges = network_edges,
directed = TRUE
)
ggraph(
investigative_graph,
layout = "fr"
) +
geom_edge_link(
aes(
width = frequency,
color = suspicious_edge,
alpha = suspicious_edge
),
arrow = arrow(
length = unit(2.5, "mm"),
type = "closed"
),
end_cap = circle(3, "mm")
) +
geom_node_point(
aes(
size = activity,
color = suspect_type,
shape = entity_type
),
alpha = 0.92
) +
geom_node_text(
aes(label = label),
repel = TRUE,
fontface = "bold",
size = 3.8
) +
scale_edge_color_manual(
values = c(
`FALSE` = "grey75",
`TRUE` = "#b22222"
),
breaks = c(FALSE, TRUE),
labels = c(
"Ordinary edge",
"Multi-indicator edge"
)
) +
scale_edge_alpha_manual(
values = c(
`FALSE` = 0.25,
`TRUE` = 0.85
)
) +
scale_edge_width_continuous(
range = c(0.3, 3.2)
) +
scale_color_manual(
values = c(
"Top person suspect" = "#b22222",
"Top agent suspect" = "#d97706",
"Connected person" = "#4f81bd",
"Connected agent" = "#7a8b80"
)
) +
scale_shape_manual(
values = c(
"Person" = 16,
"Agent" = 15
)
) +
scale_size_continuous(
range = c(3, 14)
) +
guides(
color = guide_legend(
title = "Node classification",
nrow = 2,
byrow = TRUE,
order = 1,
override.aes = list(size = 5)
),
shape = guide_legend(
title = "Entity type",
nrow = 1,
order = 2,
override.aes = list(size = 5)
),
edge_width = guide_legend(
title = "Transaction frequency",
nrow = 1,
order = 3
),
edge_color = guide_legend(
title = "Transaction indicator",
nrow = 1,
order = 4,
override.aes = list(
edge_width = 1.5,
edge_alpha = 1
)
),
size = "none",
edge_alpha = "none"
) +
labs(
title = "Investigative network around the leading people and physical agents",
subtitle = paste(
"Node size reflects activity;",
"red edges contain at least two independent transaction indicators"
),
caption = glue(
"The graph is a focused one-hop and agent-path view, not the full ",
"{comma(nrow(tx))}-transaction network."
)
) +
theme_graph(base_family = "Arial") +
theme(
legend.position = "bottom",
legend.box = "vertical",
legend.box.just = "left",
legend.direction = "horizontal",
legend.margin = margin(t = 10),
legend.spacing.y = unit(4, "pt"),
plot.caption = element_text(
hjust = 0.5,
margin = margin(t = 12)
),
plot.margin = margin(
t = 12,
r = 25,
b = 12,
l = 12
)
)The network answers two questions that the tables cannot answer alone. First, it shows whether the highest-scoring people are isolated outliers or whether they are connected through common agents and counterparties. Second, it shows which physical locations repeatedly sit between prominent senders and dispersed recipients.
A large node is not automatically suspicious. It may simply represent the main location or a customer who uses the service frequently. The concern increases when a prominent node is also joined by red edges, repeated missing identification, unusual amount behaviour or immediate relay. This distinction is especially important for agent_11, which dominates send volume but must still earn its place on the shortlist through supporting evidence.
The network completes the movement from isolated transactions to connected behaviour. The final section returns to the decision that prompted the analysis: which people and agents should be reviewed first, and what information is needed to test a lawful explanation?
The people requiring the highest-priority enhanced review are name_10, name_1119, name_143. The physical agents requiring the highest-priority control review are agent_39, agent_73, agent_246.
The answer is based on convergence rather than one unusual transaction. The leading cases combine high frequency or value with repeated immediate-relay behaviour, amount management, agent concentration or network prominence. The shortlist is intentionally small so that investigative effort can be directed toward the entities with the clearest and most explainable combination of evidence.
The purpose of the analysis is not to maximise the number of alerts. It is to direct attention toward behaviour that is unusual, repeated and explainable. A defensible review should surface the strongest connected patterns early while avoiding unnecessary escalation of ordinary remittance activity. The six shortlisted entities provide the starting point, not the final conclusion.
The analytical benchmark and typologies were informed by the following official materials: