How cases begin
Of 2,333 regulator documents read for origin statements, 75 (3%) say where a case began: a referral, a surveillance or data-unit detection, an examination. Another 866 only thank other bodies or record a firm's self-report; 1,355 (58%) are silent. Silence means the document does not say, not that nobody referred the matter.
Most of what a regulator publishes about a case says what the respondent did and what it cost them. Very little says how the regulator came to look. Reading the closing paragraphs of 2,333 documents for that one thing shows how thin the public record of origins is. Seventy-five documents (3%) say where a case began. Another 866 thank a named body for help, or record that the respondent reported itself, and 37 ASIC releases note that ASIC referred the matter on to a prosecutor. The other 1,355 (58%) are silent, and 246 of those do no more than name the staff who ran the investigation.
Silence is not evidence that a case had no referral, tip or surveillance alert behind it. It means the document does not say. A regulator that chose not to name its source, or whose house style leaves the credit paragraph out, looks the same here as a case that began cold. Everything below is a count of what the documents say, and none of it is a count of how cases actually start.
How it was measured
The records in the library each link to the regulator’s own release, order or notice. For every record the script finds the cached copy of that document, reduces it to the body text (cutting menus, footers and the whistleblower links in every CFTC page footer), splits it into sentences and runs a list of phrase patterns. The script is scripts/analysis/how-cases-begin.mjs; it reads the case files and the cache, fetches nothing, and prints every number used here.
Of the 2,428 records, 2,333 could be assessed. The 69 Ontario Capital Markets Tribunal records link to a proceeding index page that lists documents but carries no narrative, and 26 SEC administrative-proceeding records have no cached copy of the page they link to. Both groups are left out of every denominator below.
The detector looks for explicit wording only. It sorts what it finds into tiers:
- Explicit origin. A statement that a case or investigation began with, followed from or resulted from a referral, a complaint from the public, a detection by a surveillance team or an exchange, a data-analytics unit, an examination, a short-seller report or a request from a foreign regulator, or a whistleblower award.
- Self-report. The respondent reported its own conduct to the regulator, in a sentence that says it did.
- Credit. The regulator thanks, or says it was assisted by, a named agency, exchange, prosecutor or specialist unit. A sentence only counts if it names a body other than the respondent; a respondent’s cooperation, which earns a discount in the penalty, does not.
- Hand-off. ASIC says a prosecutor took the matter after an ASIC referral. This says how the criminal case began, not where the investigation came from, so it is kept apart.
A fifth thing, naming the staff who conducted the investigation, is recorded but is not an origin statement. Credit is also not origin: a body thanked for assistance may have started the matter or may only have helped with it, and the sentences rarely say which.
Figures are as of the evening of 2026-10-04, after the status and flag corrections. The documents themselves are unchanged, so the origin, credit and silence counts did not move. Three inputs did. The status of 543 of the 733 records that had status “filed” was researched and updated (190 remain filed), which reshuffles the status comparison below. The parallel-criminal-case flag was corrected (816 records now true), which changes that comparison. And one SEC litigation release lost its insider-trading tag, so that row now has 391 releases, 281 of them crediting someone.
How accurate the detector is
I read the matched sentence, and where it was ambiguous the paragraph around it, for 164 detected records: all 75 explicit-origin documents, all 21 self-reports, 8 of the 39 hand-offs and a random 60 of the 896 credit documents. In the final version every one of those was a statement of the kind claimed. That is optimistic, because the rules were adjusted over several rounds of exactly this reading: early versions fired on a sentence about a referral fee, on a respondent’s cooperation, on a person named as assisting the investigators, on an internal committee referral and on a rule telling firms to report themselves. Each of these was found by reading, then vetoed or rewritten.
For the other direction I read 134 documents the detector had passed over, stratified by document family, looking at the opening, the closing paragraphs and every sentence containing a suspect word. In three rounds, 7 hid a statement the detector of the day had missed (5%): two examination-led origins, two credits to SEC specialist units, a credit to a US Attorney’s office, a credit to a state commission and one firm bringing its own conduct to the FCA. After each round the rules were widened and the whole set re-run. The last fresh sample of 45 found one miss (2%), and that is the only measure not tuned on the same records. A reasonable summary is that at least 95% of the documents counted as silent really say nothing explicit about origin or credit, and that the explicit-origin count of 75 is probably a small undercount, not an overcount.
A second pass counted a few headline figures a different way, by searching the plain text for the thanks wording and a fixed window after it. FINRA credited in SEC litigation releases came to 329 by that method against 332 in the script; the FBI 370 against 371; the CME Group in CFTC releases 55 against 69, because the simple method looked only at the first thanks sentence.
What the documents say
The gaps are mostly a matter of document type. SEC litigation releases are short announcements that often end with two sentences, one naming the investigators and one thanking other bodies. SEC administrative orders and judge decisions are the order or decision itself, which has no such paragraph: none of the 559 orders carries a credit, and the 12 that say anything are firms reporting themselves. CFTC press releases carry a thanks paragraph in about seven in ten (233 of 338). FCA final notices describe the conduct and the penalty in numbered paragraphs, and four of the 48 record a firm reporting its own conduct to the Authority.
| Document family | Assessed | Explicit origin | Self-report only | Credit only | ASIC hand-off | Silent |
|---|---|---|---|---|---|---|
| SEC litigation releases | 1,214 | 60 | 0 | 616 | 0 | 538 |
| SEC administrative orders | 559 | 0 | 12 | 0 | 0 | 547 |
| SEC judge decisions | 61 | 1 | 0 | 0 | 0 | 60 |
| CFTC press releases | 338 | 2 | 4 | 228 | 0 | 104 |
| ASIC media releases | 113 | 12 | 1 | 1 | 37 | 62 |
| FCA final notices | 48 | 0 | 4 | 0 | 0 | 44 |
| All | 2,333 | 75 | 21 | 845 | 37 | 1,355 |
Each document is counted once, in the highest tier it reaches. Of the 538 silent SEC litigation releases, 241 name the investigators and nothing else.
The 75 explicit statements
Sixty-one are SEC documents (39, 19 and 3 in the three groups below), 12 are ASIC and 2 are CFTC; none is an FCA notice, and none is a whistleblower award.
- A data-analytics unit (39 SEC releases). These say the case came out of the Market Abuse Unit, and 34 of them name its Analysis and Detection Center, in the same stock sentence about tools that detect suspicious trading patterns. Bryan Cohen and Brian Wong are two of them. All 34 are insider-trading cases; they are 9% of the 391 SEC litigation releases tagged insider trading. The Market Abuse Unit is named somewhere in 281 SEC documents, but in most it appears only as the team that ran the investigation, which is not an origin.
- An examination (19 SEC releases). The investigation followed an SEC examination or an examination-programme referral, as in the Constantin social-media ramping case and Heckler.
- Surveillance, a referral or a complaint (12 ASIC, 2 CFTC). Of ASIC’s 12, two follow a referral from the ASX (one example), nine point to ASIC’s own market surveillance team, directly or by an internal referral (2025, 2016, Openmarkets), and one follows complaints from the public (a boiler-room raid). One wording is ambiguous about whose surveillance team is meant. The two CFTC statements are an exchange’s compliance department (Logista) and the CME Group’s market regulation department (a 2015 spoofing case).
- Other (3 SEC). A data-driven review of executives’ trading plans (Peizer), a request for assistance from a foreign authority (Waymack), and an administrative-law-judge decision that records short-seller reports and a late annual filing as what prompted two of its investigations.
Whistleblower awards appear in none of the 2,333 documents. The SEC and the CFTC both run award programmes, and 63 documents carry the standard notice that whistleblowers can receive part of the sanctions, but awards are announced separately and none is attached to these case records. One release mentions complaints from the public, the ASIC raid above, and none mentions a tip or a whistleblower as the source. That is a fact about how releases are written; it says nothing about how many cases begin that way.
The SEC’s explicit statements are also getting more common. Among SEC litigation releases, 25 of 739 filed from 2015 to 2021 (3%) carry one, against 35 of 470 filed from 2022 to 2026 (7%). The rise comes from Analysis and Detection Center statements, 12 filed in the first period and 22 in the second, while examination-led ones fell from 12 to 7. It shows a change in what the releases say, not in how many cases begin that way.
Who is thanked
Of the 2,333, 896 documents credit at least one named body, and the same few names run through them. US Attorney’s offices are named in 482, the FBI in 447 and FINRA in 333. The Department of Justice itself (61), IRS Criminal Investigation (40) and the Postal Inspection Service (27) follow. Many documents credit several kinds at once: 122 name three or more kinds of body.
Foreign regulators appear in 151: the UK FCA or its predecessor in 70, Canadian provincial and national bodies in 47, Asian regulators in 42 and European ones in 37, but those counts overlap, because a single long thanks list can name bodies from a dozen jurisdictions. Seventy-five of the 151 are SEC releases and 76 are CFTC releases. Exchanges and the National Futures Association are named in 91 documents, and 90 of the 91 are CFTC press releases: the CME Group in 69 (20% of all CFTC releases), the NFA in 19 and ICE in 7. The SEC rarely credits an exchange: one of 1,214 litigation releases does so, whereas FINRA is credited in 332 of them (27%). The SEC and CFTC credit each other in 36 CFTC releases and 5 SEC ones.
The SEC’s own units are credited in 115 documents, usually in a sentence saying a named analyst from the Market Abuse Unit, the Analysis and Detection Center, the Office of Market Intelligence or the Office of Investigative and Market Analytics helped. In 28 of those, no outside body is credited at all.
Which techniques mention whom
The mix of bodies depends on the technique, though the table below mixes this with the kind of case. SEC litigation releases have a large enough sample to compare, so the comparison is within them.
In the grid, “Insider” is insider trading, “Pump” pump and dump, “Unreg.” unregistered distributions, “Boiler” boiler rooms, “Control” undisclosed control blocks and “Promo” paid stock promotion. A release with two technique tags appears in both rows.
- Insider trading (391 SEC litigation releases) credits FINRA in 217 (55%) and US prosecutors or agents in 157 (40%); 281 (72%) credit someone. Insider trading is also the only technique with Analysis and Detection Center origin statements, 34 of them. The releases do not say what FINRA did.
- Ponzi schemes (289) almost never credit FINRA (8, 3%). They credit US prosecutors or agents in 96 (33%) and anyone in 135 (47%).
- Pump and dump (104) credit FINRA in 41 (39%), US prosecutors or agents in 33 (32%) and a foreign regulator in 19 (18%), the highest foreign share among the larger groups.
- Undisclosed control blocks (50) credit FINRA in 20 (40%) and a foreign regulator in 10 (20%).
- Unregistered distributions (66) are the least credited, with 18 (27%) naming anyone.
Among the smaller groups, 14 of 35 paid stock promotion releases credit someone and 26 of 52 boiler-room releases do; price manipulation (14 of 21) is too small to treat as a rate.
For CFTC press releases the pattern differs. Spoofing (73 releases) credits an exchange in 51 (70%), and 57 (78%) credit anyone. Price manipulation (55) credits foreign regulators in 29 (53%) and US prosecutors or agents in 26 (47%). Benchmark-submission rigging (18) credits foreign regulators in 15 and US prosecutors or agents in 15, a sample too small for more than a description. Ponzi schemes (109) credit US prosecutors or agents in 55 (50%) and an exchange in 11 (10%). So the names differ by technique: exchanges in spoofing releases, overseas authorities in price manipulation. The releases do not say what each body supplied; how spoofing gets caught covers what the records show of detection.
For ASIC the explicit origin statements are few and cluster in insider trading: 7 of 55 insider-trading releases (13%) and 3 of 36 price-manipulation ones (8%). Another 38 ASIC releases (37 with no other kind of statement) note that ASIC referred the matter to the Commonwealth prosecutor, which says who brought the criminal case and not how ASIC found it.
Two things that sit behind the technique differences
The parallel criminal case. Among SEC litigation releases, 337 of the 530 flagged as having a criminal parallel (64%) credit US prosecutors or agents, against 89 of the 684 without one (13%). In CFTC releases it is 79 of 108 (73%) against 38 of 230 (17%). Technique differences partly follow this: insider trading has 181 criminal parallels in 391 SEC releases and Ponzi schemes 142 in 289. A thanks paragraph to a prosecutor goes with a criminal case running alongside, whatever the technique. The flag was corrected on 128 records on 2026-10-04 (113 set to false because the release states no criminal case, 15 set to true), which narrowed the gap: earlier counts of 74% against 3% (SEC) and 82% against 9% (CFTC) rested on a flag that was too generous. Part of the remaining 13% and 17% without the flag may be criminal cases the release does not mention.
The stage of the case. Of SEC litigation releases with status “filed”, 86 of 157 (55%) credit someone; for “settled” it is 325 of 486 (67%), and for “judgment” 239 of 547 (44%). Many judgment releases are follow-ups to a case announced years before, and they tend to carry the investigators’ names and nothing else. Record status is a flawed measure: it records the outcome as researched, often for some defendants only, and 190 records across the library are still marked “filed” because no outcome was found.
What this does not show
- Selection. These are the cases regulators chose to bring and publish. A matter that began with a referral and ended with no action leaves no document here, and nothing in the library says how many cases begin with each source.
- Silence is not absence. A document without an origin statement does not tell you there was no referral, tip or alert. Practice varies: SEC administrative orders carry no credit paragraph, and the FCA’s notices do not describe how a matter was found. A source that regulators keep out of their releases cannot appear in this analysis, however common it is.
- Credit is not origin. Thanking FINRA or the FBI says they helped. It does not say they started the matter. The 896 credit documents are an upper bound on documents that could show an outside body involved, and the 75 explicit statements are the only ones that speak to a beginning.
- Different document types. An SEC order, an SEC release, a CFTC press release, an ASIC release and an FCA notice have different shapes, so a rate of 56% for one and 8% for another is a difference in how the document is written as much as in how cases begin. Comparisons across agencies are comparisons of documents.
- Documents not assessed. 69 Ontario records link to an index page with no narrative and 26 SEC records have no cached page, so 95 of the 2,428 records are outside the denominators. The Ontario tribunal’s decisions might say more; they were not read here.
- A detector, not a reader. It matches phrases. It found a statement the reader would accept in each of the 164 documents I read, and missed a statement in about 2% to 5% of the silent ones I read, so some documents are counted as silent that are not. Names of bodies come from a list of patterns; a body I did not list is not counted by name, and the grouping into “foreign regulator” and similar is mine.
- Technique tags are the library’s own. They come from the audit described in What reading every record found, and a record can carry several, so the technique rows overlap. Whether the thanks paragraph relates to the tagged conduct is not checked.
- Checked by AI agents, not lawyers. The records were compared with their primary documents by Claude AI agents, and the origin statements here were read by one too. No lawyer reviewed them. If a record or a count looks wrong, use the error link on its page or write to [email protected] with the slug; see the editorial policy.
Reproducing the counts
Run node scripts/analysis/how-cases-begin.mjs out.json. It reads src/content/cases/*.json and the cached documents under .cache/, takes about half a minute, and writes the full output: coverage by agency and document type, counts of each kind of statement, the named bodies, and the technique, year, criminal-parallel and status breakdowns. PDF text is extracted with the unpdf package the ingest pipeline already uses. The phrase list is in the script, along with the list of sentences it vetoes.