What reading every record found
Two audit rounds on 2 and 3 October 2026 read all 2,428 case records against regulators' own documents. Claude AI agents did the reading; no lawyer reviewed it. 487 records lost a technique tag, 345 now carry none, and money, status or names changed on many more. The cause was keyword rules matching words in boilerplate, not conduct.
The library’s technique tags were assigned by keyword rules running over regulators’ releases. Nobody read most of the documents behind them. On 2 and 3 October 2026 that changed: all 2,428 case records were read against the regulator’s own document, and where the tag, the money, the status or the names did not match, they were corrected. This post is the evidence behind the claim, made on the editorial policy page, that a checked record has been compared with its primary document.
Who did the checking, and how
The reading was done by Claude AI agents, not by people. Eight agents each took one technique group in the first round on 2 October, and ten took lists in the second round on 3 October. Each worked from a written brief: open the cached copy of the regulator’s document (a record whose document could not be read was left unchecked), decide whether it charges or finds the tagged conduct or merely mentions it, and correct the tags, money, status and names. A red flag, a quoted comparator, a respondent’s history or a description of a policy does not support a tag.
The lead session, also Claude, re-read random samples of the agents’ work and every disputed case; the maintainer made the policy calls, such as deleting non-case records. No lawyer reviewed any of it. That matters for what “checked” means: the editorial policy, the case pages and the data export all now say what that means: a comparison made by an AI model following written instructions, with samples and disputed cases read a second time. An earlier version of that wording said “a person or a documented review pass”, which was vague about who; it was corrected on 3 October when this post was written.
Two groups were sampled in round one, not read in full: about 150 of 725 insider-trading records, drawn from a random stratified sample and from targeted sets of likely false positives, and 147 of 602 Ponzi records. Round two read the remainder of both (569 and 441), plus Rule 105 and naked short selling (132 records), a structural group of reverse mergers, control blocks, free riding and benchmark rigging (111) and a miscellaneous list (128). The agents’ own tallies were approximate and overlapped, so the numbers below come from comparing the two states of the library instead.
How the numbers were computed
A script compares every case file in the git commit from the morning of 2 October, before round one, with the
file today. Records were renamed on 3 October, so it matches them by their releaseUrl. Thirteen records were
deleted in between, five and then eight CFTC annual-results summaries and non-case notices, and are excluded; 2,428
records are present in both states. The script is a short Node file in the audit working notes, not in the repository, and re-running it on the
same two states gives the same output.
Two cautions. The comparison counts everything that changed between the two commits. Other fixes landed in the same window: CFTC filing dates one day late on 91 records, rewritten titles and new addresses for 551 records. None of these touches tags, money, status or names, but tag edits made on 2 October while the front-running post was revised count. And a changed money or names field is not always a wrong one: a blank interest figure that was filled in counts, as does “et al.” split into the people named. The script cannot tell a corrected number from a completed one.
What changed
Of the 2,428 records, 487 (20%) lost at least one technique tag, 105 (4%) gained one, and 345 (14%) went from tagged to untagged. No untagged record gained a tag. Counted as tags, 533 of 2,744 (19%) were removed and 130 were added, the largest additions being 39 to undisclosed control blocks and 19 each to matched orders and unregistered distributions.
Tags were the smallest of the four kinds of change. The defendants list changed in 962 records (40%), any of the four money fields in 753 (31%), and the status in 681 (28%); 1,868 records (77%) changed in at least one of the four. Status matters most here, because the corrections policy ranks a wrong dismissal first. Fifty-two records moved off “dismissed” to a judgment, a settlement or a filed matter, and two moved onto it. The largest other changes were 173 unknown statuses that became filed and 151 judgments that became settlements, mostly consents still awaiting court approval.
The tag changes were not spread evenly.
| Technique | Tags before | Tags after | Removed | Share removed |
|---|---|---|---|---|
| Insider trading | 722 | 633 | 89 | 12% |
| Ponzi schemes | 600 | 474 | 127 | 21% |
| Unregistered distributions | 160 | 157 | 22 | 14% |
| Price manipulation | 130 | 131 | 8 | 6% |
| Paid stock promotion | 129 | 66 | 64 | 50% |
| Pump and dump | 125 | 122 | 19 | 15% |
| Spoofing | 109 | 96 | 15 | 14% |
| Wash trading | 100 | 50 | 50 | 50% |
| Rule 105 offering shorts | 84 | 85 | 0 | 0% |
| Boiler rooms | 77 | 68 | 10 | 13% |
The “after” column includes tags added in the audit. The chart adds four smaller techniques that fared worst.
Three patterns stand out. First, the error followed the vocabulary. Techniques named by an ordinary word in regulatory writing, such as a promoter, a wash sale or a fictitious sale, lost half their tags; techniques named by a specific statute and a specific rule, such as Rule 105, lost none. Second, groups differed in what was wrong. Insider trading kept most of its tags but had the worst field errors: the newest list of 142 records had wrong money, status or names in about four in five, because the first ingest took only the first dollar figure and read “settled” or “judgment” from the verb in the release. Third, the Ponzi group surprised the reviewers. Round one’s random sample of 31 Ponzi records found no wrong tag, and round two then removed the tag from 109 of 441. The difference is a definition. Round one counted a single “Ponzi-like payments” clause as support; round two did not, unless the release gave the scheme an amount or put it in the lead. The sample was too small to predict a definitional shift, and 10 to 15 clause-only records could still go either way.
Why the tags were wrong
The recurring causes, each with one record you can check.
A word in a policy description. The SEC’s order against Monness, Crespi, Hardt & Co. was tagged front running and insider trading. The phrase “research front running” appears once, in a description of what the firm’s own restricted-list policy was for. The order finds that the firm did not enforce its procedures; it does not say anyone traded ahead of a report. The front-running post tells that story at length.
A tax rule with the same name. Betterment and Wealthfront were tagged wash trading. Their orders concern tax-loss harvesting disclosures, and “wash sale” there is the Internal Revenue Service rule that disallows a loss when a substantially identical stock is bought within thirty days. It is not the manipulation of the same name.
A substring. The rule for exchange wash trading matched “wash” inside “Washington”. The audit reports record this cause without naming one record, so there is no case to link.
A list of red flags. OTC Link carried four tags: layering, spoofing, wash trading and unregistered distributions. The SEC’s order faults it for failing to monitor, investigate and report suspicious activity, and the manipulation types appear only as kinds of activity a compliance policy lists for flagging. It now carries none. The same pattern removed tags from the other suspicious-activity-report failure orders in the wash, spoofing, pump-and-dump and unregistered groups.
Somebody else’s number. Brian Hunter was recorded with a $7.5 million penalty. That figure is the penalty Amaranth, his former employer, settled for in 2009, recited in the same release. The release says the court ordered Hunter to pay $750,000. A similar slip, a CFTC headline “pay over $X” total that includes restitution entered as the penalty, affected 17 Ponzi records.
A status read from the wrong word. Wells Fargo Advisors was recorded as dismissed; the order censures it and imposes a $5 million penalty, and the wash-trading tag it carried is gone. Dismissal of a relief defendant, one count or a defunct company set the same wrong status in other records. The Polevikov record, caught in the earlier front-running check and so outside the numbers above, was wrong in just this way: the court dismissed claims against his wife as a relief defendant, and judgment was entered against him.
A clause in a fraud about something else. The SEC’s complaint against Adam and Daniel Kaplan alleges overbilled fees and misappropriation, and “Ponzi-like payments” made to some clients to conceal them. It does not charge a Ponzi scheme, and the tag was removed.
Two more causes have no record to link. Annual-results announcements from the CFTC were stored as cases under a single respondent’s name with unrelated dollar figures; some of the 13 deleted records were of this kind. And the CFTC’s index showed 110 of 351 cached releases exactly one day later than the releases themselves, which is why 91 records’ dates were corrected.
What changed in the rules and the process
Because the causes sat in the rules, fixing the records alone would have let the ingest reproduce them. The keyword rules were re-run against the audited records, treating the audited tags as the answer. Precision, the share of tags the rules assigned that the audit confirmed, rose from 76% to 92%. False positives fell from 327 to 95. The price was recall, the share of audited tags the rules found: 91% before and 90% after, with true positives down by about eight. By technique the losses were concentrated. Paid stock promotion went from 98% to 85% recall, because nine of the missed promoters use the same “stock promoter” sentence as the 56 false positives and the rule cannot tell them apart; the audit’s readers could, because they read the whole document.
That figure is weaker than it sounds. The rules were tuned on the same 1,058 records they were then measured on, so 92% describes how well the rules now reproduce the audit, not how they will do on next month’s releases. The audit set leaned toward the groups already suspected, and some audited tags are themselves judgement calls. The rule changes have not been run across the library, so the corrected records come from the reading and not from the new rules.
The ingest now refuses several things. It rejects annual-results summaries, enforcement advisories, manuals and task-force or personnel notices before classification. The tax meaning of wash sale, the vocabulary of compliance and suspicious-activity failures, and “red flag” language veto the manipulation tags. Insider trading now needs a trading anchor such as a tipper, a misappropriation or a purchase while in possession of non-public information. Bare “stock promoter” no longer tags paid promotion.
Each case page now carries a “checked against the primary document” line with the date and a “Report an error in this record” link that opens an email with the address filled in. The case explorer filters on the checked mark and exports it, and the corrections page logs every group’s findings.
What this does and does not mean
It means the counts now measure something narrower than before. A count under a technique is a count of enforcement as regulators published it, passed through our reading of what each document charges. For the records read, the error rate is measured: before the audit about one record in five carried a tag the reader decided the document did not support.
It does not mean the library is right. The reading was done once, by an AI agent, with a second look at samples and disputes. Misses are possible, and a “checked” mark is not an endorsement, as the editorial policy says. The briefs asked the agents to look for keyword false positives, so they were primed to find them, and the audit was one-directional: it asked whether a tag was supported, not whether an untagged record deserved one. Nobody has measured what the classifier never ingested. After the audit 364 records carry no technique, up from 16, because the library no longer forces a tag where the document does not support one.
The last six were read on 3 October, after the counts in the charts above were taken. The Ontario tribunal decisions for Caruso and Candusso were fetched and read: the Caruso tribunal dismissed every allegation against Caruso and his co-respondent Sidders and sanctioned only Cornish, and the Candusso tribunal found five of seven respondents liable and recomputed the penalty total to C$2.95 million. Cormark turned out not to be an insider-trading case at all, and the Tribunal found none of the allegations proven, so it lost its tag; that is one more removed tag and one more untagged record than the charts show. Swift Trade was confirmed as layering from the Court of Appeal judgment. LeadFX and CFTC release 8716-23 are not enforcement actions and carry no tag. A few records were confirmed from a regulator’s summary page alone.
Some tags remain judgement calls. Between 12 and about 40 unregistered sellers of other people’s Ponzi products, such as Woodbridge and EquiAlt promoters, kept the Ponzi tag, because the documents call the scheme Ponzi and charge a Section 5 violation, and whether the seller ran it is an editorial decision. A handful of binary-options call-centre cases were kept as weak boiler-room fits because the SEC itself uses the term. These are editorial choices; the tags on the technique pages are ours, not the regulators’. The earlier post What the enforcement record is made of shows how the library’s mix looks by family; its counts were recomputed after the audit.
If you find a record the audit got wrong, use the error link on its page or write to [email protected] with the slug and a link to the document. It will be logged on the corrections page.