ETH Zurich and Anthropic study finds LLMs can unmask anonymous accounts for a few dollars each
0
0

Researchers at ETH Zurich, AI safety group MATS, and Anthropic have developed an automated LLM pipeline that can connect pseudonymous posts to real-world identities. It was able to identify hundreds of Hacker News users at 90% precision for as low as $1 per user.
The February paper made a comeback this week on X and Reddit.
The paper, “Large-scale online deanonymization with LLMs,” was first posted on arXiv on February 18. Later, it appeared in the proceedings of the 35th USENIX Security Symposium.
The authors are Simon Lermen, Daniel Paleka, Joshua Swanson, Michael Aerni, Nicholas Carlini, and Florian Tramèr. Carlini is an employee of Anthropic, the creator of Claude, and the others are affiliated with ETH Zurich and MATS.
Four stages narrow a pool of 89,000 candidates to a shortlist
No stolen data or hacked server implicated. The agent reads stuff anyone can already see, then does the work of a patient investigator, but more cheaply and faster.
The researchers carved the attack up into four stages, which they called Extract, Search, Reason, and Calibrate. At first, a language model gets identity clues from raw posts, like a hint of a job, a location, or a way of phrasing.

Semantic embeddings then narrow a pool of up to 89,000 candidates to a shortlist. A reasoning model grades the best matches and decides if two accounts are the same person.
The Calibrate step causes the pipeline to abstain if not certain. Fewer false positives.
The pipeline runs on off-the-shelf tools, web search, embeddings, and models, e.g., GPT-5.2. Previous deanonymization attacks, such as the 2008 re-identification of the anonymized Netflix ratings, relied on structured data.
The pipeline identified 226 of 338 Hacker News users at 90% precision
To assess the method without risking real people, the team assembled 338 Hacker News users with bios that pointed to a LinkedIn page. Each had a known ground-truth identity.
Fed only comments and submissions, the agent called 226 correctly, with about 67% at 90% precision. The agent got 25 calls wrong and passed on 86 accounts.
Classical non-LLM baselines got close to zero in the paper’s larger matching tests. The experiments cost less than $2,000 in total, or $1 to $4 per profile.
Two other datasets matched Reddit users across two different movie communities and rebuilt one Reddit history by splitting a single history into two time periods. And on each, LLM methods beat the older techniques by a large margin.
The authors cite governments targeting journalists or activists, companies producing ever-sharper ad profiles, and scammers building dossiers for social-engineering pitches.
“The combination is often a unique fingerprint,” Lermen wrote in a February 24 blog post. He added that if a team of smart investigators could identify someone from their posts, LLM agents likely can too, and the cost is only going down.
Researchers did not release the code, prompts, or any real identities turned up by the system. The study was reviewed by the ethics board of ETH Zurich before release.
The paper concludes that “the practical obscurity protecting pseudonymous users online no longer holds.”
The paper’s resurgence comes a day after SEC Commissioner Hester Peirce cautioned against bulk KYC data collection. Peirce said the approach amounts to “building bigger and bigger data haystacks.”
Recent breaches at Revolut and vendors tied to hardware-wallet maker Trezor have fueled fears of “wrench attacks,” in which someone who learns who holds crypto shows up in person.
The smartest crypto minds already read our newsletter. Want in? Join them.
0
0
Securely connect the portfolio you’re using to start.





