A Framework for Evaluating AI-Enabled Social Engineering
Large language models (LLMs) can now produce persuasive, personalized phishing content at scale. Yet there is no systematic, reusable benchmark for measuring how this offensive capability varies across models or relates to general capability. We introduce ScamBench: a benchmark and evaluation pipeline for LLM-generated spear phishing. ScamBench contributes (i) a public dataset of over 16,000 personalized phishing emails generated by 20 frontier and open-weight models against 150 synthetic target profiles spanning diverse occupations, life stages, and levels of technical sophistication; (ii) a methodology in which simulated recipients and real human users predict behavioral responses to each email; and (iii) a public leaderboard ranking models by how often their generated emails persuade the target to click a link. Across the 20 models, Claude Opus 4.5 achieves the highest click-through rate at 50.3\%, while even the weakest model elicits clicks at 29.8\%. Together, our findings establish AI-enabled social engineering as an empirically measurable capability that scales with frontier models, with ScamBench providing a reusable benchmark for tracking it as models continue to advance.
Fred Heiding
Fred Heiding is the executive director of Menlo Park Intelligence and a researcher at UC Berkeley’s Center for Long-Term Cybersecurity. He was previously a doctoral and postdoctoral researcher at the Harvard School of Engineering and Applied Sciences and the Harvard Kennedy School, working with Bruce Schneier and Eric Rosenbach. He is a member of the World Economic Forum’s Centre for Cybersecurity and the Geneva Centre for Security Policy, and serves on the committee of the Harvard and MIT Technology and National Security Conference. He has taught several AI and cybersecurity classes at Harvard, regularly briefs the U.S. Congress on AI-powered cyber threats, and advises foreign governments on national cyber risks. His work has been featured at leading conferences, journals, and media outlets, including The Economist, Reuters, TIME, and Foreign Affairs. Fred has assisted in the discovery of more than 45 critical computer vulnerabilities (CVEs), and he previously made headlines for hacking the King of Sweden and the European Commissioner.