Building Trust Infrastructure for Human-in-the-Loop Adversarial ML Research: Lessons from a Multi-Channel Recruitment Pilot for AI-Phishing Benchmarks
Camilla Balbis
Mentored by Fred Heiding
Working report from the SPAR program. May not reflect the authors' current views.
Abstract
Adversarial machine learning benchmarks that evaluate large language model capabilities, such as generating persuasive phishing emails, require ongoing human-in-the-loop evaluation to produce ecologically valid data. Yet no published methodology exists for recruiting voluntary human subjects for deception-based adversarial research, where the research topic itself may trigger defensive reactions that undermine participation. This report documents a three-month, zero-budget pilot across sixteen Canadian universities that tested ten messaging strategies, four outreach channels, and a peer-network incentive model to recruit participants for ScamBench, a human-validated phishing benchmark developed by The AI and Cybersecurity Institute (TAICI). The pilot generated 449 unique contacts and secured twelve partnerships, fifteen newsletter mentions, and one high-impact Student Ambassador agreement. Key findings challenge three assumptions in the phishing and adversarial ML literatures. First, the most technically relevant audiences, cybersecurity and computer science clubs, converted at 0%, while general STEM and women-in-STEM groups achieved 8-11% conversion, suggesting that domain expertise may increase threat perception and self-selection bias rather than willingness to participate. Second, institutional newsletters outperformed direct club outreach (9.1% vs. 2.7% conversion), demonstrating that formal broadcast channels with editorial credibility can exceed targeted technical communities in recruitment efficiency. Third, a Student Ambassador model that delegates trust-building to existing peer networks projects 500-1,000 signups from a single partnership, offering a scalable alternative to the scalability-quality tradeoff observed between manual outreach (4.7% conversion) and automated scraping (1.7%). The analysis draws on the phishing user-study literature-where recruitment is frequently the weakest methodological link and participant priming significantly alters behaviour into frame recruitment design as a governance variable rather than a logistical footnote. The findings align with peer-network theories of trust-based participation while adding a new, counterintuitive result: highly technical student communities were not the most effective recruitment targets for adversarial phishing research. Recommendations include deprioritizing faculty channels (0/53 conversion), leading with "AI Safety" rather than project-specific names in cold outreach, budgeting 6-12 months for trust-building infrastructure, and investing in career-oriented incentive structures for student demographics.