Twitter defines four types of spam behavior in their policy: commercially-motivated spam, inauthentic engagements, coordinated activity, and coordinated harmful activity. These violations can be enacted by promotional accounts, false/fake accounts, bots, malicious accounts (trolls), and compromised accounts.
This report examines the prevalence of spam accounts on Twitter, evaluates Twitter's spam detection processes, and analyzes the effects of spam accounts on the platform. We will use multiple methodologies to estimate the true rate of spam accounts and assess whether Twitter's reported figures are accurate.
Sample mDAU accounts, identify spam, and extrapolate to full list
Create machine learning model to distinguish spam from legitimate users
Apply Twitter's Smyte and Botmaker workflows to mDAU data
Use state-of-the-art tools like Botometer to establish baseline
Re-estimated mDAU spam account rate of 10.32% (vs Twitter's 0.72%)
8.01% of active mDAU accounts predicted to be spam
Studies consistently report 9-21% spam accounts
Spam accounts nearly twice as active as non-spam accounts
Process collapses uncertainty toward "Good" classification
Evaluators lack access to crucial non-public account data
Insufficient investigation of coordinated inauthentic behavior
Overlooks common cross-platform amplification tactics
Cryptocurrency schemes, counterfeit products, phishing
Coordinated efforts to influence market perceptions
Distortion of discourse, amplification of false narratives
Spread of unverified medical claims, conspiracy theories
Regularly revise to address evolving spam tactics
Implement tools to identify inauthentic networks
Regularly evaluate and retrain spam detection staff
Consider off-Twitter activity in spam determination
The higher prevalence of spam accounts has significant implications for Twitter's advertising-based business model. Spam accounts generate a disproportionate amount of activity, potentially inflating engagement metrics presented to advertisers.
Additionally, the presence of spam accounts degrades the user experience for legitimate users, potentially impacting user retention and growth. Twitter must balance aggressive spam removal with the risk of falsely flagging legitimate accounts.
Our analysis suggests that Twitter's reported spam account rate significantly underestimates the true prevalence of spam on the platform. The company's spam evaluation process has several flaws that lead to systematic undercounting of spam accounts.
To address these issues, Twitter should:
Implement more robust spam detection techniques
Provide more detailed reporting on spam account prevalence
Develop advanced AI/ML tools for spam detection
Problem Statement: Prevalence of Spam Accounts on Twitter