Online social networks have become essential spaces for communication and participation but also fertile grounds for online harms such as hate speech, harassment, and disinformation. These phenomena not only affect individuals through direct psychological harm but also undermine collective trust, polarize communities, and enable the manipulation of public discourse. To counter these risks, platforms deploy moderation pipelines that define ethical guidelines, detect harmful content, apply interventions, and evaluate their effectiveness. Yet, despite increasing sophistication, moderation remains fraught with challenges. Automated toxicity detectors inherit biases from human annotations, limiting their fairness and accuracy. Detection of disinformation still struggles to capture the coordination strategies that underpin large-scale information operations. Interventions are often generic or reactive, lacking contextual adaptation to the communities and conversations in which harms occur. Finally, evaluations of moderation are mostly retrospective, producing fragmented insights and overlooking long-term or unintended effects. This thesis addresses these gaps by advancing the study of moderation across its core components: detection, intervention, and evaluation. First, to better understand coordinated disinformation, we construct novel datasets from two major state-sponsored information operations on Twitter/X and apply network-science methods to disentangle malicious coordination from organic interactions. Our analysis reveals distinct strategies used by malicious actors and clarifies their role in shaping online debates. In parallel, we examine the biases embedded in hate speech detection by analyzing a large annotated dataset and comparing the behavior of human annotators with persona-driven large language models. This dual perspective shows how demographic attributes shape labeling decisions and how such biases are reproduced in automated systems. Together, these studies improve the reliability of detection while exposing its inherent limitations. Second, we move to the deployment of moderation interventions, introducing a framework for generating contextualized counterspeech with large language models. Unlike traditional approaches that rely on generic replies, our method adapts responses to the community context, the target user, and the surrounding conversation. Through both automated measures and a large-scale crowdsourcing study, we evaluate the qualities that make counterspeech more persuasive and effective, highlighting the potential of personalization in reducing online toxicity. Third, we examine the outcomes of hard moderation by analyzing the Reddit Great Ban, a large-scale deplatforming intervention. We provide causal evidence of both intended reductions in toxicity and unintended side effects, showing how some users escalated their harmful behaviors after the intervention. Building on these insights, we then shift from description to prediction by developing models that forecast moderation outcomes before deployment. By leveraging behavioral traces, we show that user abandonment can be anticipated, opening the way to proactive moderation strategies that balance effectiveness with community sustainability. Taken together, these contributions offer a systematic exploration of the moderation pipeline. By combining coordination analysis, bias auditing, personalized intervention design, and predictive evaluation, this thesis advances our understanding of how platforms can mitigate online harms while acknowledging the complexity and trade-offs inherent in moderation. The findings underscore both the potential and the limitations of algorithmic and socio-technical solutions, paving the way for more effective, context-aware, and ethically grounded approaches to governing digital communities.

ADDRESSING ONLINE HARMS THROUGH CONTENT MODERATION: DETECTION OF MISBEHAVIOR AND INTERVENTION STRATEGIES

CIMA, LORENZO
2026

Abstract

Online social networks have become essential spaces for communication and participation but also fertile grounds for online harms such as hate speech, harassment, and disinformation. These phenomena not only affect individuals through direct psychological harm but also undermine collective trust, polarize communities, and enable the manipulation of public discourse. To counter these risks, platforms deploy moderation pipelines that define ethical guidelines, detect harmful content, apply interventions, and evaluate their effectiveness. Yet, despite increasing sophistication, moderation remains fraught with challenges. Automated toxicity detectors inherit biases from human annotations, limiting their fairness and accuracy. Detection of disinformation still struggles to capture the coordination strategies that underpin large-scale information operations. Interventions are often generic or reactive, lacking contextual adaptation to the communities and conversations in which harms occur. Finally, evaluations of moderation are mostly retrospective, producing fragmented insights and overlooking long-term or unintended effects. This thesis addresses these gaps by advancing the study of moderation across its core components: detection, intervention, and evaluation. First, to better understand coordinated disinformation, we construct novel datasets from two major state-sponsored information operations on Twitter/X and apply network-science methods to disentangle malicious coordination from organic interactions. Our analysis reveals distinct strategies used by malicious actors and clarifies their role in shaping online debates. In parallel, we examine the biases embedded in hate speech detection by analyzing a large annotated dataset and comparing the behavior of human annotators with persona-driven large language models. This dual perspective shows how demographic attributes shape labeling decisions and how such biases are reproduced in automated systems. Together, these studies improve the reliability of detection while exposing its inherent limitations. Second, we move to the deployment of moderation interventions, introducing a framework for generating contextualized counterspeech with large language models. Unlike traditional approaches that rely on generic replies, our method adapts responses to the community context, the target user, and the surrounding conversation. Through both automated measures and a large-scale crowdsourcing study, we evaluate the qualities that make counterspeech more persuasive and effective, highlighting the potential of personalization in reducing online toxicity. Third, we examine the outcomes of hard moderation by analyzing the Reddit Great Ban, a large-scale deplatforming intervention. We provide causal evidence of both intended reductions in toxicity and unintended side effects, showing how some users escalated their harmful behaviors after the intervention. Building on these insights, we then shift from description to prediction by developing models that forecast moderation outcomes before deployment. By leveraging behavioral traces, we show that user abandonment can be anticipated, opening the way to proactive moderation strategies that balance effectiveness with community sustainability. Taken together, these contributions offer a systematic exploration of the moderation pipeline. By combining coordination analysis, bias auditing, personalized intervention design, and predictive evaluation, this thesis advances our understanding of how platforms can mitigate online harms while acknowledging the complexity and trade-offs inherent in moderation. The findings underscore both the potential and the limitations of algorithmic and socio-technical solutions, paving the way for more effective, context-aware, and ethically grounded approaches to governing digital communities.
19-mag-2026
Inglese
content moderation
intervention deployment
intervention evaluation
misbehavior detection
social media analysis
Avvenuti, Marco
Cresci, Stefano
File in questo prodotto:
File Dimensione Formato  
PhD_thesis_Lorenzo_Cima_2.pdf

accesso aperto

Licenza: Creative Commons
Dimensione 19.9 MB
Formato Adobe PDF
19.9 MB Adobe PDF Visualizza/Apri

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14242/377355
Il codice NBN di questa tesi è URN:NBN:IT:UNIPI-377355