Modern NLP models often struggle with the non-standard language and misspellings in user-generated content. This thesis investigates misspelling phenomena through five complementary perspectives: I. Background: A comprehensive review of NLP strategies for handling errors, from automatic correction and robust neural architectures to recent advances in Large Language Models (LLMs). II. Linguistic: An analysis of how unintentional misspellings reflect a writer’s cognitive, cultural, or social identity. III. Orthographic: An evaluation of orthographic robustness, testing whether NLP models can replicate the human ability to comprehend scrambled words. IV. Phonetic: Using a purpose-built dataset of phonetic variants to assess performance impacts on linguistic models and LLMs. V. Visual: An investigation into misspellings based on graphical similarities (homoglyphs like "rn" vs "m"). This section evaluates the efficacy of visually-grounded character embeddings in neural networks. Taken together, these perspectives underscore the multifaceted nature of misspellings in NLP and highlight the value of interdisciplinary approaches in addressing them. While no single strategy proves sufficient on its own, integrating insights from multiple specialized domains offers promising directions for future research on building more robust and cognitively inspired NLP systems.

Misspellings in Natural Language Processing: An Interdisciplinary Investigation

SPERDUTI, GIANLUCA
2026

Abstract

Modern NLP models often struggle with the non-standard language and misspellings in user-generated content. This thesis investigates misspelling phenomena through five complementary perspectives: I. Background: A comprehensive review of NLP strategies for handling errors, from automatic correction and robust neural architectures to recent advances in Large Language Models (LLMs). II. Linguistic: An analysis of how unintentional misspellings reflect a writer’s cognitive, cultural, or social identity. III. Orthographic: An evaluation of orthographic robustness, testing whether NLP models can replicate the human ability to comprehend scrambled words. IV. Phonetic: Using a purpose-built dataset of phonetic variants to assess performance impacts on linguistic models and LLMs. V. Visual: An investigation into misspellings based on graphical similarities (homoglyphs like "rn" vs "m"). This section evaluates the efficacy of visually-grounded character embeddings in neural networks. Taken together, these perspectives underscore the multifaceted nature of misspellings in NLP and highlight the value of interdisciplinary approaches in addressing them. While no single strategy proves sufficient on its own, integrating insights from multiple specialized domains offers promising directions for future research on building more robust and cognitively inspired NLP systems.
3-mag-2026
Inglese
nlp, errori ortografici, errori, refusi
nlp, misspellings, errors, noise
Moreo Fernández, Alejandro
Sebastiani, Fabrizio
File in questo prodotto:
File Dimensione Formato  
PhdThesisSperduti.pdf

embargo fino al 05/05/2029

Licenza: Creative Commons
Dimensione 6.38 MB
Formato Adobe PDF
6.38 MB Adobe PDF

I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.14242/367070
Il codice NBN di questa tesi è URN:NBN:IT:UNIPI-367070