Modern NLP models often struggle with the non-standard language and misspellings in user-generated content. This thesis investigates misspelling phenomena through five complementary perspectives: I. Background: A comprehensive review of NLP strategies for handling errors, from automatic correction and robust neural architectures to recent advances in Large Language Models (LLMs). II. Linguistic: An analysis of how unintentional misspellings reflect a writer’s cognitive, cultural, or social identity. III. Orthographic: An evaluation of orthographic robustness, testing whether NLP models can replicate the human ability to comprehend scrambled words. IV. Phonetic: Using a purpose-built dataset of phonetic variants to assess performance impacts on linguistic models and LLMs. V. Visual: An investigation into misspellings based on graphical similarities (homoglyphs like "rn" vs "m"). This section evaluates the efficacy of visually-grounded character embeddings in neural networks. Taken together, these perspectives underscore the multifaceted nature of misspellings in NLP and highlight the value of interdisciplinary approaches in addressing them. While no single strategy proves sufficient on its own, integrating insights from multiple specialized domains offers promising directions for future research on building more robust and cognitively inspired NLP systems.
Misspellings in Natural Language Processing: An Interdisciplinary Investigation
SPERDUTI, GIANLUCA
2026
Abstract
Modern NLP models often struggle with the non-standard language and misspellings in user-generated content. This thesis investigates misspelling phenomena through five complementary perspectives: I. Background: A comprehensive review of NLP strategies for handling errors, from automatic correction and robust neural architectures to recent advances in Large Language Models (LLMs). II. Linguistic: An analysis of how unintentional misspellings reflect a writer’s cognitive, cultural, or social identity. III. Orthographic: An evaluation of orthographic robustness, testing whether NLP models can replicate the human ability to comprehend scrambled words. IV. Phonetic: Using a purpose-built dataset of phonetic variants to assess performance impacts on linguistic models and LLMs. V. Visual: An investigation into misspellings based on graphical similarities (homoglyphs like "rn" vs "m"). This section evaluates the efficacy of visually-grounded character embeddings in neural networks. Taken together, these perspectives underscore the multifaceted nature of misspellings in NLP and highlight the value of interdisciplinary approaches in addressing them. While no single strategy proves sufficient on its own, integrating insights from multiple specialized domains offers promising directions for future research on building more robust and cognitively inspired NLP systems.| File | Dimensione | Formato | |
|---|---|---|---|
|
PhdThesisSperduti.pdf
embargo fino al 05/05/2029
Licenza:
Creative Commons
Dimensione
6.38 MB
Formato
Adobe PDF
|
6.38 MB | Adobe PDF |
I documenti in UNITESI sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/20.500.14242/367070
URN:NBN:IT:UNIPI-367070