Publication Details
Abstract
Artificial intelligence (AI) and Natural Language Processing (NLP) have transformed lexicographic practices by enabling scalable and automated language analysis. While these advancements allow for the efficient processing of linguistic data, they often fall short in accurately capturing and representing culturally embedded concepts, particularly in underrepresented languages like Uzbek. Despite progress in semantic models and named entity recognition, there remains a significant lack of cultural sensitivity in AI-based lexicographic systems, primarily due to limited annotated corpora and challenges in modeling context-specific meanings. This study examines how NLP techniques—specifically tokenization, word embedding, and named entity recognition—function in representing cultural concepts in English and Uzbek, and evaluates their strengths and limitations. Findings show that while English-language models handle idiomatic and cultural terms with moderate success, models for Uzbek exhibit considerable deficiencies due to morphological complexity and corpus scarcity. Both languages face issues with accurately capturing idiomatic expressions and culturally loaded entities, leading to semantic distortion in automated outputs. The paper introduces a comparative framework grounded in semantic theory and cultural linguistics, providing practical examples of misrepresentation and highlighting the need for culturally annotated corpora and cross-cultural NLP modeling. To achieve culturally competent AI lexicography, interdisciplinary collaboration is essential. Future systems must integrate domain-specific resources, cultural annotations, and linguistic diversity to ensure that AI technologies do not reduce language to mechanistic processing but preserve its cultural and emotional richness.