Cross-Linguistic Overlap English-Spanish Vocabulary: A Data-Driven Approach to Lexical Similarity
Keywords:
Cross-language similarities, shared vocabulary, cognates, loanwords, second language acquisitionAbstract
This study presents a computational methodology for identifying shared vocabulary between English and Spanish to support second language acquisition. Similarity indexes (S.I.) were calculated to determine the percentage of orthographic and semantic overlap between the two languages. Both Spanish and English use the Roman script, which allowed for string-based comparison. English has received multiple loanwords from Latin, from which Spanish also derives. This study analyses the most frequent vocabulary in both languages to assess the similarity level and extract similar lexical items. The results show that a lexical similarity level of 53.13% corresponding to 1594 shared lexical terms was calculated using a lexico-statistical computational method to compare the 3000 highest frequency terms of the basic vocabulary of English and Spanish, identifying shared vocabulary as a pedagogical tool.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Maria Isabel Maldonado Garcia (Translator)

This work is licensed under a Creative Commons Attribution 4.0 International License.
Linguistic Forum – A Journal of Linguistics is an open-access journal. All articles are published under the Creative Commons Attribution 4.0 International (CC BY 4.0) License. Authors retain copyright while granting the journal the right of first publication. The CC BY 4.0 License permits unrestricted use, distribution, reproduction, adaptation, and commercial reuse in any medium, provided the original work is properly cited and appropriate credit is given to the author(s) and the journal.
