Contaduria y Administracion | 2026
Authors: Cuevas J.P.; Cuevas-Rasgado A.D.; Reyes-Ortiz J.A.; Bravo M.
DOI: 10.22201/fca.24488410e.2026.5282
Journal: Contaduria y Administracion
Year: 2026
Publisher: Universidad Nacional Autonoma de Mexico
Document Type: Article
Open Access: All Open Access; Gold Open Access
Cited by: 0
The scientific papers are written in natural language, with a significant proportion being in Spanish, and have no structure processable by computers, which results in tedious and time-consuming manual analysis. Thus, managing scientific texts in Spanish is a challenge that requires advanced computational methods. Therefore, this paper presents a novel methodology that includes three Information Retrieval (IR) approaches based on Natural Language Processing (NLP). The main aim is the information management from scientific documents in Spanish. The IR approaches implemented in the methodology are based on textual, probabilistic, and semantic similarity to retrieve documents regarding a question. The proposed methodology is applied to the scientific Spanish literature generated during the COVID-19 pandemic. An evaluation process based on 100 queries over 249,474 scientific documents to accurate the recoverability of relevant documents was carried out. The results show that the probabilistic approach implemented in the methodology achieved an 85% f-measure, supported by the Latent Dirichlet Allocation (LDA) topic discovery algorithm. Finally, the proposed methodology is considered domain- independent to retrieve documents in Spanish. © 2019 Universidad Nacional Autónoma de México, Facultad de Contaduría y Administración. This is an open access article under the CC BY-NC-SA (https://creativecommons.org/licenses/by-nc-sa/4.0/)
information management; natural language processing in spanish texts; recovering scientific documents; semantic and probabilistic approaches; text-similarity