GBU College Logo
Sign in with GBU Microsoft

Methodology for the management of scientific literature in Spanish through information retrieval approaches using natural language processing; [Metodología para la gestión de la literatura científica en español mediante enfoques de recuperación de información utilizando el procesamiento del lenguaje natural]

Contaduria y Administracion | 2026

Paper Details

Authors: Cuevas J.P.; Cuevas-Rasgado A.D.; Reyes-Ortiz J.A.; Bravo M.

DOI: 10.22201/fca.24488410e.2026.5282

Journal: Contaduria y Administracion

Year: 2026

Publisher: Universidad Nacional Autonoma de Mexico

Document Type: Article

Open Access: All Open Access; Gold Open Access

Cited by: 0

Abstract

The scientific papers are written in natural language, with a significant proportion being in Spanish, and have no structure processable by computers, which results in tedious and time-consuming manual analysis. Thus, managing scientific texts in Spanish is a challenge that requires advanced computational methods. Therefore, this paper presents a novel methodology that includes three Information Retrieval (IR) approaches based on Natural Language Processing (NLP). The main aim is the information management from scientific documents in Spanish. The IR approaches implemented in the methodology are based on textual, probabilistic, and semantic similarity to retrieve documents regarding a question. The proposed methodology is applied to the scientific Spanish literature generated during the COVID-19 pandemic. An evaluation process based on 100 queries over 249,474 scientific documents to accurate the recoverability of relevant documents was carried out. The results show that the probabilistic approach implemented in the methodology achieved an 85% f-measure, supported by the Latent Dirichlet Allocation (LDA) topic discovery algorithm. Finally, the proposed methodology is considered domain- independent to retrieve documents in Spanish. © 2019 Universidad Nacional Autónoma de México, Facultad de Contaduría y Administración. This is an open access article under the CC BY-NC-SA (https://creativecommons.org/licenses/by-nc-sa/4.0/)

Keywords

information management; natural language processing in spanish texts; recovering scientific documents; semantic and probabilistic approaches; text-similarity