Rubén Manrique

dblp:204/3652 · also Rubén Francisco Manrique · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-8742-2094ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 9 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Improving Low-Resource Translation with Dictionary-Guided Fine-Tuning and RL: A Spanish-to-Wayuunaiki Study
abstract
Low-resource machine translation remains a significant challenge for large language models (LLMs), which often lack exposure to these languages during pretraining and have limited parallel data for fine-tuning. We propose a novel approach that enhances translation for low-resource languages by integrating an external dictionary tool and training models end-to-end using reinforcement learning, in addition to supervised fine-tuning. Focusing on the Spanish–Wayuunaiki language pair, we frame translation as a tool-augmented decision-making problem in which the model can selectively consult a bilingual dictionary during generation. Our method combines supervised instruction tuning with Group Relative Policy Optimization (GRPO), enabling the model to learn both when and how to use the tool effectively. BLEU similarity scores are used as rewards to guide this learning process. Preliminary results show that our tool-augmented models achieve up to +3.37 BLEU improvement over previous work and an 18% relative gain compared to a supervised baseline without dictionary access, on the Spanish–Wayuunaiki test set from the AmericasNLP 2025 Shared Task. We also conduct ablation studies to assess the effects of model architecture and training strategy, comparing Qwen2.5-0.5B-Instruct with other models such as LLaMA and a prior NLLB-based system. These findings highlight the promise of combining LLMs with external tools and the role of reinforcement learning in improving translation quality in low-resource language settings.
Manuel Mosquera, Melissa Robles, Johan R. Portela, Rubén Manrique
AAAI4
2026 CFFitST: Classification few-shot fit sentence transformer
Daniel Fernando Gómez-Barrera, Luccas Rojas Becerra, Juan Pinzón Roncancio, David Ortiz Almanza, Juan Arboleda, Mario Linares-Vásquez, Rubén Manrique
Sci. Comput. Program.7
2025 AI-Generated Code Detection: An Examination of Current Tools in Education
Juan Esteban Cuellar Argotty, Rubén Manrique
ITS (1)2
2025 Manchita: An AI-Powered Gamified Learning Environment
Nicolás Klopstock, Ernesto José Duarte, Edier Becerra, Rubén Manrique
ITS (1)4
2025 AI-Powered Tutoring for Novice Programmers: Supporting Students Throughout Project Development
Juan Diego Lugo Sánchez, Brena Marques Ribeiro, Rubén Manrique, Kelly Garcés
ITS (1)3
2025 Evaluating Generalization Capabilities of LLM-Based Agents in Mixed-Motive Scenarios Using Concordia
abstract
Large Language Model (LLM) agents have demonstrated impressive capabilities for social interaction and are increasingly being deployed in situations where they might engage with both human and artificial agents. These interactions represent a critical frontier for LLM-based agents, yet existing evaluation methods fail to measure how well these capabilities generalize to novel social situations. In this paper, we introduce a method for evaluating the ability of LLM-based agents to cooperate in zero-shot, mixed-motive environments using Concordia, a natural language multi-agent simulation environment. Our method measures general cooperative intelligence by testing an agent's ability to identify and exploit opportunities for mutual gain across diverse partners and contexts. We present empirical results from the NeurIPS 2024 Concordia Contest, where agents were evaluated on their ability to achieve mutual gains across a suite of diverse scenarios ranging from negotiation to collective action problems. Our findings reveal significant gaps between current agent capabilities and the robust generalization required for reliable cooperation, particularly in scenarios demanding persuasion and norm enforcement.
Chandler Smith, Marwa Abdulhai, Manfred Diaz, Marko Tesic, Rakshit S. Trivedi, Alexander Vezhnevets, Lewis Hammond, Jesse Clifton, Minsuk Chang, Edgar A. Duéñez-Guzmán, John P. Agapiou, Jayd Matyas, Danny Karmon, Beining Zhang, Jim Dilkes, Akash Kundu, Emanuel Tewolde, Jebish Purbey, Ram Mohan Rao Kadiyala, Siddhant Gupta, Aliaksei Korshuk, Buyantuev Alexander, Ilya Makarov, Rolando Fernandez, Zhihan Wang, Caroline Wang, Jiaxun Cui, Lingyun Xiao, Yoonchang Sung, Muhammad Arrasy Rahman, Peter Stone 0001, Yipeng Kang, Hyeonggeun Yun, Ananya, Taehun Cha, Elizaveta Tennant, Olivia Macmillan-Scott, Marta Segura, Diana Riazi, Fuyang Cui, Sriram Ganapathi, Toryn Q. Klassen, Nico Schiavone, Mogtaba Alim, Sheila A. McIlraith, Manuel Ríos, Oswaldo Peña, Manuela Chacon-Chamorro, Rubén Manrique, Luis Felipe Giraldo, Nicanor Quijano, Fangwei Zhong, Wenming Tu, Zhaowei Zhang 0001, Zixia Jia, Zilong Zheng, Chichen Lin, Weijian Fan, Chenao Liu, Sneheel Sarangi, Shuqing Shi, Yali Du 0001, Avinaash Anand Kulandaivel, Yang Liu 0266, Ruiyang Wu 0007, Chetan Talele, Sunjia Lu, Gema Parreno, Shamika Dhuri, Bain McHale, Tim Baarslag, Dylan Hadfield-Menell, Natasha Jaques, José Hernández-Orallo, Joel Z. Leibo
NeurIPS54
2024 Analysis of Machine Learning Models for Academic Performance Prediction
Andres Benitez Amaya, Harold E. Castro, Rubén Manrique
ITS (2)3
2024 Preserving Heritage: Developing a Translation Tool for Indigenous Dialects
abstract
The preservation and understanding of indigenous languages emerge as crucial, given their substantial contribution to the cultural and linguistic heritage of communities. Despite their undeniable value, these languages are threatened by extinction due to a dwindling number of native speakers and the predominance of oral traditions over written forms. In this context, this study aims to contribute to the conservation of these languages through the development of a Spanish-indigenous language translator. This research employs neural machine translation technology, investigating three distinct approaches: a translation model based on transformers, finetuning with a Finnish translator, and finetuning with a multilingual translator. The results obtained from these methodologies are promising, demonstrating competitive viability when compared to the limited existing research in this field of study.
Melissa Robles, Cristian A. Martínez, Juan Camilo Prieto, Sara Palacios, Rubén Manrique
WSDM5
2023 HABLA: A Dataset of Latin American Spanish Accents for Voice Anti-spoofing
Pablo Andrés Tamayo Flórez, Rubén Manrique, Bernardo Pereira Nunes
INTERSPEECH2
2022 Log mining for course recommendation in limited information scenarios
Juan Camilo Sanguino, Rubén Manrique, Olga Mariño, Mario Linares, Nicolás Cardozo
EDM2
2022 Not Another Hardcoded Solution to the Student Dropout Prediction Problem: A Novel Approach Using Genetic Algorithms for Feature Selection
Yixin Cheng, Bernardo Pereira Nunes, Rubén Manrique
ITS3
2020 'A Little Knowledge is a Dangerous Thing': A method to automatically detect knowledge compartmentalization and oversimplification
abstract
Simplification is a common practice to allow students to understand complex concepts. However, such practice may lead to oversimplification and knowledge compartmentalization problems. This paper proposes a method to minimize the effects of oversimplification and knowledge compartmentalization through a five-step processing chain, following instructional design principles, to foster advanced knowledge acquisition. The proposed method was applied to online courses and a discussion for an example of a Data Science course is provided to show its applicability.
Crystiam Kelle Pereira, Bernardo Pereira Nunes, Sean W. M. Siqueira, Rubén Manrique, Jerry Fernandes Medeiros
ICALT4
2019 Towards the Identification of Concept Prerequisites Via Knowledge Graphs
abstract
Learning basic concepts before complex ones is a natural form of learning. This paper addresses the specific problem of identifying concept prerequisites to inform about the basic knowledge required to understand a particular concept. Briefly, given a target concept c, the goal is to (a) find candidate concepts in a Knowledge Graph (KG) that serve as possible prerequisite for c; and, (b) evaluate the prerequisite relation between the target and candidates concepts via a supervised learning model. Our approach explores the DBpedia Knowledge Graph and its semantic relations to find candidate concepts as well as a pruning step to reduce the candidate concept set. Finally, we employ supervised learning algorithms to evaluate and generate a list of prerequisites for the target concept. A ground truth created based on expert knowledge is used to validate our approach, exhibiting promising results with a precision varying between 83% and 92.9%.
Rubén Manrique, Bernardo Pereira Nunes, Olga Mariño, Nicolás Cardozo, Sean W. M. Siqueira
ICALT1
2019 Using Query Reformulation to Compare Learning Behaviors in Web Search Engines
abstract
Web search engines have gained importance as tools capable of connecting informal and self-learning with formal learning by aiding individuals in retrieving relevant information through the formulation and modification of their queries. Understand the differences between query states and their transitions becomes increasingly important, as doing so makes the optimization of search engines' results according to educational uses and needs possible. This paper introduces the ESKiP Taxonomy of Query States, a classification framework validated in an experiment involving two different query log datasets. It enables the comparison between the behaviors of users in search for knowledge (learners) and users performing transactional or factual searches in Web search engines.
Marcelo Tibau, Sean W. M. Siqueira, Bernardo Pereira Nunes, Terhi Nurmikko-Fuller, Rubén Manrique
ICALT5
2019 An Analysis of Student Representation, Representative Features and Classification Algorithms to Predict Degree Dropout
abstract
Identifying and monitoring students who are likely to dropout is a vital issue for universities. Early detection allows institutions to intervene, addressing problems and retaining students. Prior research into the early detection of at-risk students has opted for the use of predictive models, but a comprehensive assessment of the suitability of different algorithms and approaches is complicated by the large number of variable features that constitute a student's educational experience. Predictive models vary in terms of their amplitude, temporality and the learning algorithms employed. While amplitude refers to the ability of the model to operate on multiple degrees, temporality is often considered due to the natural temporal aspect of the data. In the absence of a comparative framework of learning algorithms, the aim of this paper has been to provide such an analysis, based on a proposed classification of strategies for predicting dropouts in Higher Education Institutions. Three different student representations are implemented (namely Global Feature-Based, Local Feature-Based, and Time Series) in conjunction with the appropriate learning algorithms for each of them. A description of each approach, as well as its implementation process, are presented in this paper as technical contributions. An experiment based on a dataset of student information from two degrees, namely Business Administration and Architecture, acquired through an automated management system from a university in Brazil is used. Our findings can be summarized as: (i) of the three proposed student representations, the Local Feature-Based was the most suitable approach for predicting dropout. In addition to providing high quality results, the Local Feature-Based representations are simple to build, and the construction of the model is less expensive when compared to more complex ones; (ii) as a conclusion of the results obtained via Local Feature-Based, dropout can be said to be accurately predicted using grades of a few core courses, so there is no need for a complex features extraction process; (iii) considering temporal aspects of the data does not seem to contribute to the prediction performance although it increases computational costs as the model complexity increases.
Rubén Manrique, Bernardo Pereira Nunes, Olga Mariño, Marco A. Casanova, Terhi Nurmikko-Fuller
LAK1
2018 Investigating Learning Resources Precedence Relations via Concept Prerequisite Learning
abstract
The identification of prerequisite relationships among concepts is a fundamental step toward the organization of knowledge for educational purposes. In the context of a learning process, simplest concepts that are requirements to understand and address more complex concepts should be presented first. Therefore, the identification of prerequisite relationships is a fundamental step for effective course design and automatic learning path generation systems. Although there have been recent advances in machine learning methods for the automatic identification of prerequisite relationships between concepts, little research has been done on whether these automatic strategies can be extended to establish precedence relationships among learning resources. The precedence relation between two learning resources establishes which of the resources must be presented first. In this paper, we approach this problem and propose a strategy to identify the precedence relation. Given two learning resources our strategy analyzes prerequisites among the concepts addressed by the learning resources to estimate the precedence relation. A set of 1588 pairs of learning resources extracted from MOOCs refined by human experts is used to evaluate the strategy. The experimental results show that it is possible to identify the precedence relation between learning resources through the automatic identification of prerequisite relationships between concepts.
Rubén Manrique, Juan Sebastián Sosa, Olga Mariño, Bernardo Pereira Nunes, Nicolás Cardozo
WI1
2017 Towards automatic learning content sequence via linked open data
abstract
The paradigm of lifelong learning supported by technology is redefining the way we learn as well as the way we search and consume the ever growing corpus of information available in the Web to acquire knowledge on a particular subject. This research addresses the problem of finding and organizing learning content to support self-directed learners in achieving a learning goal through the search, selection and sequencing of Web content that might or might not have been conceived as learning resources. We plan to build an automatic process driven by the knowledge available in datasets belonging to the Linked Open Initiative and open non-structured information such as courses syllabi and books table of contents. Our proposed service have two main components: (i) a graph of interrelated learning concepts from which is possible infer what concepts must be addressed first before others in the learning process (prerequisite relationships), and (ii) a component for the creation of learning resources sequences based on a learning goal and a learner profile.
Rubén Manrique
WI1
2017 How does the size of a document affect linked open data user modeling strategies?
abstract
Semantic user modeling techniques use a representation based on concepts that are linked to a Knowledge Base (KB). Current research uses Linked Open Data (LOD) because of its comprehensive interlinked datasets, which allow excellent cross-domain modeling capabilities. LOD semantic user profiles have been employed in the context of Social Networks, which require user's posts as input. Less attention has been paid to other domains for which input documents differ from short and concise posts. In this paper, we perform a comparative study of different LOD semantic user modeling techniques by taking different types of documents as input: short, medium, and long texts. We selected recommending academic documents based on modeling the user's research interests as the evaluation scenario. Academic documents' titles, abstracts, and the body of text were used, respectively, for short, medium, and long documents. Our results showed that expansion strategies work best for short and medium documents while filtering strategies are more appropriate when the whole document is used as input. Finally, we explored diverse alternatives if documents did not include a summary or abstract, and we concluded that, in this case, the two best alternatives are a filtering strategy over the whole text and the use of TextRank algorithm to build a set of key sentences to be used as input of an expansion strategy.
Rubén Manrique, Olga Mariño
WI1