Juan Manuel Rodriguez

dblp:22/8690 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-1130-8065ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 10 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Deep Learning Techniques for Text-Based Emotional Response Generation: A Systematic Review
abstract
This work presents a systematic literature review of deep learning techniques for text-based emotional response generation, a particular topic in the area of emotional dialogue systems. In particular, this work focus on systems that support text-based interaction where the system detects the user's emotion and responds accordingly. We have scouted over 6000 articles for this review, selecting 74 of them and reviewing them to identify datasets, methods, evaluation techniques, and challenges in the area. Our analysis points out an overwhelming bias toward researching these problems in English and Chinese, with few to no resources for other languages. Although several methodologies are proposed in the literature, including Seq2Seq and Conditional Variational Autoencoders, more recent research has been focused on the Transformer architecture. It is recognized that there are manual metrics that do not have an equivalent to automatic metrics, and there are variations in human evaluation results in comparison to modern automated metrics. With this paper, we aim to summarize the state-of-the-art of deep learning techniques for emotional response generation while identifying open research questions.
Germán Lescano, Juan Manuel Rodriguez, Rosanna Costaguta
IEEE Trans. Affect. Comput.2
2025 SpanishTweetsCOVID-19: A Social Media Enriched Covid-19 Twitter Spanish Dataset
Antonela Tommasel, Juan Manuel Rodriguez
IEEE Big Data2
2025 PrivEval: a tool for interactive evaluation of privacy metrics in synthetic data generation
abstract
Synthetic data generation (SDG) is the process of generating a new synthetic dataset based on the statistical properties of a confidential existing dataset. Differential privacy is the property of a SDG mechanism that establishes how protected individuals whose sensitive data is part of the confidential dataset are, when sharing such data. To ensure a SDG is differentially private, noise is injected into the statistics learned from the dataset. Depending on the amount of noise injected, we witness a trade-off between privacy and utility. Privacy is then measured via a set of privacy metrics that usually establish a lower bound on a few aspects of the privacy-utility trade-off. Therefore, it is not possible to assess privacy based only on one metric. To close this gap, we demonstrate PrivEval, a tool to assist users in evaluating the privacy properties of a synthetic dataset. PrivEval implements several privacy metrics and validates them on both a single user and the overall dataset. Besides, PrivEval checks assumptions behind each metric. Hence, PrivEval is a first step to bridge the gap between privacy experts and the general public to make privacy estimation more transparent.
Frederik M. Trudslev, Matteo Lissandrini, Juan Manuel Rodriguez, Martin Bøgsted, Daniele Dell'Aglio
Proc. VLDB Endow.3
2025 Countering the Spread: An Approach to Identify Misinformation Spreaders in Social Media
abstract
Although social media generally provides a safe and enjoyable experience, it can also serve as a quick and easy means for spreading false news, misinformation, and other harmful content. These contents have been proven effective in influencing people’s beliefs and behaviors, spanning from influencing political opinions to directly impacting public health, particularly during events such as the COVID-19 pandemic. Then, it becomes crucial to take proactive measures to identify the misinformation spreaders to mitigate their impact and influence over society. Existing approaches primarily focus on identifying spreaders by analyzing features such as writing style, content, user profiles, and engagement statistics. However, since fake or deceiving content is frequently crafted to closely resemble authentic information, traditional techniques alone prove insufficient to effectively identify either fake content or its spreaders. In this context, this work introduces a deep-learning model tailored for detecting misinformation spreaders in social media. Our model not only incorporates content-based features but also integrates patterns of social interactions and information propagation structures. By considering the multifaceted nature of misinformation spread, our approach provides a more holistic and accurate means of identifying its spreaders. An experimental evaluation focusing on COVID-related data yielded promising results, demonstrating a significant performance improvement compared to other techniques in the literature. Thus, this research contributes to the ongoing efforts to develop robust tools for reducing the adverse effects of misinformation in social media.
Antonela Tommasel, Juan Manuel Rodriguez
IEEE Trans. Comput. Soc. Syst.2
2024 Beyond Words: A Preliminary Study for Multimodal Hate Speech Detection
abstract
Hate speech involves harmful expressions that either directly attack or endorse hatred toward a group or individual, often based on identity characteristics such as ethnicity, religion, or sexual orientation. The lack of regulation regarding hate speech raises concerns about the vulnerability of targeted groups and the potential escalation of violence. Detecting and eliminating hate speech from social media is crucial to prevent harm, negative psychological effects, and further attacks. The complexity of identifying hate speech lies in the indirect or implicit nature of its expressions, which can be influenced by context and the use of sarcasm or irony. Detection becomes even more challenging in multimodal scenarios, such as memes, where text and images are combined to convey meaning. To address this problem, this study analyzes the design of a multimodal hate speech detection technique that integrates textual and visual dimensions. This study explores how to define the textual and visual representations of memes and explores fusion strategies. The preliminary experiments conducted over publicly available datasets achieved encouraging results when compared to more complex approaches in the literature, while also highlighting the challenges of the task.
Sofía Barceló, Magalí Boulanger, Antonela Tommasel, Juan Manuel Rodriguez
CLEI4
2024 Does the Performance of Text-to-Image Retrieval Models Generalize Beyond Captions-as-a-Query?
Juan Manuel Rodriguez, Nima Tavassoli, Eliezer Levy, Gil Lederman, Dima Sivov, Matteo Lissandrini, Davide Mottin
ECIR (4)1
2022 Tracking the evolution of crisis processes and mental health on social media during the COVID-19 pandemic
abstract
The COVID-19 pandemic has affected all aspects of society, bringing health hazards and posing challenges to public order, governments, and mental health. This study examines the stages of crisis response and recovery as a sociological problem by operationalising a well-known model of crisis stages in terms of a psycho-linguistic analysis. Based on an extensive collection of Twitter data spanning from March to August 2020 in Argentina, we present a thematic study on the differences in language used in social media posts and look at indicators that reveal the distinctive stages of a crisis and the country response thereof. The analysis was combined with a study of the temporal prevalence of mental health related conversations and emotions. This approach can provide insights for public health policy design to monitor and eventually intervene during the different stages of a crisis, thus improving the adverse mental health effects on the population.
Antonela Tommasel, Jorge Andrés Díaz Pace, Daniela Godoy, Juan Manuel Rodriguez
Behav. Inf. Technol.4
2021 I Want to Break Free! Recommending Friends from Outside the Echo Chamber
abstract
Recommender systems serve as mediators of information consumption and propagation. In this role, these systems have been recently criticized for introducing biases and promoting the creation of echo chambers and filter bubbles, thus lowering the diversity of both content and potential new social relations users are exposed to. Some of these issues are a consequence of the fundamental concepts on which recommender systems are based on. Assumptions like the homophily principle might lead users to content that they already like or friends they already know, which can be naïve in the era of ideological uniformity and fake news. A significant challenge in this context is how to effectively learn the dynamic representations of users based on the content they share and their echo chamber or community interactions to recommend potentially relevant and diverse friends from outside the network of influence of the users’ echo chamber. To address this, we devise FRediECH (a Friend RecommenDer for breakIng Echo CHambers), an echo chamber-aware friend recommendation approach that learns users and echo chamber representations from the shared content and past users’ and communities’ interactions. Comprehensive evaluations over Twitter data showed that our approach achieved better performance (in terms of relevance and novelty) than state-of-the-art alternatives, validating its effectiveness.
Antonela Tommasel, Juan Manuel Rodriguez, Daniela Godoy
RecSys2
2020 DewSim: A trace-driven toolkit for simulating mobile device clusters in Dew computing environments
abstract
Summary Dew computing is an emerging computing paradigm, which aims at minimizing the dependency over existing internetwork back‐haul, ie, being dependent on processing resources offered by remote servers. Smartphones and tablets ubiquity and powerful computing hardware motivated researchers to investigate the way of providing Dew computing services by exploiting the aggregated capabilities of devices in a vicinity, a smart device cluster. Consequently, research on resource management is necessary to learn how to scavenge resources from such a cluster, deal with devices heterogeneity, limitations, and dynamic resource availability. Simulation is commonly practiced for studying resource management in other distributed computing research fields, specially due to the complexity involved in the set up of experiments. However, a free‐to‐use purpose specific toolkit for studying smart device clusters do not exist or have been documented. Current simulation efforts do not allow researchers to faithfully represent key singularities of such environment, which are energy depletion and nondedicated nature of computing resources. We propose a trace‐based toolkit built on modular software artifacts to speed up research in resource management techniques in Dew environments. A trace‐driven methodology is adopted to assure practical value of simulated scenarios. The toolkit comprises a device profiler application for Android to capture generic battery and CPU traces from real devices, a profile mixer to create user interaction baseline traces through generic ones, and an extensible engine to simulate the execution of workloads configurable via text files. Verification and validation tests were run to show correctness and reliability of our simulation approach.
Matías Hirsch, Cristian Mateos, Juan Manuel Rodriguez, Alejandro Zunino
Softw. Pract. Exp.3
2017 An empirical evaluation of a simple energy aware scheduler for mobile grids
abstract
Nowadays mobile devices are multi-core computers with considerable unused capabilities. Therefore, several researchers have considered harnessing the power of these battery-powered devices for distributed computing. Although their evergrowing capabilities, the fact that mobile devices run on battery poses a major challenge for applying traditional distributed computing techniques. Particularly, researchers aimed at using mobile devices as resources for executing computationally intensive task. Different job scheduling algorithms were proposed with this aim, but many of them require information that is unavailable or difficult to obtain in real-life environments, such as how much energy would require a job to be finished. In this context, Simple Energy Aware Scheduler (SEAS) is a scheduling technique for computational intensive Mobile Grids that only require easily accessible information. It was proposed in 2010 and it has been the base for a range of research work. Despite being described as easily implementable in real-life scenarios, SEAS and other SEAS-improvements works have always been evaluated using simulations. In this work, we present a distributed computing platform for mobile devices that support SEAS and empirical evaluation of the SEAS scheduler. The obtained result supports previous simulation results and by extension further validating other SEAS-based results.
Alexander Perez Campos, Juan Manuel Rodriguez, Alejandro Zunino
CLEI2
2017 A performance comparison of data-aware heuristics for scheduling jobs in mobile grids
abstract
Given mobile devices ubiquity and capabilities, some researchers now consider them as resource providers of distributed environments called mobile Grids for running resource intensive software. Therefore, job scheduling has to deal with device singularities, such as energy constraints, mobility and unstable connectivity. Many existing schedulers consider at least one of these aspects, but their applicability strongly depends on information that is unavailable or difficult to estimate accurately, like job execution time. Other efforts do not assume knowing job CPU requirements but ignore energy consumption due to data transfer operations, which is not realistic for data-intensive applications. This work, on the contrary, considers the last as non negligible and known by the scheduler. Under these assumptions, we conduct a performance study of several traditional scheduling heuristics adapted to this environment, which are applied with the known information of jobs but evaluated along with job information unknown to the scheduler. Experiments are performed via a simulation software that employs hardware profiles derived from real mobile devices. Our goal is to contribute to better understand both the capabilities and limitations of this kind of schedulers in the incipient area of mobile Grids.
Matías Hirsch, Cristian Mateos, Juan Manuel Rodriguez, Alejandro Zunino, Yisel Gari, David A. Monge
CLEI3
2017 Word embeddings for improving REST services discoverability
abstract
Web Services have become essential to the software industry as they provide reusable, remotely-accessible functionality and data, thus accelerating client application development and relieving users from maintaining the consumed services. Since Web Services-particularly their descriptions-must be discovered before being consumed, many discovery approaches based on classic Information Retrieval techniques, which store and process textual service descriptions, have arisen. These discovery approaches are affected by natural language ambiguity, such as synonymy and homonymy. Such issues are known as term mismatch, since descriptions relevant to a keyword-based query can be retrieved only if they share many words. Recently, Word Embeddings emerged and tried to cope with this problem by representing words in a language as vectors in a continuous vector space. An interesting property of these vectors is that two different words with similar meaning are represented by vectors that are close in the space. Word Embeddings aim at categorising and quantifying words so that semantic relationships can be established through simple vector distance measures. In this paper, we exploit Word Embeddings to find hidden relationships between service descriptions and queries for the case of REST services, a recent alternative to SOAP-oriented services. The results showed improvements over classical service retrieval techniques such as Vector Space Model or Latent Semantic Analysis of up to 20% in Precision, 39% in Recall, 35% in F-Measure and 10% in NDCG.
Ignacio Lizarralde, Juan Manuel Rodriguez, Cristian Mateos, Alejandro Zunino
CLEI2
2017 A Two-Phase Energy-Aware Scheduling Approach for CPU-Intensive Jobs in Mobile Grids
Matías Hirsch, Juan Manuel Rodriguez, Cristian Mateos, Alejandro Zunino
J. Grid Comput.2
2017 A multi-core computing approach for large-scale multi-label classification
abstract
Large scale multi-label learning, i.e. the problem of determining the associated set of labels for an instance, is gaining relevance in recent years due to the emergence of several real-world applications. Most notably, the exponential growth of the Social Web where a resource can be labeled by mil lions of users using one or more tags, i.e. a resource can be associated to several labels at the same time. A well-known approach for multi-label classification is the Binary Relevance (BR) algorithm which trains a binary classifier for each label independently. However, the serial implementation of BR is not suitable for medium or large datasets due to the time and computational resources required for training. For example, training classifiers for mid-size datasets using MULAN implementation of BR might take several weeks. This paper discusses a parallel implementation of the MULAN BR technique that harnesses the computational power of nowadays multi-core processors. Our implementation presents a speed-up in the training phase of up to 12 times when compared to the original MULAN implementation. In addition, the cross-validation technique of MULAN had huge RAM requirements, making it unusable with large datasets. Therefore, we have overcome this limitation by using compact data structures and taking advantage of disk caching. We have also compared our implementation against scikit-learn, a popular tool for data mining and data analysis, showing significant improvements in speed-up.
Juan Manuel Rodriguez, Daniela Godoy, Cristian Mateos, Alejandro Zunino
Intell. Data Anal.1
2016 Battery-aware centralized schedulers for CPU-bound jobs in mobile Grids
Matías Hirsch, Juan Manuel Rodriguez, Alejandro Zunino, Cristian Mateos
Pervasive Mob. Comput.2
2015 Improving REST Service Discovery with Unsupervised Learning Techniques
abstract
Discovery and replacement are two of the main features of Service Oriented Computing. There has been much research on these topics for traditional SOAP-based Web Services, particularly on discovery. Although the original proposal for REST services lacks this feature, some researchers have studied how to perform discovery for REST services using both IR based techniques and semantic techniques. This work presents a novel IR-based discovery approach for REST services described via WADL files. Our approach takes advantage of unsupervised machine learning techniques for improving discovering results. In particular, the approach relies on clustering algorithms, such as K-means or X-means, to reduce the search space for a given query. The experimental results show that using an appropriate clustering technique, our approach achieves nearly 4 times higher F-measure than a traditional IR-based search engine, namely Apache Lucene. Additionally, the paper reports other metrics, such as Recall, Precision, Precision at-10 and Recall at-10, that also point out that the proposed approach outperforms Lucene. Finally, another important contribution is a set of queries and WADL files gathered from the Internet that can be used for evaluating future discovery proposals.
Juan Manuel Rodriguez, Alejandro Zunino, Cristian Mateos, Felix Oscar Segura, Emmanuel Rodriguez
CISIS1
2015 Assisting developers to build high-quality code-first Web Service APIs
Juan Manuel Rodriguez, Cristian Mateos, Alejandro Zunino
J. Web Eng.1
2015 A tool to improve code-first Web services discoverability through text mining techniques
abstract
Summary Service‐oriented development is challenging mainly because Web service developers tend to disregard the importance of the exposed service APIs, which are specified using Web Service Description Language (WSDL) documents. Methodologically, WSDL documents can be either manually generated or inferred from service implementations using WSDL generation tools. The latter option, called code first, is the most used approach in the industry. However, it is known that there are some bad practices in service implementations or defects in WSDL generation tools that may cause WSDL documents to present WSDL anti‐patterns, which in turn compromise the chances of documents of being discovered and understood. In this paper, we present a software tool that assists developers in obtaining WSDL documents with as few WSDL anti‐patterns as possible. The tool combines text mining and meta‐programming techniques to process service implementations and is developed as an Eclipse plug‐in. An evaluation of the tool by using a data‐set of real service implementations in terms of anti‐pattern avoidance accuracy and discovery performance by using classical Information Retrieval metrics—Precision‐at‐n, Recall and Normalized Discounted Cumulative Gain—is also reported.Copyright © 2014 John Wiley & Sons, Ltd.
Cristian Mateos, Juan Manuel Rodriguez, Alejandro Zunino
Softw. Pract. Exp.2
2013 An Approach for Web Service Discoverability Anti-Patterns Detection
Juan Manuel Rodriguez, Marco Crasso, Alejandro Zunino
J. Web Eng.1
2013 Best practices for describing, consuming, and discovering web services: a comprehensive toolset
abstract
SUMMARY The service‐oriented computing (SOC) paradigm has recently gained a lot of attention in the software industry because SOC represents a novel and a fresh way of architecting distributed applications. SOC is usually materialized via web services, which allows developers to structure applications exposing a clear, public interface to their capabilities. Although conceptually and technologically mature, SOC still lacks adequate development support from a methodological point of view. In this paper, we present the EasySOC project, a set of guidelines to simplify the development of service‐oriented applications and services. EasySOC is a synthesized catalog of best SOC development practices that arise as a result of several years of research in fundamental Services Computing topics, that is, Web Service Description Language‐based technical specification, Web Service discovery, and Web Service outsourcing. In addition, we describe a materialization of the guidelines for the Java language, which has been implemented as a plug‐in for the Eclipse IDE. We believe that both the practical nature of the guidelines and the availability of this software that enforces them may help software practitioners to rapidly exploit our ideas for building real SOC applications. Copyright © 2012 John Wiley & Sons, Ltd.
Juan Manuel Rodriguez, Marco Crasso, Cristian Mateos, Alejandro Zunino
Softw. Pract. Exp.1
2010 Improving Web Service descriptions for effective service discovery
Juan Manuel Rodriguez, Marco Crasso, Alejandro Zunino, Marcelo R. Campo
Sci. Comput. Program.1