VLDB 2026 Research / reviewers in the wild / expert
Carlos Rodríguez 0001
dblp:89/2560-1
· DBLP profile ↗
16ranked-venue papers
5as first author
3since 2021 · last 2025
0000-0002-7263-821XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 12 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 10 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 5 · 2 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | VestiME: A RAG-based Multimodal Fashion Outfit Recommender SystemabstractFashion plays a significant role in cultures around the world and in the global economy. Similarly to what has happened with other consumer products, the Internet has trans-formed the way people shop for fashion, transitioning from an in-store experience to an increasingly online one thanks to ever-evolving e-commerce. While much of today’s e-commerce experience for fashion still relies on a keyword-search-based approach for web and mobile catalogs, the advent of Generative Artificial Intelligence (GenAI) and multimodal Large Language Models (LLMs) has brought unprecedented opportunities for searching, exploring, and recommending fashion products online. This paper aims to leverage state-of-the-art GenAI and multimodal LLM technologies to provide a different customer experience based on conversational, intent-driven exploration of fashion catalogs and outfit recommendations through our proposed solution, VestiME. The results of our user study involving fashion experts show the potential of these highly interactive, conversational, and multimodal e-commerce experiences for fashion. Amanda Maduro, Carlos Rodríguez 0001, Sebastian Ortiz-Chamorro |
CLEI | 2 |
| 2021 | An Early Alert System for Software Vulnerabilities based on Vulnerability Repositories and Social NetworksabstractThe huge amount of information regarding software vulnerabilities, the multiple and heterogeneous information sources, and the lack of awareness about the dangers of software vulnerabilities, exacerbates the risks of security threats being materialized. In this complex context, this paper approaches the problem of managing early alerts for software vulnerablities by leveraging existing vulnerability information found in vulnerability repositories and social networks. To this end, we propose a solution based on techniques that stem from automated retrieval of information about vulneratilities from the above sources, userdefined preferences regarding their technological environment and intelligent vulnerability tagging. Our user studies reveal the feasibility of our approach as a tool for managing early alerts regarding software vulnerabilities and keeping security professionals aware of them. Néstor Fabián Riveros, Carlos Rodríguez 0001 |
CLEI | 2 |
| 2021 | Why Some Bug-bounty Vulnerability Reports are Invalid?: Study of bug-bounty reports and developing an out-of-scope taxonomy modelabstractBackground: Despite the increasing popularity of bug-bounty platforms in industry, little empirical evidence exists to identify the nature of invalid vulnerability reports. Mitigation of invalid reports is a serious concern of organisations running or using bug-bounty platforms as well as security researchers. Aims: In this work we aim to identify: (i) why some reports are considered as invalid? (ii) what are the characteristics of reports considered as invalid due to being out-of-scope? Method: We conducted an empirical study on disclosed invalid reports in HackerOne to examine the reasons these reports are marked as invalid and we found that out-of-scope is the leading reason. Since all out-of-scope reports were rejected according to the programs policy page we studied all programs policy pages in two major bug-bounty platforms to understand the characteristics of an out-of-scope report. We developed a generalised out-of-scope taxonomy model and we used our model to further analyse HackerOne out-of-scope reports to find the leading attributes of this model that contributes to the fate of these reports. Results: We identified out-of-scope followed by false-positive as two main reasons for a report to be deemed invalid. We found that the attribute of vulnerability type in our taxonomy model is the leading characteristic of out-of-scope reports. We also identified the top 9 out-of-interest vulnerability types according to policy pages. Conclusions: Our study can help bug-bounty platforms and researchers to better understand the nature of invalid reports. Our finding about the importance of vulnerability type in validating reports can be used to justify future works to develop automated classification techniques based on vulnerability types to better triage invalid reports. Our top 9 out-of-interest vulnerability types can be used as a blacklist to automatically classify possibly an out-of-scope report. Finally our generalised out-of-scope taxonomy model can guide organisations as a base model to create their policy page and tailor it as they need. Saman Shafigh, Boualem Benatallah, Carlos Rodríguez 0001, Mortada Al-Banna |
ESEM | 3 |
| 2020 | State Machine Based Human-Bot Conversation Model and Services
Shayan Zamanirad, Boualem Benatallah, Carlos Rodríguez 0001, Mohammad-Ali Yaghoub-Zadeh-Fard, Sara Bouguelia, Hayet Brabra |
CAiSE | 3 |
| 2020 | API Topics Issues in Stack Overflow Q&As Posts: An Empirical StudyabstractApplication Programming Interfaces (APIs) have become one of the key assets within modern businesses, facilitating the linking and integration of intra- and inter-organizational data and systems in the context of complex and heterogeneous technology ecosystems. APIs allow organizations to monetize data, build profitable partnerships and foster innovation and growth. Understanding APIs and their usage are therefore key to building solutions for enabling successful business operations. This paper aims at understanding API topic issues posted on Stack Overflow (SO), a Community Question Answering (CQA) site for programmers. We conduct an empirical analysis on a sample of 400 randomly-selected Q&As threads to help identify API-related issues and their main topics. A thematic analysis performed on this sample reveals eight main topics related to APIs, among which API usage, debugging, API constraints and API security emerged as the major ones. We also exemplify the types of support provided by SO community in addressing each of the identified topics and discuss possible venues on how to further leverage this knowledge. George Ajam, Carlos Rodríguez 0001, Boualem Benatallah |
CLEI | 2 |
| 2020 | API Topic Issues Indexing, Exploration and Discovery for API Community KnowledgeabstractApplication Programming Interface (API) is a core technology that facilitates developers' productivity by enabling the reuse of software components. Understanding APIs and gaining knowledge about their usage are therefore fundamental needs for developers that impact a wide range of software development activities. This paper presents an approach to enable API users to explore, discover and learn about APIs through API topic issues discussed in Stack Overflow (SO), a widely used programming, community question-answering (CQA) site. Our work proposes an integrated API Knowledge Base (KB) and indexing technique that combines both SO API-related posts as well as other API learning resources collected from the Web (e.g., API video-tutorials from Youtube). The resulting indexed and enriched API community knowledge can be queried in a API-topic-issue driven manner using a simple yet powerful domain-specific language (DSL). We demonstrate the feasibility of our approach through Scout-bot, our tool for exploration and discovery of API topic issues. George Ajam, Carlos Rodríguez 0001, Boualem Benatallah |
CLEI | 2 |
| 2020 | Learning Word Representation for the Cyber Security Vulnerability DomainabstractThere have been ever-increasing amounts of security vulnerabilities discovered and reported in recent years. Much of the information related to these vulnerabilities is currently available to the public, in the form of rich, textual data (e.g. vulnerability reports). Many of the state-of-the-art techniques used today to process such textual data rely on so-called word embeddings. As of today, several pre-trained embeddings have been created, many of which rely on general-purpose training datasets such as Google News and Wikipedia. More recently, other domain-specific word embeddings have been created (e.g. in the context of software development) to cope with terminology and ambiguity limitations of existing general-purpose embeddings. The availability of word embeddings for specialised domains is critical for the effectiveness of domain-specific tasks that rely on this technique. In this paper, we propose a word embedding for the cyber security vulnerability domain. We train our embedding model on multiple, rich and heterogeneous security vulnerability information sources publicly available on the web. The benefits of such specialised word embedding are demonstrated through a qualitative comparison of word similarity and the exemplary task of matching security professionals to vulnerability discovery tasks posted to bug bounty programs. We also introduce a new dataset of words pairs similarity with a human judgement that can be used as a benchmark. Our experimental results show that, in the context of cyber security, our domain-specific word embedding outperforms existing pre-trained embeddings built on general-purpose and software engineering datasets. Sara Mumtaz, Carlos Rodríguez 0001, Boualem Benatallah, Mortada Al-Banna, Shayan Zamanirad |
IJCNN | 2 |
| 2020 | Automatic Action Extraction for Short Text Conversation Using Unsupervised Learning
Senthil Ganesan Yuvaraj, Shayan Zamanirad, Boualem Benatallah, Carlos Rodríguez 0001 |
WISE (2) | 4 |
| 2019 | Security Vulnerability Information Service with Natural Language Query Support
Carlos Rodríguez 0001, Shayan Zamanirad, Reza Nouri, Kirtana Darabal, Boualem Benatallah, Mortada Al-Banna |
CAiSE | 1 |
| 2019 | Expert2Vec: Experts Representation in Community Question Answering for Question Routing
Sara Mumtaz, Carlos Rodríguez 0001, Boualem Benatallah |
CAiSE | 2 |
| 2017 | Programming bots by synthesizing natural language expressions into API invocationsabstractAt present, bots are still in their preliminary stages of development. Many are relatively simple, or developed ad-hoc for a very specific use-case. For this reason, they are typically programmed manually, or utilize machine-learning classifiers to interpret a fixed set of user utterances. In reality, real world conversations with humans require support for dynamically capturing users expressions. Moreover, bots will derive immeasurable value by programming them to invoke APIs for their results. Today, within the Web and Mobile development community, complex applications are being stringed together with a few lines of code - all made possible by APIs. Yet, developers today are not as empowered to program bots in much the same way. To overcome this, we introduce BotBase, a bot programming platform that dynamically synthesizes natural language user expressions into API invocations. Our solution is two faceted: Firstly, we construct an API knowledge graph to encode and evolve APIs; secondly, leveraging the above we apply techniques in NLP, ML and Entity Recognition to perform the required synthesis from natural language user expressions into API calls. Shayan Zamanirad, Boualem Benatallah, Moshe Chai Barukh, Fabio Casati, Carlos Rodríguez 0001 |
ASE | 5 |
| 2016 | REST APIs: A Large-Scale Analysis of Compliance with Principles and Best Practices
Carlos Rodríguez 0001, Marcos Báez, Florian Daniel, Fabio Casati, Juan Carlos Trabucco, Luigi Canali, Gianraffaele Percannella |
ICWE | 1 |
| 2016 | Analysis and improvement of business process models using spreadsheets
Jorge Saldivar, Carla Vairetti, Carlos Rodríguez 0001, Florian Daniel, Fabio Casati, Rosa Alarcón |
Inf. Syst. | 3 |
| 2016 | Mining and Quality Assessment of Mashup Model Patterns with the Crowd: A Feasibility StudyabstractPattern mining, that is, the automated discovery of patterns from data, is a mathematically complex and computationally demanding problem that is generally not manageable by humans. In this article, we focus on small datasets and study whether it is possible to mine patterns with the help of the crowd by means of a set of controlled experiments on a common crowdsourcing platform. We specifically concentrate on mining model patterns from a dataset of real mashup models taken from Yahoo! Pipes and cover the entire pattern mining process, including pattern identification and quality assessment. The results of our experiments show that a sensible design of crowdsourcing tasks indeed may enable the crowd to identify patterns from small datasets (40 models). The results, however, also show that the design of tasks for the assessment of the quality of patterns to decide which patterns to retain for further processing and use is much harder (our experiments fail to elicit assessments from the crowd that are similar to those by an expert). The problem is relevant in general to model-driven development (e.g., UML, business processes, scientific workflows), in that reusable model patterns encode valuable modeling and domain knowledge, such as best practices, organizational conventions, or technical choices, that modelers can benefit from when designing their own models. Carlos Rodríguez 0001, Florian Daniel, Fabio Casati |
ACM Trans. Internet Techn. | 1 |
| 2014 | Crowd-Based Mining of Reusable Process Model Patterns
Carlos Rodríguez 0001, Florian Daniel, Fabio Casati |
BPM | 1 |
| 2013 | SOA-enabled compliance management: instrumenting, assessing, and analyzing service-based business processes
Carlos Rodríguez 0001, Daniel Schleicher, Florian Daniel, Fabio Casati, Frank Leymann, Sebastian Wagner 0001 |
Serv. Oriented Comput. Appl. | 1 |