EDBT 2026 Demo / reviewers in the wild / expert
Marcos Báez
dblp:45/112
· DBLP profile ↗
14ranked-venue papers in the field
3as first author
8since 2021 · last 2023
0000-0003-1666-2474ORCID · verified
Domains — venue-derived; a paper can count in several
Business Process & Enterprise Data · 5Database Systems & Data Management · 3 (1 first)Information Retrieval & Web Search · 3 (2 first)Data Mining & Knowledge Discovery · 2Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Adaptive search query generation and refinement in systematic literature review
Maisie Badami, Boualem Benatallah, Marcos Báez |
Inf. Syst. | 3 |
| 2022 | Systematic Literature Review Search Query Refinement Pipeline: Incremental Enrichment and Adaptation
Maisie Badami, Boualem Benatallah, Marcos Báez |
CAiSE | 3 |
| 2022 | Context Knowledge-Aware Recognition of Composite Intents in Task-Oriented Human-Bot Conversations
Sara Bouguelia, Hayet Brabra, Boualem Benatallah, Marcos Báez, Shayan Zamanirad, Hamamache Kheddouci |
CAiSE | 4 |
| 2022 | Crowdsourcing Syntactically Diverse Paraphrases with Diversity-Aware Prompts and Workflows
Jorge Ramírez, Marcos Báez, Auday Berro, Boualem Benatallah, Fabio Casati |
CAiSE | 2 |
| 2022 | Supporting Natural Language Interaction with the Web
Marcos Báez, Cinzia Cappiello, Claudia Maria Cutrupi, Maristella Matera, Isabella Possaghi, Emanuele Pucci, Gianluca Spadone, Antonella Pasquale |
ICWE | 1 |
| 2021 | Reusable Abstractions and Patterns for Recognising Compositional Conversational Flows
Sara Bouguelia, Hayet Brabra, Shayan Zamanirad, Boualem Benatallah, Marcos Báez, Hamamache Kheddouci |
CAiSE | 5 |
| 2021 | On the Impact of Predicate Complexity in Crowdsourced Classification TasksabstractThis paper explores and offers guidance on a specific and relevant problem in task design for crowdsourcing: how to formulate a complex question used to classify a set of items. In micro-task markets, classification is still among the most popular tasks. We situate our work in the context of information retrieval and multi-predicate classification, i.e., classifying a set of items based on a set of conditions. Our experiments cover a wide range of tasks and domains, and also consider crowd workers alone and in tandem with machine learning classifiers. We provide empirical evidence into how the resulting classification performance is affected by different predicate formulation strategies, emphasizing the importance of predicate formulation as a task design dimension in crowdsourcing. Jorge Ramírez, Marcos Báez, Fabio Casati, Luca Cernuzzi, Boualem Benatallah, Ekaterina A. Taran, Veronika A. Malanina |
WSDM | 2 |
| 2021 | An Extensible and Reusable Pipeline for Automated Utterance ParaphrasesabstractIn this demonstration paper we showcase an extensible and reusable pipeline for automatic paraphrase generation , i.e., reformulating sentences using different words. Capturing the nuances of human language is fundamental to the effectiveness of Conversational AI systems, as it allows them to deal with the different ways users can utter their requests in natural language. Traditional approaches to utterance paraphrasing acquisition, such as hiring experts or crowd-sourcing, involve processes that are often costly or time consuming, and with their own trade-offs in terms of quality. Automatic paraphrasing is emerging as an attractive alternative that promises a fast, scalable and cost-effective process. In this paper we showcase how our extensible and reusable pipeline for automated utterance paraphrasing can support the development of Conversational AI systems by integrating and extending existing techniques under an unified and configurable framework. Auday Berro, Mohammad-ali Yaghub Zade Fard, Marcos Báez, Boualem Benatallah, Khalid Benabdeslem |
Proc. VLDB Endow. | 3 |
| 2020 | Challenges and strategies for running controlled crowdsourcing experimentsabstractThis paper reports on the challenges and lessons we learned while running controlled experiments in crowdsourcing platforms. Crowdsourcing is becoming an attractive technique to engage a diverse and large pool of subjects in experimental research, allowing researchers to achieve levels of scale and completion times that would otherwise not be feasible in lab settings. However, the scale and flexibility comes at the cost of multiple and sometimes unknown sources of bias and confounding factors that arise from technical limitations of crowdsourcing platforms and from the challenges of running controlled experiments in the “wild”. In this paper, we take our experience in running systematic evaluations of task design as a motivating example to explore, describe, and quantify the potential impact of running uncontrolled crowdsourcing experiments and derive possible coping strategies. Among the challenges identified, we can mention sampling bias, controlling the assignment of subjects to experimental conditions, learning effects, and reliability of crowdsourcing results. According to our empirical studies, the impact of potential biases and confounding factors can amount to a 38% loss in the utility of the data collected in uncontrolled settings; and it can significantly change the outcome of experiments. These issues ultimately inspired us to implement CrowdHub, a system that sits on top of major crowdsourcing platforms and allows researchers and practitioners to run controlled crowdsourcing projects. Jorge Ramírez, Marcos Báez, Fabio Casati, Luca Cernuzzi, Boualem Benatallah |
CLEI | 2 |
| 2020 | Automatic Generation of Chatbots for Conversational Web Browsing
Pietro Chittò, Marcos Báez, Florian Daniel, Boualem Benatallah |
ER | 2 |
| 2019 | Understanding the Impact of Text Highlighting in Crowdsourcing TasksabstractText classification is one of the most common goals of machine learning (ML) projects, and also one of the most frequent human intelligence tasks in crowdsourcing platforms. ML has mixed success in such tasks depending on the nature of the problem, while crowd-based classification has proven to be surprisingly effective, but can be expensive. Recently, hybrid text classification algorithms, combining human computation and machine learning, have been proposed to improve accuracy and reduce costs. One way to do so is to have ML highlight or emphasize portions of text that it believes to be more relevant to the decision. Humans can then rely only on this text or read the entire text if the highlighted information is insufficient. In this paper, we investigate if and under what conditions highlighting selected parts of the text can (or cannot) improve classification cost and/or accuracy, and in general how it affects the process and outcome of the human intelligence tasks. We study this through a series of crowdsourcing experiments running over different datasets and with task designs imposing different cognitive demands. Our findings suggest that highlighting is effective in reducing classification effort but does not improve accuracy - and in fact, low-quality highlighting can decrease it. Jorge Ramírez, Marcos Báez, Fabio Casati, Boualem Benatallah |
HCOMP | 2 |
| 2016 | REST APIs: A Large-Scale Analysis of Compliance with Principles and Best Practices
Carlos Rodríguez 0001, Marcos Báez, Florian Daniel, Fabio Casati, Juan Carlos Trabucco, Luigi Canali, Gianraffaele Percannella |
ICWE | 2 |
| 2011 | Knowledge Spaces
Marcos Báez, Fabio Casati, Maurizio Marchese |
ICWE | 1 |
| 2009 | Universal Resource Lifecycle ManagementabstractThis paper presents a model and a tool that allows Web users to define, execute, and manage lifecycles for any artifact available on the Web. In the paper we show the need for lifecycle management of Web artifacts, and we show in particular why it is important that non-programmers are also able to do this. We then discuss why current models do not allow this, and we present a model and a system implementation that achieves lifecycle management for any URI-identifiable and accessible object. The most challenging parts of the work lie in the definition of a simple but universal model and system (and in particular in allowing universality and simplicity to coexist) and in the ability to hide from the lifecycle modeler the complexity intrinsic in having to access and manage a variety of resources, which differ in nature, in the operations that are allowed on them, and in the protocols and data formats required to access them. Marcos Báez, Fabio Casati, Maurizio Marchese |
ICDE | 1 |