VLDB 2026 Research / reviewers in the wild / expert
Marco Túlio Ribeiro
dblp:21/10105
· DBLP profile ↗
24ranked-venue papers
9as first author
14since 2021 · last 2023
0000-0002-3301-1297ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 first-authorHuman-computer interaction and ubiquitous computing · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Targeted Data Generation: Finding and Fixing Model WeaknessesabstractEven when aggregate accuracy is high, stateof-the-art NLP models often fail systematically on specific subgroups of data, resulting in unfair outcomes and eroding user trust.Additional data collection may not help in addressing these weaknesses, as such challenging subgroups may be unknown to users, and underrepresented in the existing and new data.We propose Targeted Data Generation (TDG), a framework that automatically identifies challenging subgroups, and generates new data for those subgroups using large language models (LLMs) with a human in the loop.TDG estimates the expected benefit and potential harm of data augmentation for each subgroup, and selects the ones most likely to improve withingroup performance without hurting overall performance.In our experiments, TDG 1 significantly improves the accuracy on challenging subgroups for state-of-the-art sentiment analysis and natural language inference models, while also improving overall test accuracy. Zexue He, Marco Túlio Ribeiro, Fereshte Khani |
ACL (1) | 2 |
| 2023 | Supporting Human-AI Collaboration in Auditing LLMs with LLMsabstractLarge language models (LLMs) are increasingly becoming all-powerful and pervasive via deployment in sociotechnical systems. Yet these language models, be it for classification or generation, have been shown to be biased, behave irresponsibly, causing harm to people at scale. It is crucial to audit these language models rigorously before deployment. Existing auditing tools use either or both humans and AI to find failures. In this work, we draw upon literature in human-AI collaboration and sensemaking, and interview research experts in safe and fair AI, to build upon the auditing tool: AdaTest [36], which is powered by a generative LLM. Through the design process we highlight the importance of sensemaking and human-AI communication to leverage complementary strengths of humans and generative models in collaborative auditing. To evaluate the effectiveness of AdaTest++, the augmented tool, we conduct user studies with participants auditing two commercial language models: OpenAI’s GPT-3 and Azure’s sentiment analysis model. Qualitative analysis shows that AdaTest++ effectively leverages human strengths such as schematization, hypothesis testing. Further, with our tool, users identified a variety of failures modes, covering 26 different topics over 2 tasks, that have been shown in formal audits and also those previously under-reported. Charvi Rastogi, Marco Túlio Ribeiro, Nicholas King, Harsha Nori, Saleema Amershi |
AIES | 2 |
| 2023 | Adaptive Testing of Computer Vision ModelsabstractVision models often fail systematically on groups of data that share common semantic characteristics (e.g., rare objects or unusual scenes), but identifying these failure modes is a challenge. We introduce AdaVision, an interactive process for testing vision models which helps users identify and fix coherent failure modes. Given a natural language description of a coherent group, AdaVision retrieves relevant images from LAION-5B with CLIP. The user then labels a small amount of data for model correctness, which is used in successive retrieval rounds to hill-climb towards high-error regions, refining the group definition. Once a group is saturated, AdaVision uses GPT-3 to suggest new group descriptions for the user to explore. We demonstrate the usefulness and generality of AdaVision in user studies, where users find major bugs in state-of-the-art classification, object detection, and image captioning models. These user-discovered groups have failure rates 2-3x higher than those surfaced by automatic error clustering methods. Finally, finetuning on examples found with AdaVision fixes the discovered bugs when evaluated on unseen examples, without degrading in-distribution accuracy, and while also improving performance on out-of-distribution datasets. Irena Gao, Gabriel Ilharco, Scott M. Lundberg, Marco Túlio Ribeiro |
ICCV | 4 |
| 2023 | Editing models with task arithmetic
Gabriel Ilharco, Marco Túlio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, Ali Farhadi |
ICLR | 2 |
| 2023 | ScatterShot: Interactive In-context Example Curation for Text TransformationabstractThe in-context learning capabilities of LLMs like GPT-3 allow annotators to customize an LLM to their specific tasks with a small number of examples. However, users tend to include only the most obvious patterns when crafting examples, resulting in underspecified in-context functions that fall short on unseen cases. Further, it is hard to know when “enough” examples have been included even for known patterns. In this work, we present ScatterShot, an interactive system for building high-quality demonstration sets for in-context learning. ScatterShot iteratively slices unlabeled data into task-specific patterns, samples informative inputs from underexplored or not-yet-saturated slices in an active learning manner, and helps users label more efficiently with the help of an LLM and the current example set. In simulation studies on two text perturbation scenarios, ScatterShot sampling improves the resulting few-shot functions by 4-5 percentage points over random sampling, with less variance as more examples are added. In a user study, ScatterShot greatly helps users in covering different patterns in the input space and labeling in-context examples more efficiently, resulting in better in-context learning and less user effort. Sherry Tongshuang Wu, Hua Shen 0005, Daniel S. Weld, Jeffrey Heer, Marco Túlio Ribeiro |
IUI | 5 |
| 2023 | Collaborative Alignment of NLP ModelsabstractDespite substantial advancements, Natural Language Processing (NLP) models often require post-training adjustments to enforce business rules, rectify undesired behavior, and align with user values.
These adjustments involve operationalizing "concepts"—dictating desired model responses to certain inputs.
However, it's difficult for a single entity to enumerate and define all possible concepts, indicating a need for a multi-user, collaborative model alignment framework.
Moreover, the exhaustive delineation of a concept is challenging, and an improper approach can create shortcuts or interfere with original data or other concepts.
To address these challenges, we introduce CoAlign, a framework that enables multi-user interaction with the model, thereby mitigating individual limitations.
CoAlign aids users in operationalizing their concepts using Large Language Models, and relying on the principle that NLP models exhibit simpler behaviors in local regions.
Our main insight is learning a \emph{local} model for each concept, and a \emph{global} model to integrate the original data with all concepts.
We then steer a large language model to generate instances within concept boundaries where local and global disagree.
Our experiments show CoAlign is effective at helping multiple users operationalize concepts and avoid interference for a variety of scenarios, tasks, and models. Fereshte Khani, Marco Túlio Ribeiro |
NeurIPS | 2 |
| 2023 | What Did My AI Learn? How Data Scientists Make Sense of Model BehaviorabstractData scientists require rich mental models of how AI systems behave to effectively train, debug, and work with them. Despite the prevalence of AI analysis tools, there is no general theory describing how people make sense of what their models have learned. We frame this process as a form of sensemaking and derive a framework describing how data scientists develop mental models of AI behavior. To evaluate the framework, we show how existing AI analysis tools fit into this sensemaking process and use it to design AIFinnity , a system for analyzing image-and-text models. Lastly, we explored how data scientists use a tool developed with the framework through a think-aloud study with 10 data scientists tasked with using AIFinnity to pick an image captioning model. We found that AIFinnity ’s sensemaking workflow reflected participants’ mental processes and enabled them to discover and validate diverse AI behaviors. Ángel Alexander Cabrera, Marco Túlio Ribeiro, Bongshin Lee, Robert DeLine, Adam Perer, Steven Mark Drucker |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2022 | Do Feature Attribution Methods Correctly Attribute Features?abstractFeature attribution methods are popular in interpretable machine learning. These methods compute the attribution of each input feature to represent its importance, but there is no consensus on the definition of "attribution", leading to many competing methods with little systematic evaluation, complicated in particular by the lack of ground truth attribution. To address this, we propose a dataset modification procedure to induce such ground truth. Using this procedure, we evaluate three common methods: saliency maps, rationales, and attentions. We identify several deficiencies and add new perspectives to the growing body of evidence questioning the correctness and reliability of these methods applied on datasets in the wild. We further discuss possible avenues for remedy and recommend new attribution methods to be tested against ground truth before deployment. The code and appendix are available at https://yilunzhou.github.io/feature-attribution-evaluation/. Yilun Zhou, Serena Booth, Marco Túlio Ribeiro, Julie A. Shah |
AAAI | 3 |
| 2022 | Adaptive Testing and Debugging of NLP ModelsabstractCurrent approaches to testing and debugging NLP models rely on highly variable human creativity and extensive labor, or only work for a very restrictive class of bugs.We present AdaTest, a process which uses large scale language models (LMs) in partnership with human feedback to automatically write unit tests highlighting bugs in a target model.Such bugs are then addressed through an iterative text-fixretest loop, inspired by traditional software development.In experiments with expert and non-expert users and commercial / research models for 8 different tasks, AdaTest makes users 5-10x more effective at finding bugs than current approaches, and helps users effectively fix bugs without adding new bugs. Marco Túlio Ribeiro, Scott M. Lundberg |
ACL (1) | 1 |
| 2022 | Fixing Model Bugs with Natural Language PatchesabstractCurrent approaches for fixing systematic problems in NLP models (e.g., regex patches, finetuning on more data) are either brittle, or labor-intensive and liable to shortcuts.In contrast, humans often provide corrections to each other through natural language.Taking inspiration from this, we explore natural language patches-declarative statements that allow developers to provide corrective feedback at the right level of abstraction, either overriding the model ("if a review gives 2 stars, the sentiment is negative") or providing additional information the model may lack ("if something is described as the bomb, then it is good").We model the task of determining if a patch applies separately from the task of integrating patch information, and show that with a small amount of synthetic data, we can teach models to effectively use real patches on real data-1 to 7 patches improve accuracy by ~1-4 accuracy points on different slices of a sentiment analysis dataset, and F1 by 7 points on a relation extraction dataset.Finally, we show that finetuning on as many as 100 labeled examples may be needed to match the performance of a small set of language patches. Shikhar Murty, Christopher D. Manning, Scott M. Lundberg, Marco Túlio Ribeiro |
EMNLP | 4 |
| 2022 | ExSum: From Local Explanations to Model UnderstandingabstractInterpretability methods are developed to understand the working mechanisms of blackbox models, which is crucial to their responsible deployment.Fulfilling this goal requires both that the explanations generated by these methods are correct and that people can easily and reliably understand them.While the former has been addressed in prior work, the latter is often overlooked, resulting in informal model understanding derived from a handful of local explanations.In this paper, we introduce explanation summary (EXSUM), a mathematical framework for quantifying model understanding, and propose metrics for its quality assessment.On two domains, EXSUM highlights various limitations in the current practice, helps develop accurate model understanding, and reveals easily overlooked properties of the model.We also connect understandability to other properties of explanations such as human alignment, robustness, and counterfactual similarity and plausibility. Yilun Zhou, Marco Túlio Ribeiro, Julie A. Shah |
NAACL-HLT | 2 |
| 2021 | Polyjuice: Generating Counterfactuals for Explaining, Evaluating, and Improving ModelsabstractTongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, Daniel Weld. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Sherry Tongshuang Wu, Marco Túlio Ribeiro, Jeffrey Heer, Daniel S. Weld |
ACL/IJCNLP (1) | 2 |
| 2021 | Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team PerformanceabstractMany researchers motivate explainable AI with studies showing that human-AI team performance on decision-making tasks improves when the AI explains its recommendations. However, prior studies observed improvements from explanations only when the AI, alone, outperformed both the human and the best team. Can explanations help lead to complementary performance, where team accuracy is higher than either the human or the AI working solo? We conduct mixed-method user studies on three datasets, where an AI with accuracy comparable to humans helps participants solve a task (explaining itself in some conditions). While we observed complementary improvements from AI augmentation, they were not increased by explanations. Rather, explanations increased the chance that humans will accept the AI’s recommendation, regardless of its correctness. Our result poses new challenges for human-centered AI: Can we develop explanatory approaches that encourage appropriate trust in AI, and therefore help generate (or improve) complementary performance? Gagan Bansal, Sherry Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Túlio Ribeiro, Daniel S. Weld |
CHI | 7 |
| 2021 | Beyond Accuracy: Behavioral Testing of NLP Models with Checklist (Extended Abstract)abstractAlthough measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches for evaluating models either focus on individual tasks or on specific behaviors. Inspired by principles of behavioral testing in software engineering, we introduce CheckList, a task-agnostic methodology for testing NLP models. CheckList includes a matrix of general linguistic capabilities and test types that facilitate comprehensive test ideation, as well as a software tool to generate a large and diverse number of test cases quickly. We illustrate the utility of CheckList with tests for three tasks, identifying critical failures in both commercial and state-of-art models. In a user study, a team responsible for a commercial sentiment analysis model found new and actionable bugs in an extensively tested model. In another user study, NLP practitioners with CheckList created twice as many tests, and found almost three times as many bugs as users without it. Marco Túlio Ribeiro, Sherry Tongshuang Wu, Carlos Guestrin, Sameer Singh 0001 |
IJCAI | 1 |
| 2020 | Beyond Accuracy: Behavioral Testing of NLP Models with CheckListabstractAlthough measuring held-out accuracy has been the primary approach to evaluate generalization, it often overestimates the performance of NLP models, while alternative approaches for evaluating models either focus on individual tasks or on specific behaviors.Inspired by principles of behavioral testing in software engineering, we introduce CheckList, a taskagnostic methodology for testing NLP models.CheckList includes a matrix of general linguistic capabilities and test types that facilitate comprehensive test ideation, as well as a software tool to generate a large and diverse number of test cases quickly.We illustrate the utility of CheckList with tests for three tasks, identifying critical failures in both commercial and state-of-art models.In a user study, a team responsible for a commercial sentiment analysis model found new and actionable bugs in an extensively tested model.In another user study, NLP practitioners with CheckList created twice as many tests, and found almost three times as many bugs as users without it. Marco Túlio Ribeiro, Sherry Tongshuang Wu, Carlos Guestrin, Sameer Singh 0001 |
ACL | 1 |
| 2020 | SQuINTing at VQA Models: Introspecting VQA Models With Sub-QuestionsabstractExisting VQA datasets contain questions with varying levels of complexity. While the majority of questions in these datasets require perception for recognizing existence, properties, and spatial relationships of entities, a significant portion of questions pose challenges that correspond to reasoning tasks - tasks that can only be answered through a synthesis of perception and knowledge about the world, logic and / or reasoning. Analyzing performance across this distinction allows us to notice when existing VQA models have consistency issues - they answer the reasoning questions correctly but fail on associated low-level perception questions. For example, in Figure 1, models answer the complex reasoning question “Is the banana ripe enough to eat?” correctly, but fail on the associated perception question “Are the bananas mostly green or yellow?” indicating that the model likely answered the reasoning question correctly but for the wrong reason. We quantify the extent to which this phenomenon occurs by creating a new Reasoning split of the VQA dataset and collecting VQAintrospect, a new dataset1 which currently consists of 200K new perception questions which serve as sub questions corresponding to the set of perceptual tasks needed to effectively answer the complex reasoning questions in the Reasoning split. Our evaluation shows that state-of-the-art VQA models have comparable performance in answering perception and reasoning questions, but suffer from consistency problems. To address this shortcoming, we propose an approach called Sub-Question Importance-aware Network Tuning (SQuINT), which encourages the model to attend to the same parts of the image when answering the reasoning question and the perception sub question. We show that SQuINT improves model consistency by ~7%, also marginally improving performance on the Reasoning questions in VQA, while also displaying better attention maps. Ramprasaath R. Selvaraju, Purva Tendulkar, Devi Parikh, Eric Horvitz, Marco Túlio Ribeiro, Besmira Nushi, Ece Kamar |
CVPR | 5 |
| 2020 | Intelligible and Explainable Machine Learning: Best Practices and Practical ChallengesabstractLearning methods such as boosting and deep learning have made ML models harder to understand and interpret. This puts data scientists and ML developers in the position of often having to make a tradeoff between accuracy and intelligibility. Research in IML (Interpretable Machine Learning) and XAI (Explainable AI) focus on minimizing this trade-off by developing more accurate interpretable models and by developing new techniques to explain black-box models. Such models and techniques make it easier for data scientists, engineers and model users to debug models and achieve important objectives such as ensuring the fairness of ML decisions and the reliability and safety of AI systems. In this tutorial, we present an overview of various interpretability methods and provide a framework for thinking about how to choose the right explanation method for different real-world scenarios. We will focus on the application of XAI in practice through a variety of case studies from domains such as healthcare, finance, and bias and fairness. Finally, we will present open problems and research directions for the data mining and machine learning community. What audience will learn: When and how to use a variety of machine learning interpretability methods through case studies of real-world situations. The difference between glass-box and black-box explanation methods and when to use them. How to use open source interpretability toolkits that are now available Rich Caruana, Scott M. Lundberg, Marco Túlio Ribeiro, Harsha Nori, Samuel Jenkins |
KDD | 3 |
| 2019 | Are Red Roses Red? Evaluating Consistency of Question-Answering ModelsabstractAlthough current evaluation of questionanswering systems treats predictions in isolation, we need to consider the relationship between predictions to measure true understanding.A model should be penalized for answering "no" to "Is the rose red?" if it answers "red" to "What color is the rose?".We propose a method to automatically extract such implications for instances from two QA datasets, VQA and SQuAD, which we then use to evaluate the consistency of models.Human evaluation shows these generated implications are well formed and valid.Consistency evaluation provides crucial insights into gaps in existing models, and retraining with implicationaugmented data improves consistency on both synthetic and human-generated implications. Marco Túlio Ribeiro, Carlos Guestrin, Sameer Singh 0001 |
ACL (1) | 1 |
| 2019 | Errudite: Scalable, Reproducible, and Testable Error AnalysisabstractThough error analysis is crucial to understanding and improving NLP models, the common practice of manual, subjective categorization of a small sample of errors can yield biased and incomplete conclusions.This paper codifies model and task agnostic principles for informative error analysis, and presents Errudite, an interactive tool for better supporting this process.First, error groups should be precisely defined for reproducibility; Errudite supports this with an expressive domainspecific language.Second, to avoid spurious conclusions, a large set of instances should be analyzed, including both positive and negative examples; Errudite enables systematic grouping of relevant instances with filtering queries.Third, hypotheses about the cause of errors should be explicitly tested; Errudite supports this via automated counterfactual rewriting.We validate our approach with a user study, finding that Errudite (1) enables users to perform high quality and reproducible error analyses with less effort, (2) reveals substantial ambiguities in prior published error analyses practices, and (3) enhances the error analysis experience by allowing users to test and revise prior beliefs. Sherry Tongshuang Wu, Marco Túlio Ribeiro, Jeffrey Heer, Daniel S. Weld |
ACL (1) | 2 |
| 2018 | Anchors: High-Precision Model-Agnostic ExplanationsabstractWe introduce a novel model-agnostic system that explains the behavior of complex models with high-precision rules called anchors, representing local, "sufficient" conditions for predictions. We propose an algorithm to efficiently compute these explanations for any black-box model with high-probability guarantees. We demonstrate the flexibility of anchors by explaining a myriad of different models for different domains and tasks. In a user study, we show that anchors enable users to predict how a model would behave on unseen instances with less effort and higher precision, as compared to existing linear explanations or no explanations. Marco Túlio Ribeiro, Sameer Singh 0001, Carlos Guestrin |
AAAI | 1 |
| 2018 | Semantically Equivalent Adversarial Rules for Debugging NLP modelsabstractComplex machine learning models for NLP are often brittle, making different predictions for input instances that are extremely similar semantically. To automatically detect this behavior for individual instances, we present semantically equivalent adversaries (SEAs) – semantic-preserving perturbations that induce changes in the model’s predictions. We generalize these adversaries into semantically equivalent adversarial rules (SEARs) – simple, universal replacement rules that induce adversaries on many instances. We demonstrate the usefulness and flexibility of SEAs and SEARs by detecting bugs in black-box state-of-the-art models for three domains: machine comprehension, visual question-answering, and sentiment analysis. Via user studies, we demonstrate that we generate high-quality local adversaries for more instances than humans, and that SEARs induce four times as many mistakes as the bugs discovered by human experts. SEARs are also actionable: retraining models using data augmentation significantly reduces bugs, while maintaining accuracy. Marco Túlio Ribeiro, Sameer Singh 0001, Carlos Guestrin |
ACL (1) | 1 |
| 2016 | "Why Should I Trust You?": Explaining the Predictions of Any ClassifierabstractDespite widespread adoption, machine learning models remain mostly black boxes. Understanding the reasons behind predictions is, however, quite important in assessing trust, which is fundamental if one plans to take action based on a prediction, or when choosing whether to deploy a new model. Such understanding also provides insights into the model, which can be used to transform an untrustworthy model or prediction into a trustworthy one. Marco Túlio Ribeiro, Sameer Singh 0001, Carlos Guestrin |
KDD | 1 |
| 2014 | Multiobjective Pareto-Efficient Approaches for Recommender SystemsabstractRecommender systems are quickly becoming ubiquitous in applications such as e-commerce, social media channels, and content providers, among others, acting as an enabling mechanism designed to overcome the information overload problem by improving browsing and consumption experience. A typical task in many recommender systems is to output a ranked list of items, so that items placed higher in the rank are more likely to be interesting to the users. Interestingness measures include how accurate, novel, and diverse are the suggested items, and the objective is usually to produce ranked lists optimizing one of these measures. Suggesting items that are simultaneously accurate, novel, and diverse is much more challenging, since this may lead to a conflicting-objective problem, in which the attempt to improve a measure further may result in worsening other measures. In this article, we propose new approaches for multiobjective recommender systems based on the concept of Pareto efficiency—a state achieved when the system is devised in the most efficient manner in the sense that there is no way to improve one of the objectives without making any other objective worse off. Given that existing multiobjective recommendation algorithms differ in their level of accuracy, diversity, and novelty, we exploit the Pareto-efficiency concept in two distinct manners: (i) the aggregation of ranked lists produced by existing algorithms into a single one, which we call Pareto-efficient ranking, and (ii) the weighted combination of existing algorithms resulting in a hybrid one, which we call Pareto-efficient hybridization. Our evaluation involves two real application scenarios: music recommendation with implicit feedback (i.e., Last.fm) and movie recommendation with explicit feedback (i.e., MovieLens). We show that the proposed Pareto-efficient approaches are effective in suggesting items that are likely to be simultaneously accurate, diverse, and novel. We discuss scenarios where the system achieves high levels of diversity and novelty without compromising its accuracy. Further, comparison against multiobjective baselines reveals improvements in terms of accuracy (from 10.4% to 10.9%), novelty (from 5.7% to 7.5%), and diversity (from 1.6% to 4.2%). Marco Túlio Ribeiro, Nivio Ziviani, Edleno Silva de Moura, Itamar Hata, Anísio Lacerda, Adriano Veloso |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2012 | Pareto-efficient hybridization for multi-objective recommender systemsabstractPerforming accurate suggestions is an objective of paramount importance for effective recommender systems. Other important and increasingly evident objectives are novelty and diversity, which are achieved by recommender systems that are able to suggest diversified items not easily discovered by the users. Different recommendation algorithms have particular strengths and weaknesses when it comes to each of these objectives, motivating the construction of hybrid approaches. However, most of these approaches only focus on optimizing accuracy, with no regard for novelty and diversity. The problem of combining recommendation algorithms grows significantly harder when multiple objectives are considered simultaneously. For instance, devising multi-objective recommender systems that suggest items that are simultaneously accurate, novel and diversified may lead to a conflicting-objective problem, where the attempt to improve an objective further may result in worsening other competing objectives. In this paper we propose a hybrid recommendation approach that combines existing algorithms which differ in their level of accuracy, novelty and diversity. We employ an evolutionary search for hybrids following the Strength Pareto approach, which isolates hybrids that are not dominated by others (i.e., the so called Pareto frontier). Experimental results on two recommendation scenarios show that: (i) we can combine recommendation algorithms in order to improve an objective without significantly hurting other objectives, and (ii) we allow for adjusting the compromise between accuracy, diversity and novelty, so that the recommendation emphasis can be adjusted dynamically according to the needs of different users. Marco Túlio Ribeiro, Anísio Lacerda, Adriano Veloso, Nivio Ziviani |
RecSys | 1 |