VLDB 2026 Research / reviewers in the wild / expert
Jorge Henrique Piazentin Ono
dblp:170/3072 · also Jorge Piazentin Ono
· DBLP profile ↗
12ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-2424-0186ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VCR: Interpretable and interactive debugging of object detection models with visual concepts
Jie Jeff Xu, Saahir Dhanani, Jorge Henrique Piazentin Ono, Liu Ren 0001, Kexin Rong 0001 |
Inf. Syst. | 3 |
| 2025 | InterChat: Enhancing Generative Visual Analytics using Multimodal InteractionsabstractAbstract The rise of Large Language Models (LLMs) and generative visual analytics systems has transformed data‐driven insights, yet significant challenges persist in accurately interpreting users analytical and interaction intents. While language inputs offer flexibility, they often lack precision, making the expression of complex intents inefficient, error‐prone, and time‐intensive. To address these limitations, we investigate the design space of multimodal interactions for generative visual analytics through a literature review and pilot brainstorming sessions. Building on these insights, we introduce a highly extensible workflow that integrates multiple LLM agents for intent inference and visualization generation. We develop InterChat, a generative visual analytics system that combines direct manipulation of visual elements with natural language inputs. This integration enables precise intent communication and supports progressive, visually driven exploratory data analyses. By employing effective prompt engineering, and contextual interaction linking, alongside intuitive visualization and interaction designs, InterChat bridges the gap between user interactions and LLM‐driven visualizations, enhancing both interpretability and usability. Extensive evaluations, including two usage scenarios, a user study, and expert feedback, demonstrate the effectiveness of InterChat. Results show significant improvements in the accuracy and efficiency of handling complex visual analytics tasks, highlighting the potential of multimodal interactions to redefine user engagement and analytical depth in generative visual analytics. Juntong Chen, Jiang Wu 0012, Jiajing Guo, Vikram Mohanty, Jorge Henrique Piazentin Ono, Liu Ren 0001, Dongyu Liu |
Comput. Graph. Forum | 6 |
| 2025 | VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and AnalysisabstractAbstract Real‐world machine learning models require rigorous evaluation before deployment, especially in safety‐critical domains like autonomous driving and surveillance. The evaluation of machine learning models often focuses on data slices, which are subsets of the data that share a set of characteristics. Data slice finding automatically identifies conditions or data subgroups where models underperform, aiding developers in mitigating performance issues. Despite its popularity and effectiveness, data slicing for vision model validation faces several challenges. First, data slicing often needs additional image metadata or visual concepts, and falls short in certain computer vision tasks, such as object detection. Second, understanding data slices is a labor‐intensive and mentally demanding process that heavily relies on the expert's domain knowledge. Third, data slicing lacks a human‐in‐the‐loop solution that allows experts to form hypothesis and test them interactively. To overcome these limitations and better support the machine learning operations lifecycle, we introduce VISLIX, a novel visual analytics framework that employs state‐of‐the‐art foundation models to help domain experts analyze slices in computer vision models. Our approach does not require image metadata or visual concepts, automatically generates natural language insights, and allows users to test data slice hypothesis interactively. We evaluate VISLIX with an expert study and three use cases, that demonstrate the effectiveness of our tool in providing comprehensive insights for validating object detection models. Xinyuan Yan, Xiwei Xuan, Jorge Henrique Piazentin Ono, Jiajing Guo, Vikram Mohanty, Arvind Kumar Shekar, Liang Gou, Bei Wang 0001, Liu Ren 0001 |
Comput. Graph. Forum | 3 |
| 2025 | AttributionScanner: A Visual Analytics System for Model Validation With Metadata-Free Slice FindingabstractData slice finding is an emerging technique for validating machine learning (ML) models by identifying and analyzing subgroups in a dataset that exhibit poor performance, often characterized by distinct feature sets or descriptive metadata. However, in the context of validating vision models involving unstructured image data, this approach faces significant challenges, including the laborious and costly requirement for additional metadata and the complex task of interpreting the root causes of underperformance. To address these challenges, we introduce AttributionScanner, an innovative human-in-the-loop Visual Analytics (VA) system, designed for metadata-free data slice finding. Our system identifies interpretable data slices that involve common model behaviors and visualizes these patterns through an Attribution Mosaic design. Our interactive interface provides straightforward guidance for users to detect, interpret, and annotate predominant model issues, such as spurious correlations (model biases) and mislabeled data, with minimal effort. Additionally, it employs a cutting-edge model regularization technique to mitigate the detected issues and enhance the model's performance. The efficacy of AttributionScanner is demonstrated through use cases involving two benchmark datasets, with qualitative and quantitative evaluations showcasing its substantial effectiveness in vision model validation, ultimately leading to more reliable and accurate models. Xiwei Xuan, Jorge Henrique Piazentin Ono, Liang Gou, Kwan-Liu Ma, Liu Ren 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2025 | VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data LabelsabstractThe advances in multi-modal foundation models (FMs) (e.g., CLIP and LLaVA) have facilitated the auto-labeling of large-scale datasets, enhancing model performance in challenging downstream tasks such as open-vocabulary object detection and segmentation. However, the quality of FM-generated labels is less studied as existing approaches focus more on data quantity over quality. This is because validating large volumes of data without ground truth presents a considerable challenge in practice. Existing methods typically rely on limited metrics to identify problematic data, lacking a comprehensive perspective, or apply human validation to only a small data fraction, failing to address the full spectrum of potential issues. To overcome these challenges, we introduce VISTA, a visual analytics framework that improves data quality to enhance the performance of multi-modal models. Targeting the complex and demanding domain of open-vocabulary image segmentation, VISTA integrates multi-phased data validation strategies with human expertise, enabling humans to identify, understand, and correct hidden issues within FM-generated labels. Through detailed use cases on two benchmark datasets and expert reviews, we demonstrate VISTA's effectiveness from both quantitative and qualitative perspectives. Xiwei Xuan, Jorge Henrique Piazentin Ono, Liang Gou, Kwan-Liu Ma, Liu Ren 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | USE: Universal Segment Embeddings for Open-Vocabulary Image SegmentationabstractThe open-vocabulary image segmentation task involves partitioning images into semantically meaningful segments and classifying them with flexible text-defined categories. The recent vision-based foundation models such as the Segment Anything Model (SAM) have shown superior performance in generating class-agnostic image segments. The main challenge in open-vocabulary image segmentation now lies in accurately classifying these segments into text-defined categories. In this paper, we introduce the Universal Segment Embedding (USE) framework to address this challenge. This framework is comprised of two key components: 1) a data pipeline designed to efficiently curate a large amount of segment-text pairs at various granularities, and 2) a universal segment embedding model that enables precise segment classification into a vast range of text-defined categories. The USE model can not only help open-vocabulary image segmentation but also facilitate other downstream tasks (e.g., querying and ranking). Through comprehensive experimental studies on semantic segmentation and part segmentation benchmarks, we demonstrate that the USE framework outperforms state-of-the-art open-vocabulary segmentation methods. Xiwei Xuan, Clint Sebastian, Jorge Henrique Piazentin Ono, Sima Behpour, Thang Doan, Liang Gou, Han-Wei Shen, Liu Ren 0001 |
CVPR | 5 |
| 2024 | Slicing, Chatting, and Refining: A Concept-Based Approach for Machine Learning Model Validation with ConceptSlicerabstractAs machine learning (ML) gains wider adoption in real-world applications, the validation of ML models becomes fundamental for its productization, particularly in safety-critical applications. Recently, data slice finding has emerged as a popular method for validating ML models, but it requires additional metadata or cross-modal embeddings for the slices to be interpretable. We propose ConceptSlicer, an integrated workflow that facilitates the slicing of computer vision models using visual concepts. This approach breaks down the image dataset into interpretable visual concepts, serving as metadata in the slice finding process. Our system offers insights into model issues and enables a deeper understanding of computer vision models’ strengths and weaknesses. We evaluate ConceptSlicer through interviews with eight domain experts and machine learning practitioners, and fine-tune the ML models based on their feedback. Our study also highlights varied attitudes towards large foundational models, encouraging contemplation of the challenges and opportunities presented by this technological advancement. Xiaoyu Zhang 0014, Jorge Henrique Piazentin Ono, Liang Gou, Mrinmaya Sachan, Kwan-Liu Ma, Liu Ren 0001 |
IUI | 2 |
| 2024 | Demonstration of VCR: A Tabular Data Slicing Approach to Understanding Object Detection Model PerformanceabstractIn this demonstration, we present VCR, an automated slice discovery method (SDM) for object detection models that helps practitioners identify and explain specific scenarios in which their models exhibit systematic errors. VCR leverages the capabilities of vision foundation models to generate segment-level visual concepts that serve as interpretable explanation primitives. By integrating these visual concepts with additional image metadata in a tabular format, VCR uses a scalable frequent itemset mining-based technique to identify common patterns associated with model performance. We will demonstrate VCR's capabilities through three usage scenarios. First, users can explore the automatically extracted visual concepts and their associated labels. Second, users can run slice finding on a large object detection dataset and visually inspect the results to discover systematic errors. Finally, users can iteratively refine their slicing results by providing feedback on the granularity of visual concepts and the quality of the generated labels. These scenarios will illustrate how VCR can aid practitioners in discovering non-trivial gaps in their models' performance, providing actionable insights for model improvement. Jie Jeff Xu, Saahir Dhanani, Jorge Henrique Piazentin Ono, Liu Ren 0001, Kexin Rong 0001 |
Proc. VLDB Endow. | 3 |
| 2023 | SliceTeller: A Data Slice-Driven Approach for Machine Learning Model ValidationabstractReal-world machine learning applications need to be thoroughly evaluated to meet critical product requirements for model release, to ensure fairness for different groups or individuals, and to achieve a consistent performance in various scenarios. For example, in autonomous driving, an object classification model should achieve high detection rates under different conditions of weather, distance, etc. Similarly, in the financial setting, credit-scoring models must not discriminate against minority groups. These conditions or groups are called as "Data Slices". In product MLOps cycles, product developers must identify such critical data slices and adapt models to mitigate data slice problems. Discovering where models fail, understanding why they fail, and mitigating these problems, are therefore essential tasks in the MLOps life-cycle. In this paper, we present SliceTeller, a novel tool that allows users to debug, compare and improve machine learning models driven by critical data slices. SliceTeller automatically discovers problematic slices in the data, helps the user understand why models fail. More importantly, we present an efficient algorithm, SliceBoosting, to estimate trade-offs when prioritizing the optimization over certain slices. Furthermore, our system empowers model developers to compare and analyze different model versions during model iterations, allowing them to choose the model version best suitable for their applications. We evaluate our system with three use cases, including two real-world use cases of product development, to demonstrate the power of SliceTeller in the debugging and improvement of product-quality ML models. Xiaoyu Zhang 0014, Jorge Henrique Piazentin Ono, Huan Song, Liang Gou, Kwan-Liu Ma, Liu Ren 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | PipelineProfiler: A Visual Analytics Tool for the Exploration of AutoML PipelinesabstractIn recent years, a wide variety of automated machine learning (AutoML) methods have been proposed to generate end-to-end ML pipelines. While these techniques facilitate the creation of models, given their black-box nature, the complexity of the underlying algorithms, and the large number of pipelines they derive, they are difficult for developers to debug. It is also challenging for machine learning experts to select an AutoML system that is well suited for a given problem. In this paper, we present the Pipeline Profiler, an interactive visualization tool that allows the exploration and comparison of the solution space of machine learning (ML) pipelines produced by AutoML systems. PipelineProfiler is integrated with Jupyter Notebook and can be combined with common data science tools to enable a rich set of analyses of the ML pipelines, providing users a better understanding of the algorithms that generated them as well as insights into how they can be improved. We demonstrate the utility of our tool through use cases where PipelineProfiler is used to better understand and improve a real-world AutoML system. Furthermore, we validate our approach by presenting a detailed analysis of a think-aloud experiment with six data scientists who develop and evaluate AutoML tools. Jorge Henrique Piazentin Ono, Sonia Castelo Quispe, Roque Lopez, Enrico Bertini, Juliana Freire, Cláudio T. Silva |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2019 | HistoryTracker: Minimizing Human Interactions in Baseball Game AnnotationabstractThe sport data tracking systems available today are based on specialized hardware (high-definition cameras, speed radars, RFID) to detect and track targets on the field. While effective, implementing and maintaining these systems pose a number of challenges, including high cost and need for close human monitoring. On the other hand, the sports analytics community has been exploring human computation and crowdsourcing in order to produce tracking data that is trustworthy, cheaper and more accessible. However, state-of-the-art methods require a large number of users to perform the annotation, or put too much burden into a single user. We propose HistoryTracker, a methodology that facilitates the creation of tracking data for baseball games by warm-starting the annotation process using a vast collection of historical data. We show that HistoryTracker helps users to produce tracking data in a fast and reliable way. Jorge Henrique Piazentin Ono, Arvi Gjoka, Justin Salamon, Carlos A. Dietrich, Cláudio T. Silva |
CHI | 1 |
| 2018 | Baseball Timeline: Summarizing Baseball Plays Into a Static VisualizationabstractAbstract In sports, Play Diagrams are the standard way to represent and convey information. They are widely used by coaches, managers, journalists and fans in general. There are situations where diagrams may be hard to understand, for example, when several actions are packed in a certain region of the field or there are just too many actions to be transformed in a clear depiction of the play. The representation of how actions develop through time, in particular, may be hardly achieved on such diagrams. The time, and the relationship among the actions of the players through time, is critical on the depiction of complex plays. In this context, we present a study on how player actions may be clearly depicted on 2D diagrams. The study is focused on Baseball plays, a sport where diagrams are heavily used to summarize the actions of the players. We propose a new and simple approach to represent spatiotemporal information in the form of a timeline. We designed our visualization with a requirement driven approach, conducting interviews and fulfilling the needs of baseball experts and expert‐fans. We validate our approach by presenting a detailed analysis of baseball plays and conducting interviews with four domain experts. Jorge Henrique Piazentin Ono, Carlos A. Dietrich, Cláudio T. Silva |
Comput. Graph. Forum | 1 |