VLDB 2026 Research / reviewers in the wild / expert
Xiaoyu Zhang 0014
dblp:12/5927-14
· DBLP profile ↗
12ranked-venue papers
7as first author
11since 2021 · last 2026
0000-0002-8057-3997ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 7 · 5 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PersonaMail: Learning and Adapting Personal Communication Preferences for Context-Aware Email WritingabstractLLM-assisted writing has seen rapid adoption in interpersonal communication, yet current systems often fail to capture the subtle tones essential for effectiveness. Email writing exemplifies this challenge: effective messages require careful alignment with intent, relationship, and context beyond mere fluency. Through formative studies, we identified three key challenges: articulating nuanced communicative intent, making modifications at multiple levels of granularity, and reusing effective tone strategies across messages. We developed PersonaMail, a system that addresses these gaps through structured communication factor exploration, granular editing controls, and adaptive reuse of successful strategies. Our evaluation compared PersonaMail against standard LLM interfaces, and showed improved efficiency in both immediate and repeated use, alongside higher user satisfaction. We contribute design implications for AI-assisted communication systems that prioritize interpersonal nuance over generic text generation. Qiuyuan Ren, Felicia Fang-Yi Tan, Yang Chen 0054, Xiaoyu Zhang 0014, Shengdong Zhao 0001 |
IUI | 5 |
| 2025 | "It's impressive, but in practice...": Experiencing a Realistic Digital Transformation in and beyond the ClassroomabstractSerious games, particularly board games, have long been employed in production management education to teach various concepts.While they have demonstrated educational effectiveness, their integration with emerging Industry 4.0 technologies remains limited.Furthermore, there is a lack of empirical research on how industry practitioners apply these digitization technologies in the workplace.To bridge this gap, we designed a course that integrates digital technologies into a traditional board game.We conducted two studies to evaluate both knowledge gains within the classroom and knowledge transfer back into the manufacturing industry.Our results show an improved understanding of the synergies between production management principles and Industry 4.0 technologies, as well as the real-world challenges students face when attempting to transfer this knowledge.Our work contributes pedagogical and practical perspectives on how technology-enhanced serious games can extend learning in and beyond the classroom. Xiaoyu Zhang 0014, Alexander Albers, Torbjørn H. Netland |
CHI | 1 |
| 2025 | MRISA: A Visual Analytics Approach of Locomotion Policies Comparison for Robotics TrainingabstractIn the field of robotics, the process of locomotion control policy training is inherently iterative and exploratory. Practitioners often switch between multiple simulation and data analysis tools to observe robot postures and behaviors, track part movements, and compare reward data, which is a tedious process. To better understand their hurdles and requirements, we interviewed five robotics experts and analyzed representative figures from recent robotics publications to identify prevailing challenges and strategies for comparing and communicating locomotion policies. The main challenges include the lack of integrated simulation and visualization, difficulty in comparing multiple policies simultaneously, and the time-intensive process of creating polished visual presentations. Based on the insights, we introduce MRISA (Multi-Robot Interactive Simulation and Analysis Platform), an interactive tool designed to support exploratory analysis on pre-trained locomotion policies. MRISA integrates features including direct observation of one or multiple robots’ behaviors in a simulator, trajectories visualization with customized anchors, key measurement inspection in a timeline view, and key frames capturing. A user evaluation with 14 domain practitioners demonstrated that MRISA provides immediate insights, enabling practitioners to intuitively explore multiple dimensions of locomotion policies. Fan Shi 0002, Xiaoyu Zhang 0014, April Yi Wang |
Graphics Interface | 3 |
| 2024 | How to Engage your Readers? Generating Guiding Questions to Promote Active ReadingabstractUsing questions in written text is an effective strategy to enhance readability.However, what makes an active reading question good, what the linguistic role of these questions is, and what is their impact on human reading remains understudied.We introduce GUIDINGQ, a dataset of 10K in-text questions from textbooks and scientific articles.By analyzing the dataset, we present a comprehensive understanding of the use, distribution, and linguistic characteristics of these questions.Then, we explore various approaches to generate such questions using language models.Our results highlight the importance of capturing inter-question relationships and the challenge of question position identification in generating these questions.Finally, we conduct a human study to understand the implication of such questions on reading comprehension.We find that the generated questions are of high quality and are almost as effective as human-written questions in terms of improving readers' memorization and comprehension.github.com/eth-lre/engage-your-readers Questions in titles:How do Philosophers arrive at truth?Is there no quantum form of Einstein Gravity?Why do house-hunting ants recruit in both directions? Peng Cui 0006, Vilém Zouhar, Xiaoyu Zhang 0014, Mrinmaya Sachan |
ACL (1) | 3 |
| 2024 | Slicing, Chatting, and Refining: A Concept-Based Approach for Machine Learning Model Validation with ConceptSlicerabstractAs machine learning (ML) gains wider adoption in real-world applications, the validation of ML models becomes fundamental for its productization, particularly in safety-critical applications. Recently, data slice finding has emerged as a popular method for validating ML models, but it requires additional metadata or cross-modal embeddings for the slices to be interpretable. We propose ConceptSlicer, an integrated workflow that facilitates the slicing of computer vision models using visual concepts. This approach breaks down the image dataset into interpretable visual concepts, serving as metadata in the slice finding process. Our system offers insights into model issues and enables a deeper understanding of computer vision models’ strengths and weaknesses. We evaluate ConceptSlicer through interviews with eight domain experts and machine learning practitioners, and fine-tune the ML models based on their feedback. Our study also highlights varied attitudes towards large foundational models, encouraging contemplation of the challenges and opportunities presented by this technological advancement. Xiaoyu Zhang 0014, Jorge Henrique Piazentin Ono, Liang Gou, Mrinmaya Sachan, Kwan-Liu Ma, Liu Ren 0001 |
IUI | 1 |
| 2023 | ConceptEVA: Concept-Based Interactive Exploration and Customization of Document SummariesabstractWith the most advanced natural language processing and artificial intelligence approaches, effective summarization of long and multi-topic documents—such as academic papers—for readers from different domains still remains a challenge. To address this, we introduce ConceptEVA, a mixed-initiative approach to generate, evaluate, and customize summaries for long and multi-topic documents. ConceptEVA incorporates a custom multi-task longformer encoder decoder to summarize longer documents. Interactive visualizations of document concepts as a network reflecting both semantic relatedness and co-occurrence help users focus on concepts of interest. The user can select these concepts and automatically update the summary to emphasize them. We present two iterations of ConceptEVA evaluated through an expert review and a within-subjects study. We find that participants’ satisfaction with customized summaries through ConceptEVA is higher than their own manually-generated summary, while incorporating critique into the summaries proved challenging. Based on our findings, we make recommendations for designing summarization systems incorporating mixed-initiative interactions. Xiaoyu Zhang 0014, Jianping Kelvin Li, Po-Wei Chi, Senthil K. Chandrasegaran, Kwan-Liu Ma |
CHI | 1 |
| 2023 | LabelVizier: Interactive Validation and Relabeling for Technical Text AnnotationsabstractWith the rapid accumulation of text data produced by data-driven techniques, the task of extracting "data annotations"—concise, high-quality data summaries from unstructured raw text—has become increasingly important. The recent advances in weak supervision and crowd-sourcing techniques provide promising solutions to efficiently create annotations (labels) for large-scale technical text data. However, such annotations may fail in practice because of the change in annotation requirements, application scenarios, and modeling goals, where label validation and relabeling by domain experts are required. To approach this issue, we present LabelVizier, a human-in-the-loop workflow that incorporates domain knowledge and user-specific requirements to reveal actionable insights into annotation flaws, then produce better-quality labels for large-scale multi-label datasets. We implement our workflow as an interactive notebook to facilitate flexible error profiling, in-depth annotation validation for three error types, and efficient annotation relabeling on different data scales. We evaluated our workflow in assisting the validation and relabelling of technical text annotation with two use cases and four expert reviews. The results show that LabelVizier is applicable in various application scenarios, and users with different knowledge backgrounds have diverse preferences for the tool usage. Xiaoyu Zhang 0014, Xiwei Xuan, Alden Dima, Thurston Sexton, Kwan-Liu Ma |
PacificVis | 1 |
| 2023 | SliceTeller: A Data Slice-Driven Approach for Machine Learning Model ValidationabstractReal-world machine learning applications need to be thoroughly evaluated to meet critical product requirements for model release, to ensure fairness for different groups or individuals, and to achieve a consistent performance in various scenarios. For example, in autonomous driving, an object classification model should achieve high detection rates under different conditions of weather, distance, etc. Similarly, in the financial setting, credit-scoring models must not discriminate against minority groups. These conditions or groups are called as "Data Slices". In product MLOps cycles, product developers must identify such critical data slices and adapt models to mitigate data slice problems. Discovering where models fail, understanding why they fail, and mitigating these problems, are therefore essential tasks in the MLOps life-cycle. In this paper, we present SliceTeller, a novel tool that allows users to debug, compare and improve machine learning models driven by critical data slices. SliceTeller automatically discovers problematic slices in the data, helps the user understand why models fail. More importantly, we present an efficient algorithm, SliceBoosting, to estimate trade-offs when prioritizing the optimization over certain slices. Furthermore, our system empowers model developers to compare and analyze different model versions during model iterations, allowing them to choose the model version best suitable for their applications. We evaluate our system with three use cases, including two real-world use cases of product development, to demonstrate the power of SliceTeller in the debugging and improvement of product-quality ML models. Xiaoyu Zhang 0014, Jorge Henrique Piazentin Ono, Huan Song, Liang Gou, Kwan-Liu Ma, Liu Ren 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2022 | VAC-CNN: A Visual Analytics System for Comparative Studies of Deep Convolutional Neural NetworksabstractThe rapid development of Convolutional Neural Networks (CNNs) in recent years has triggered significant breakthroughs in many machine learning (ML) applications. The ability to understand and compare various CNN models available is thus essential. The conventional approach with visualizing each model's quantitative features, such as classification accuracy and computational complexity, is not sufficient for a deeper understanding and comparison of the behaviors of different models. Moreover, most of the existing tools for assessing CNN behaviors only support comparison between two models and lack the flexibility of customizing the analysis tasks according to user needs. This paper presents a visual analytics system, VAC-CNN (Visual Analytics for Comparing CNNs), that supports the in-depth inspection of a single CNN model as well as comparative studies of two or more models. The ability to compare a larger number of (e.g., tens of) models especially distinguishes our system from previous ones. With a carefully designed model visualization and explaining support, VAC-CNN facilitates a highly interactive workflow that promptly presents both quantitative and qualitative information at each analysis stage. We demonstrate VAC-CNN's effectiveness for assisting novice ML practitioners in evaluating and comparing multiple CNN models through two use cases and one preliminary evaluation study using the image classification tasks on the ImageNet dataset. Xiwei Xuan, Xiaoyu Zhang 0014, Oh-Hyun Kwon, Kwan-Liu Ma |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2021 | A Visual Analytics Approach for the Diagnosis of Heterogeneous and Multidimensional Machine Maintenance DataabstractAnalysis of large, high-dimensional, and heterogeneous datasets is challenging as no one technique is suitable for visualizing and clustering such data in order to make sense of the underlying information. For instance, heterogeneous logs detailing machine repair and maintenance in an organization often need to be analyzed to diagnose errors and identify abnormal patterns, formalize root-cause analyses, and plan preventive maintenance. Such real-world datasets are also beset by issues such as inconsistent and/or missing entries. To conduct an effective diagnosis, it is important to extract and understand patterns from the data with support from analytic algorithms (e.g., finding that certain kinds of machine complaints occur more in the summer) while involving the human-in-the-loop. To address these challenges, we adopt existing techniques for dimensionality reduction (DR) and clustering of numerical, categorical, and text data dimensions, and introduce a visual analytics approach that uses multiple coordinated views to connect DR + clustering results across each kind of the data dimension stated. To help analysts label the clusters, each clustering view is supplemented with techniques and visualizations that contrast a cluster of interest with the rest of the dataset. Our approach assists analysts to make sense of machine maintenance logs and their errors. Then the gained insights help them carry out preventive maintenance. We illustrate and evaluate our approach through use cases and expert studies respectively, and discuss generalization of the approach to other heterogeneous data. Xiaoyu Zhang 0014, Takanori Fujiwara, Senthil K. Chandrasegaran, Michael Brundage, Thurston Sexton, Alden Dima, Kwan-Liu Ma |
PacificVis | 1 |
| 2021 | ConceptScope: Organizing and Visualizing Knowledge in Documents based on Domain OntologyabstractCurrent text visualization techniques typically provide overviews of document content and structure using intrinsic properties such as term frequencies, co-occurrences, and sentence structures. Such visualizations lack conceptual overviews incorporating domain-relevant knowledge, needed when examining documents such as research articles or technical reports. To address this shortcoming, we present ConceptScope, a technique that utilizes a domain ontology to represent the conceptual relationships in a document in the form of a Bubble Treemap visualization. Multiple coordinated views of document structure and concept hierarchy with text overviews further aid document analysis. ConceptScope facilitates exploration and comparison of single and multiple documents respectively. We demonstrate ConceptScope by visualizing research articles and transcripts of technical presentations in computer science. In a comparative study with DocuBurst, a popular document visualization tool, ConceptScope was found to be more informative in exploring and comparing domain-specific documents, but less so when it came to documents that spanned multiple disciplines. Xiaoyu Zhang 0014, Senthil K. Chandrasegaran, Kwan-Liu Ma |
CHI | 1 |
| 2020 | Text-to-Viz: Automatic Generation of Infographics from Proportion-Related Natural Language StatementsabstractCombining data content with visual embellishments, infographics can effectively deliver messages in an engaging and memorable manner. Various authoring tools have been proposed to facilitate the creation of infographics. However, creating a professional infographic with these authoring tools is still not an easy task, requiring much time and design expertise. Therefore, these tools are generally not attractive to casual users, who are either unwilling to take time to learn the tools or lacking in proper design expertise to create a professional infographic. In this paper, we explore an alternative approach: to automatically generate infographics from natural language statements. We first conducted a preliminary study to explore the design space of infographics. Based on the preliminary study, we built a proof-of-concept system that automatically converts statements about simple proportion-related statistics to a set of infographics with pre-designed styles. Finally, we demonstrated the usability and usefulness of the system through sample results, exhibits, and expert reviews. Weiwei Cui 0001, Xiaoyu Zhang 0014, Yun Wang 0012, Bei Chen 0008, Lei Fang 0004, Jian-Guang Lou, Dongmei Zhang 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |