VLDB 2026 Research / reviewers in the wild / expert
Zhongyi Zhou
dblp:01/2600
· DBLP profile ↗
20ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | See2Refine: Vision-Language Feedback Improves LLM-Based eHMI Action DesignersabstractDing Xia, Xinyue Gui, Mark Colley, Fan Gao, Zhongyi Zhou, Dongyuan Li, Renhe Jiang, Takeo Igarashi. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ding Xia, Xinyue Gui, Mark Colley, Zhongyi Zhou, Dongyuan Li, Renhe Jiang, Takeo Igarashi |
ACL (1) | 5 |
| 2026 | Peeking Ahead of the Field Study: Exploring VLM Personas as Support Tools for Embodied Studies in HCIabstractField studies are irreplaceable but costly, time-consuming, and error-prone, which need careful preparation. Inspired by rapid-prototyping in manufacturing, we propose a fast, low-cost evaluation method using Vision-Language Model (VLM) personas to simulate outcomes comparable to field results. While LLMs show human-like reasoning and language capabilities, autonomous vehicle (AV)-pedestrian interaction requires spatial awareness, emotional empathy, and behavioral generation. This raises our research question: To what extent can VLM personas mimic human responses in field studies? We conducted parallel studies: 1) one real-world study with 20 participants, and 2) one video-study using 20 VLM personas, both on a street-crossing task. We compared their responses and interviewed five HCI researchers on potential applications. Results show that VLM personas mimic human response patterns (e.g., average crossing times of 5.25 s vs. 5.07 s) lack the behavioral variability and depth. They show promise for formative studies, field study preparation, and human data augmentation. Xinyue Gui, Ding Xia, Mark Colley, Vishal Chauhan, Anubhav, Zhongyi Zhou, Ehsan Javanmardi, Stela Hanbyeol Seo, Chia-Ming Chang 0003, Manabu Tsukada, Takeo Igarashi |
CHI | 7 |
| 2026 | AgentHands: Generating Interactive Hand Gestures for Spatially Grounded Agent Conversations in XRabstractCommunicating spatial tasks via text or speech creates “a mental mapping gap” that limits an agent’s expressiveness. Inspired by co-speech gestures in face-to-face conversation, we propose AgentHands, an LLM-powered XR system that equips agents with hands to render responses clearer and more engaging. Guided by a design taxonomy distilled from a formative study (N=10), we implement a novel pipeline to generate and render a hand agent that augments conversational responses with synchronized, space-aware, and interactive hand gestures: using a meta-instruction, AgentHands generates verbal responses embedded with GestureEvents aligned to specific words; each event specifies gesture type and parameters. At runtime, a parser converts events into time-stamped poses and motions, driving an animation system that renders expressive hands synchronized with speech. In a within-subjects study (N=12), AgentHands increased engagement and made spatially grounded conversations easier to follow compared to a speech-only baseline. Ziyi Liu 0004, David Li 0001, Zhongyi Zhou, David Kim 0002, Ruofei Du, Xun Qian |
CHI | 3 |
| 2026 | Improving Low-Vision Chart Accessibility via On-Cursor Visual ContextabstractDespite widespread use, charts remain largely inaccessible for Low-Vision Individuals (LVI). Reading charts requires viewing data points within a global context, which is difficult for LVI who may rely on magnification or experience a partial field of vision. We aim to improve exploration by providing visual access to critical context. To inform this, we conducted a formative study with five LVI. We identified four fundamental contextual elements common across chart types: axes, legend, grid lines, and the overview. We propose two pointer-based interaction methods to provide this context: Dynamic Context, a novel focus+context interaction, and Mini-map, which adapts overview+detail principles for LVI. In a study with N=22 LVI, we compared both methods and evaluated their integration to current tools. Our results show that Dynamic Context had significant positive impact on access, usability, and effort reduction; however, worsened visual load. Mini-map strengthened spatial understanding, but was less preferred for this task. We offer design insights to guide the development of future systems that support LVI with visual context while balancing visual load. Yotam Sechayk, Hennes Rave, Max Rädler, Mark Colley, Zhongyi Zhou, Ariel Shamir, Takeo Igarashi |
CHI | 5 |
| 2025 | Vision-Based Multimodal Interfaces: A Survey and Taxonomy for Enhanced Context-Aware System DesignabstractFigure 1: We review and categorize VMIs aimed at enhancing context awareness.Our key contribution is a Macro-Micro-Macro level (whole-detail-whole) system design framework, providing actionable references from a Data Modality-Driven perspective: (1) Macro-level contextual factors: considerations for context understanding (Section 3); (2) Micro-level system foundations: input data modality (visual + other modalities), data integration stages, multimodal data processing and evaluation strategies (Sections 4, 5); (3) Macro-level design synthesis: application domains, design considerations and key challenges (Sections 6, 7). Yongquan Hu, Xinya Gong, Zhongyi Zhou, Samitha Elvitigala, Florian 'Floyd' Mueller, Wen Hu 0001, Aaron J. Quigley |
CHI | 4 |
| 2025 | InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMsabstractVisual programming has the potential of providing novice programmers with a low-code experience to build customized processing pipelines.Existing systems typically require users to build pipelines from scratch, implying that novice users are expected to set up and link appropriate nodes from a blank workspace.In this paper, we introduce InstructPipe, an AI assistant for prototyping machine learning (ML) pipelines with text instructions.We contribute two large language model (LLM) modules and a code interpreter as part of our framework.The LLM modules generate pseudocode for a target pipeline, and the interpreter renders the pipeline in the node-graph editor for further human-AI collaboration.Both technical and user evaluation (N=16) shows that InstructPipe empowers users to streamline their ML pipeline workfow, reduce their learning curve, and leverage open-ended commands to spark innovative ideas. Zhongyi Zhou, Vrushank Phadnis, Xiuxiu Yuan, Xun Qian, Kristen Wright, Mark Sherwood, Jason Mayes, Yiyi Huang, Zheng Xu 0002, Yinda Zhang 0001, Johnny Lee, Alex Olwal, David Kim 0002, Ram Iyengar, Na Li 0034, Ruofei Du |
CHI | 1 |
| 2025 | HealthGenie: A Knowledge-Driven LLM Framework for Tailored Dietary GuidanceabstractSeeking dietary guidance often requires navigating complex nutritional knowledge while considering individual health needs. To address this, we present HealthGenie, an interactive platform that leverages the interpretability of knowledge graphs (KGs) and the conversational power of large language models (LLMs) to deliver tailored dietary recommendations alongside integrated nutritional visualizations for fast, intuitive insights. Upon receiving a user query, HealthGenie performs intent refinement and maps user's needs to a curated nutritional knowledge graph. The system then retrieves and visualizes relevant subgraphs, while offering detailed, explainable recommendations. Users can interactively adjust preferences to further tailor results. A within-subject study and quantitative analysis show that HealthGenie reduces cognitive load and interaction effort while supporting personalized, health-aware decision-making. Xinjie Zhao 0004, Ding Xia, Zhongyi Zhou, Rui Yang 0016, Jinghui Lu, Chanjun Park, Irene Li |
CIKM | 4 |
| 2025 | ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action ModelabstractZhongyi Zhou, Yichen Zhu, Minjie Zhu, Junjie Wen, Ning Liu, Zhiyuan Xu, Weibin Meng, Yaxin Peng, Chaomin Shen, Feifei Feng, Yi Xu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zhongyi Zhou, Yichen Zhu 0001, Minjie Zhu, Ning Liu 0007, Weibin Meng, Yaxin Peng, Chaomin Shen 0001, Feifei Feng |
EMNLP | 1 |
| 2025 | Fresh-CL: Feature Realignment through Experts on Hypersphere in Continual LearningabstractContinual Learning enables models to learn and adapt to new tasks while retaining prior knowledge. Introducing new tasks, however, can naturally lead to feature entanglement across tasks, limiting the model’s capability to distinguish between new domain data. In this work, we propose a method called Feature Realignment through Experts on hyperSpHere in Continual Learning (Fresh-CL). By leveraging predefined and fixed simplex equiangular tight frame (ETF) classifiers on a hypersphere, our model improves feature separation both intra and inter tasks. However, the projection to a simplex ETF shifts with new tasks, disrupting structured feature representation of previous tasks and degrading performance. Therefore, we propose a dynamic extension of ETF through mixture of experts, enabling adaptive projections onto diverse subspaces to enhance feature representation. Experiments on 11 datasets demonstrate a 2% improvement in accuracy compared to the strongest baseline, particularly in fine-grained datasets, confirming the efficacy of combining ETF and MoE to improve feature distinction in continual learning scenarios. Zhongyi Zhou, Yaxin Peng, Pin Yi, Minjie Zhu, Chaomin Shen 0001 |
ICASSP | 1 |
| 2025 | DiffusionVLA: Scaling Robot Foundation Models via Unified Diffusion and AutoregressionabstractIn this paper, we present DiffusionVLA, a novel framework that integrates autoregressive reasoning with diffusion policies to address the limitations of existing methods: while autoregressive Vision-Language-Action (VLA) models lack precise and robust action generation, diffusion-based policies inherently lack reasoning capabilities. Central to our approach is autoregressive reasoning — a task decomposition and explanation process enabled by a pre-trained VLM — to guide diffusion-based action policies. To tightly couple reasoning with action generation, we introduce a reasoning injection module that directly embeds self-generated reasoning phrases into the policy learning process. The framework is simple, flexible, and efficient, enabling seamless deployment across diverse robotic platforms.
We conduct extensive experiments using multiple real robots to validate the effectiveness of DiVLA. Our tests include a challenging factory sorting task, where DiVLA successfully categorizes objects, including those not seen during training. The reasoning injection module enhances interpretability, enabling explicit failure diagnosis by visualizing the model’s decision process. Additionally, we test DiVLA on a zero-shot bin-picking task, achieving \textbf{63.7\% accuracy on 102 previously unseen objects}. Our method demonstrates robustness to visual changes, such as distractors and new backgrounds, and easily adapts to new embodiments. Furthermore, DiVLA can follow novel instructions and retain conversational ability. Notably, DiVLA is data-efficient and fast at inference; our smallest DiVLA-2B runs 82Hz on a single A6000 GPU. Finally, we scale the model from 2B to 72B parameters, showcasing improved generalization capabilities with increased model size. Yichen Zhu 0001, Minjie Zhu, Zhibin Tang, Zhongyi Zhou, Chaomin Shen 0001, Yaxin Peng, Feifei Feng |
ICML | 6 |
| 2025 | ChatVLA-2: Vision-Language-Action Model with Open-World ReasoningabstractVision-language-action (VLA) models have emerged as the next generation of models in robotics. However, despite leveraging powerful pre-trained Vision-Language Models (VLMs), existing end-to-end VLA systems often lose key capabilities during fine-tuning as the model adapts to specific robotic tasks. We argue that a generalizable VLA model should retain and expand upon the VLM's core competencies: 1) **Open-world reasoning** - the VLA should inherit the knowledge from VLM, i.e., recognize anything that the VLM can recognize, capable of solving math problems, possessing visual-spatial intelligence, 2) **Reasoning following** – effectively translating the open-world reasoning into actionable steps for the robot. In this work, we introduce **ChatVLA-2**, a novel mixture-of-expert VLA model coupled with a specialized three-stage training pipeline designed to preserve the VLM’s original strengths while enabling actionable reasoning. To validate our approach, we design a math-matching task wherein a robot interprets math problems written on a whiteboard and picks corresponding number cards from a table to solve equations. Remarkably, our method exhibits exceptional mathematical reasoning and OCR capabilities, despite these abilities not being explicitly trained within the VLA. Furthermore, we demonstrate that the VLA possesses strong spatial reasoning skills, enabling it to interpret novel directional instructions involving previously unseen objects. Overall, our method showcases reasoning and comprehension abilities that significantly surpass state-of-the-art imitation learning methods such as OpenVLA, DexVLA, and $\pi_0$. This work represents a substantial advancement toward developing truly generalizable robotic foundation models endowed with robust reasoning capacities. Zhongyi Zhou, Yichen Zhu 0001, Zhibin Tang, Yaxin Peng, Chaomin Shen 0001 |
NeurIPS | 1 |
| 2024 | AffViT: Fast Affine Medical Image Registration with Convolutional Vision Transformer
Chaomin Shen 0001, Zhongyi Zhou |
PRICAI (3) | 3 |
| 2024 | Dynamic Alignment and Fusion of Multimodal Physiological Patterns for Stress RecognitionabstractStress has been identified as one of major causes of health issues. To detect the stress levels with higher accuracy, fusion of multimodal physiological signals is a promising technique. However, there is an asynchrony between physiological signals observed from different perspectives. Exploring the temporal alignment relationship between modalities is helpful to improve the quality of multimodal fusion. This paper proposes an end-to-end multimodal stress detection model based on Bidirectional Cross- and Self-modal Attention (BCSA) mechanism. Specifically, we first construct different feature extractors based on the characteristics of Blood Volume Pulse (BVP) and Electrodermal Activity (EDA) to complete automated temporal feature extraction. Secondly, cross-modal attention is used to seek the alignment relationship between the two modalities and fully fuse cross-modal information. The self-modal attention is used to attenuate noise and redundant information, highlight important information and obtain salient stress representations. Finally, the stress representations of the two modalities are processed separately, and the mean square error (MSE) is used to narrow the gap between them. Experimental results on the UBFC-Phys dataset and WESAD dataset show that the proposed model can effectively improve the accuracy of stress recognition, and outperforms several state-of-the-art methods. Xiaowei Zhang 0001, Zhongyi Zhou, Qiqi Zhao, Sipo Zhang, Rui Li 0105, Bin Hu 0001 |
IEEE Trans. Affect. Comput. | 3 |
| 2024 | Discriminative Joint Knowledge Transfer With Online Updating Mechanism for EEG-Based Emotion RecognitionabstractDomain adaptation (DA) has aroused a wide concern in electroencephalogram (EEG)-based cross-subject emotion recognition tasks. However, many existing DA algorithms focus more on transferability rather than discriminability. In addition, these algorithms typically rely on iterative optimization with pseudo-labels to attain the optimal model. In this study, a novel method with an online updating mechanism named discriminative joint knowledge transfer (DJKT) is proposed. A precise calculation of discriminative information for different emotional states within and across subjects is achieved by leveraging a small number of labeled target-domain samples. Furthermore, to accommodate the time-varying EEG, we extend the passive-aggressive (PA) algorithm to enable online adaptation of the emotion recognition model, thereby enhancing its suitability for real-world scenarios. Extensive experiments conducted on the SJTU emotion EEG dataset (SEED) and SEED-IV demonstrate the effectiveness of our approach. First, comprehensive incorporation of the discriminative information improves the performance of transfer learning significantly. In comparison with several state-of-the-art methods, DJKT exhibits significantly improved emotion recognition performance in both single-source to single-target (STS) and multisource to single-target (MTS) scenarios. Second, the online adjustment strategy effectively addresses the time-varying characteristics of EEG signals, leading to a more robust and stable model. Xiaowei Zhang 0001, Zhongyi Zhou, Qiqi Zhao, Kechen Hou, Sipo Zhang, Yanmeng Cui |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | Emotion Recognition From Multimodal Physiological Signals via Discriminative Correlation Fusion With a Temporal Alignment MechanismabstractModeling correlations between multimodal physiological signals [e.g., canonical correlation analysis (CCA)] for emotion recognition has attracted much attention. However, existing studies rarely consider the neural nature of emotional responses within physiological signals. Furthermore, during fusion space construction, the CCA method maximizes only the correlations between different modalities and neglects the discriminative information of different emotional states. Most importantly, temporal mismatches between different neural activities are often ignored; therefore, the theoretical assumptions that multimodal data should be aligned in time and space before fusion are not fulfilled. To address these issues, we propose a discriminative correlation fusion method coupled with a temporal alignment mechanism for multimodal physiological signals. We first use neural signal analysis techniques to construct neural representations of the central nervous system (CNS) and autonomic nervous system (ANS). respectively. Then, emotion class labels are introduced in CCA to obtain more discriminative fusion representations from multimodal neural responses, and the temporal alignment between the CNS and ANS is jointly optimized with a fusion procedure that applies the Bayesian algorithm. The experimental results demonstrate that our method significantly improves the emotion recognition performance. Additionally, we show that this fusion method can model the underlying mechanisms in human nervous systems during emotional responses, and our results are consistent with prior findings. This study may guide a new approach for exploring human cognitive function based on physiological signals at different time scales and promote the development of computational intelligence and harmonious human-computer interactions. Kechen Hou, Xiaowei Zhang 0001, Qiqi Zhao, Wenjie Yuan 0001, Zhongyi Zhou, Sipo Zhang, Chen Li 0051, Jian Shen 0004, Bin Hu 0001 |
IEEE Trans. Cybern. | 6 |
| 2023 | SoundTraveller: Exploring Abstraction and Entanglement in Timbre Creation Interfaces for SynthesizersabstractTimbre exploration and creation are key tasks in electronic music composition. Modern synthesizers can produce thousands of unique timbres, but this complexity hinders musicians’ ability to explore these timbres effectively. We contribute SoundTraveller, an interactive timbre exploration system aimed at fostering electronic musicians’ creative processes. SoundTraveller allows the user to explore the timbral space using two modes: evolutionary and morphing, with which they can generate hundreds of unique timbres without the need to edit individual parameters. Our user study confirmed that SoundTraveller supported participants’ exploration, decreased their cognitive load, and increased their perceived creativity. Through analysis of our interview study, we contribute design considerations for timbre exploration systems, and for exploring large aesthetic parameter spaces more generally. Finally, we discuss how systems like SoundTraveller can fit into the existing workflows of electronic music composers, and how shared agency with creative technologies can impact the creative process. Zefan Sramek, Arissa J. Sato, Zhongyi Zhou, Simo Hosio, Koji Yatani |
Conference on Designing Interactive Systems | 3 |
| 2022 | Bringing Rolling Shutter Images Alive with Dual Reversed Distortion
Zhihang Zhong, Mingdeng Cao, Xiao Sun 0001, Zhirong Wu, Zhongyi Zhou, Yinqiang Zheng, Stephen Lin 0001, Imari Sato |
ECCV (7) | 5 |
| 2022 | Gesture-aware Interactive Machine Teaching with In-situ Object AnnotationsabstractInteractive Machine Teaching (IMT) systems allow non-experts to easily create Machine Learning (ML) models. However, existing vision-based IMT systems either ignore annotations on the objects of interest or require users to annotate in a post-hoc manner. Without the annotations on objects, the model may misinterpret the objects using unrelated features. Post-hoc annotations cause additional workload, which diminishes the usability of the overall model building process. In this paper, we develop LookHere, which integrates in-situ object annotations into vision-based IMT. LookHere exploits users’ deictic gestures to segment the objects of interest in real time. This segmentation information can be additionally used for training. To achieve the reliable performance of this object segmentation, we utilize our custom dataset called HuTics, including 2040 front-facing images of deictic gestures toward various objects by 170 people. The quantitative results of our user study showed that participants were 16.3 times faster in creating a model with our system compared to a standard IMT system with a post-hoc annotation process while demonstrating comparable accuracies. Additionally, models created by our system showed a significant accuracy improvement (ΔmIoU = 0.466) in segmenting the objects of interest compared to those without annotations. Zhongyi Zhou, Koji Yatani |
UIST | 1 |
| 2018 | An Image-Based Approach for Defect Detection on Decorative Sheets
Boyu Zhou, Zhongyi Zhou, Xinyi Le |
ICONIP (4) | 3 |
| 2015 | Profiling and Understanding Virtualization Overhead in CloudabstractVirtualization is a key technology for cloud data centers to implement infrastructure as a service (IaaS) and to provide flexible and cost-effective resource sharing. It introduces an additional layer of abstraction that produces resource utilization overhead. Disregarding this overhead may cause serious reduction of the monitoring accuracy of the cloud providers and may cause degradation of the VM performance. However, there is no previous work that comprehensively investigates the virtualization overhead. In this paper, we comprehensively measure and study the relationship between the resource utilizations of virtual machines (VMs) and the resource utilizations of the device driver domain, hypervisor and the physical machine (PM) with diverse workloads and scenarios in the Xen virtualization environment. We examine data from the real-world virtualized deployment to characterize VM workloads and assess their impact on the resource utilizations in the system. We show that the impact of virtualization overhead depends on the workloads, and that virtualization overhead is an important factor to consider in cloud resource provisioning. Based on the measurements, we build a regression model to estimate the resource utilization overhead of the PM resulting from providing virtualized resource to the VMs and from managing multiple VMs. Finally, our trace-driven real-world experimental results show the high accuracy of our model in predicting PM resource consumptions in the cloud datacenter, and the importance of considering the virtualization overhead in cloud resource provisioning. Liuhua Chen, Shilkumar Patel, Haiying Shen, Zhongyi Zhou |
ICPP | 4 |