VLDB 2026 Research / reviewers in the wild / expert
Zhicheng Guo
dblp:233/7924
· DBLP profile ↗
16ranked-venue papers
1as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cosine Similarity is Almost All You Need (for Prototypical-Part Models)abstractPrototypical-part networks are a popular interpretable alternative to black-box deep learning models for computer vision because of their faithful, prototype-based self-explanations. However, in practice, they have proven difficult to train because they are highly sensitive to hyperparameter tuning and difficult to comprehend because they contain a large number of prototypes. We show that replacing ℓ2distance with an angular prototype similarity in the original ProtoPNet greatly improves robustness to hyperparameter selection and is sufficient to produce accuracy and sparsity competitive with state-of-the-art on many backbones and datasets. We also show cosine similarity leads to superior accuracy for five different ProtoPNet architectures (ProtoPNet, TesNet, Deformable ProtoPNet, ProtoTree, and ST-ProtoPNet). Finally, we demonstrate ProtoPNet with cosine similarity produces better semantics than ℓ2: prototypes from cosine models score better on prototype quality metrics and are perceived as more similar 3:2 in a user study.1 Luke Moffett, Frank Willard, Maximillian Machado, Emmanuel Mokel, Jon Donnelly, Zhicheng Guo, Adam Costarino, Julia Yang, Giyoung Kim, Alina Barnett, Cynthia Rudin |
WACV | 6 |
| 2026 | Multi-granularity transformer contrastive learning and feature reconstruction for prediction of disease-related miRNAs
Ping Xuan, Zhicheng Guo, Tiangang Zhang |
BMC Bioinform. | 2 |
| 2026 | Synchronous detection of strawberry fruit and stem in natural scene using scale edge fusion network and semi-supervised learning
Qi Wu 0017, Junjie Wan, Zhicheng Guo, Shizhuang Weng, Wenshen Jia |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | Rashomon Sets for Prototypical-Part Networks: Editing Interpretable Models in Real-TimeabstractInterpretability is critical for machine learning models in high-stakes settings because it allows users to verify the model’s reasoning. In computer vision, prototypical part models (ProtoPNets) have become the dominant model type to meet this need. Users can easily identify flaws in ProtoPNets, but fixing problems in a ProtoPNet requires slow, difficult retraining that is not guaranteed to resolve the issue. This problem is called the "interaction bottleneck." We solve the interaction bottleneck for ProtoPNets by simultaneously finding many equally good ProtoPNets (i.e., a draw from a "Rashomon set"). We show that our framework – called Proto-RSet – quickly produces many accurate, diverse ProtoPNets, allowing users to correct problems in real time while maintaining performance guarantees with respect to the training set. We demonstrate the utility of this method in two settings: 1) removing synthetic bias introduced to a bird-identification model and 2) debugging a skin cancer identification model. This tool empowers nonmachine-learning experts, such as clinicians or domain experts, to quickly refine and correct machine learning models without repeated retraining by machine learning experts. Jon Donnelly, Zhicheng Guo, Alina Barnett, Hayden McTavish, Cynthia Rudin |
CVPR | 2 |
| 2025 | "What is Different Between These Datasets?" A Framework for Explaining Data Distribution ShiftsabstractThe performance of machine learning models relies heavily on the quality of input data, yet real-world applications often face significant data-related challenges. A common issue arises when curating training data or deploying models: two datasets from the same domain may exhibit differing distributions. While many techniques exist for detecting such distribution shifts, there is a lack of comprehensive methods to explain these differences in a human-understandable way beyond opaque quantitative metrics. To bridge this gap, we propose a versatile framework of interpretable methods for comparing datasets. Using a variety of case studies, we demonstrate the effectiveness of our approach across diverse data modalities—including tabular data, text data, images, time-series signals – in both low and high-dimensional settings. These methods complement existing techniques by providing actionable and interpretable insights to better understand and address distribution shifts. Varun Babbar, Zhicheng Guo, Cynthia Rudin |
J. Mach. Learn. Res. | 2 |
| 2024 | EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language ModelsabstractVision-language models (VLMs) have recently shown promising results in traditional downstream tasks. Evaluation studies have emerged to assess their abilities, with the majority focusing on the third-person perspective, and only a few addressing specific tasks from the first-person per-spective. However, the capability of VLMs to “think” from a first-person perspective, a crucial attribute for advancing autonomous agents and robotics, remains largely unexplored. To bridge this research gap, we introduce EgoThink, a novel visual question-answering benchmark that encompasses six core capabilities with twelve detailed dimensions. The benchmark is constructed using selected clips from ego-centric videos, with manually annotated question-answer pairs containing first-person information. To comprehensively assess VLMs, we evaluate twenty-one popular VLMs on EgoThink. Moreover, given the open-ended format of the answers, we use GPT-4 as the automatic judge to compute single-answer grading. Experimental results indicate that although GPT-4V leads in numerous dimensions, all evaluated VLMs still possess considerable potential for improvement in first-person perspective tasks. Meanwhile, enlarging the number of trainable parameters has the most significant impact on model performance on EgoThink. In conclusion, EgoThink serves as a valuable addition to existing evaluation benchmarks for VLMs, providing an indispensable resource for future research in the realm of embodied artificial intelligence and robotics. Sijie Cheng, Zhicheng Guo, Kechen Fang, Peng Li 0030, Huaping Liu 0001, Yang Liu 0005 |
CVPR | 2 |
| 2024 | Iterative Translation Refinement with Large Language ModelsabstractWe propose iteratively prompting a large language model to self-correct a translation, with inspiration from their strong language capability as well as a human-like translation approach. Interestingly, multi-turn querying reduces the output’s string-based metric scores, but neural metrics suggest comparable or improved quality after two or more iterations. Human evaluations indicate better fluency and naturalness compared to initial translations and even human references, all while maintaining quality. Ablation studies underscore the importance of anchoring the refinement to the source and a reasonable seed translation for quality considerations. We also discuss the challenges in evaluation and relation to human performance and translationese. Pinzhen Chen, Zhicheng Guo, Barry Haddow, Kenneth Heafield |
EAMT (1) | 2 |
| 2024 | Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?abstractMultilingual large language models are designed, claimed, and expected to cater to speakers of varied languages.We hypothesise that the current practices of fine-tuning and evaluating these models may not perfectly align with this objective owing to a heavy reliance on translation, which cannot cover languagespecific knowledge but can introduce translation defects.It remains unknown whether the nature of the instruction data has an impact on the model output; conversely, it is questionable whether translated test sets can capture such nuances.Due to the often coupled practices of using translated data in both stages, such imperfections could have been overlooked.This work investigates these issues using controlled native or translated data during the instruction tuning and evaluation stages.We show that native or generation benchmarks reveal a notable difference between native and translated instruction data especially when model performance is high, whereas other types of test sets cannot.The comparison between round-trip and single-pass translations reflects the importance of knowledge from language-native resources.Finally, we demonstrate that regularization is beneficial to bridging this gap on structured but not generative tasks. 1 Pinzhen Chen, Simon Yu, Zhicheng Guo, Barry Haddow |
EMNLP | 3 |
| 2024 | Position: Towards Unified Alignment Between Agents, Humans, and EnvironmentabstractThe rapid progress of foundation models has led to the prosperity of autonomous agents, which leverage the universal capabilities of foundation models to conduct reasoning, decision-making, and environmental interaction. However, the efficacy of agents remains limited when operating in intricate, realistic environments. In this work, we introduce the principles of Unified Alignment for Agents (UA$^2$), which advocate for the simultaneous alignment of agents with human intentions, environmental dynamics, and self-constraints such as the limitation of monetary budgets. From the perspective of UA$^2$, we review the current agent research and highlight the neglected factors in existing agent benchmarks and method candidates. We also conduct proof-of-concept studies by introducing realistic features to WebShop, including user profiles demonstrating intentions, personalized reranking reflecting complex environmental dynamics, and runtime cost statistics as self-constraints. We then follow the principles of UA$^2$ to propose an initial design of our agent and benchmark its performance with several candidate baselines in the retrofitted WebShop. The extensive experimental results further prove the importance of the principles of UA$^2$. Our research sheds light on the next steps of autonomous agent research with improved general problem-solving abilities. Zonghan Yang, Kaiming Liu, Fangzhou Xiong, Yile Wang 0001, Zeyuan Yang 0002, Zhenhe Zhang, Fuwen Luo, Zhicheng Guo, Peng Li 0030, Yang Liu 0005 |
ICML | 12 |
| 2024 | Learning From Alarms: A Robust Learning Approach for Accurate Photoplethysmography-Based Atrial Fibrillation Detection Using Eight Million Samples Labeled With Imprecise Arrhythmia AlarmsabstractAtrial fibrillation (AF) is a common cardiac arrhythmia with serious health consequences if not detected and treated early. Detecting AF using wearable devices with photoplethysmography (PPG) sensors and deep neural networks has demonstrated some success using proprietary algorithms in commercial solutions. However, to improve continuous AF detection in ambulatory settings towards a population-wide screening use case, we face several challenges, one of which is the lack of large-scale labeled training data. To address this challenge, we propose to leverage AF alarms from bedside patient monitors to label concurrent PPG signals, resulting in the largest PPG-AF dataset so far (8.5 M 30-second records from 24,100 patients) and demonstrating a practical approach to build large labeled PPG datasets. Furthermore, we recognize that the AF labels thus obtained contain errors because of false AF alarms generated from imperfect built-in algorithms from bedside monitors. Dealing with label noise with unknown distribution characteristics in this case requires advanced algorithms. We, therefore, introduce and open-source a novel loss design, the cluster membership consistency (CMC) loss, to mitigate label errors. By comparing CMC with state-of-the-art methods selected from a noisy label competition, we demonstrate its superiority in handling label noise in PPG data, resilience to poor-quality signals, and computational efficiency. Zhicheng Guo, Cynthia Rudin, Amit J. Shah, Duc H. Do, Randall J. Lee, Gari D. Clifford, Fadi B. Nahab, Xiao Hu 0002 |
IEEE J. Biomed. Health Informatics | 2 |
| 2024 | Multiple information perception-based attention in YOLO for underwater object detection
Huibing Wang, Tianxiang Cui, Zhicheng Guo, Xianping Fu |
Vis. Comput. | 4 |
| 2024 | Publisher Correction: Multiple information perception-based attention in YOLO for underwater object detection
Huibing Wang, Tianxiang Cui, Zhicheng Guo, Xianping Fu |
Vis. Comput. | 4 |
| 2023 | A Spatial Hierarchical Reasoning Network for Remote Sensing Visual Question AnsweringabstractFor visual question answering on remote sensing (RSVQA), current methods scarcely consider geospatial objects typically with large-scale differences and positional sensitive properties. Besides, modeling and reasoning the relationships between entities have rarely been explored, which leads to one-sided and inaccurate answer predictions. In this article, a novel method called spatial hierarchical reasoning network (SHRNet) is proposed, which endows a remote sensing (RS) visual question answering (VQA) system with enhanced visual–spatial reasoning capability. Specifically, a hash-based spatial multiscale visual representation module is first designed to encode multiscale visual features embedded with spatial positional information. Then, spatial hierarchical reasoning is conducted to learn the high-order inner group object relations across multiple scales under the guidance of linguistic cues. Finally, a visual-question (VQ) interaction module is employed to learn an effective image–text joint embedding for the final answer predicting. Experimental results on three public RS VQA datasets confirm the effectiveness and superiority of our model SHRNet. Zixiao Zhang, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Fang Liu 0001, Yuxuan Li 0004, Zhicheng Guo |
IEEE Trans. Geosci. Remote. Sens. | 8 |
| 2023 | A Universal Quaternion Hypergraph Network for Multimodal Video Question AnsweringabstractFusion and interaction of multimodal features are essential for video question answering. Structural information composed of the relationships between different objects in videos is very complex, which restricts understanding and reasoning. In this paper, we propose a quaternion hypergraph network (QHGN) for multimodal video question answering, to simultaneously involve multimodal features and structural information. Since quaternion operations are suitable for multimodal interactions, four components of the quaternion vectors are applied to represent the multimodal features. Furthermore, we construct a hypergraph based on the visual objects detected in the video. Most importantly, the quaternion hypergraph convolution operator is theoretically derived to realize multimodal and relational reasoning. Question and candidate answers are embedded in quaternion space, and a Q&A reasoning module is creatively designed for selecting the answer accurately. Moreover, the unified framework can be extended to other video-text tasks with different quaternion decoders. Experimental evaluations on the TVQA dataset and DramaQA dataset show that our method achieves state-of-the-art performance. Zhicheng Guo, Jiaxuan Zhao, Licheng Jiao, Xu Liu 0006, Fang Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2022 | A Hybrid Solid State Transformer (HSST) based on Two-Stage Medium Voltage SSTabstractThe concept of the Hybrid Solid State Transformer (HSST) has been introduced in earlier works – an amalgam of a large power low frequency transformer and fractional power solid state transformer. This offers an optimum compromise between cost and controllability, and thanks to its fractional approach, high efficiency. This paper presents an Input Parallel Output Series HSST, using a medium voltage two-stage isolated SST. The predicted peak efficiency of the SST is 96.8% and of the HSST is 98.5%. The HSST unit has been tested to verify the concept and operability. Sanjay Rajendran, Zhicheng Guo, Alex Q. Huang |
IECON | 2 |
| 2017 | A high step-up modular DC/DC converter for photovoltaic generation integrated into DC gridsabstractA high step-up-ratio DC/DC converter for photovoltaic energy integrated into high voltage DC grids is proposed in this paper. The converter is mainly composed of modular DC/DC conversion modules which are configured as input-parallel-output-series structure. The proposed module is consisted of two power stages, with a boost circuit as the first stage and a narrow-switching-frequency-variation LLC circuit as the second stage. The interleaved technique of the boost stage effectively reduces the input and output ripples without adding extra components. Full-range soft-switching property is achieved to the LLC stage to minimize switching losses. Theoretical analysis is carried out for the system voltage gain and design principle of the LLC resonant tank. The simulation and experimental results of a 3kW prototype circuit is presented to verify the theoretical analysis and principles for the proposed module. Guoen Cao, Zhicheng Guo |
IECON | 2 |