EDBT 2026 Demo / reviewers in the wild / expert
Yuchen Zhou 0002
dblp:39/10084-2
· DBLP profile ↗
16ranked-venue papers
9as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Logic Unseen: Revealing the Logical Blindspots of Vision-Language ModelsabstractVision-Language Models (VLMs), exemplified by CLIP, have emerged as foundational for multimodal intelligence. However, their capacity for logical understanding remains significantly underexplored, resulting in critical **''logical blindspots''** that limit their reliability in practical applications. To systematically diagnose this, we introduce **LogicBench**, a comprehensive benchmark with over 50,000 vision-language pairs across 9 logical categories and 4 diverse scenarios: images, videos, anomaly detection, and medical diagnostics. Our evaluation reveals that existing VLMs, even the state-of-the-art ones, fall at over 40 accuracy points below human performance, particularly in challenging tasks like Causality and Conditionality, highlighting their reliance on surface semantics over critical logical structures. To bridge this gap, we propose **LogicCLIP**, a novel training framework designed to boost VLMs' logical sensitivity through advancements in both data generation and optimization objectives. LogicCLIP utilizes logic-aware data generation and a contrastive learning strategy that combines coarse-grained alignment, a fine-grained multiple-choice objective, and a novel logical structure-aware objective. Extensive experiments demonstrate LogicCLIP's substantial improvements in logical comprehension across all LogicBench domains, significantly outperforming baselines. Moreover, LogicCLIP retains, and often surpasses, competitive performance on general vision-language benchmarks, demonstrating that the enhanced logical understanding does not come at the expense of general alignment. We believe LogicBench and LogicCLIP will be important resources for advancing VLM logical capabilities. Yuchen Zhou 0002, Jiayu Tang, Shuo Yang 0006, Xiaoyan Xiao, Yuqin Dai, Chao Gou, Xiaobo Xia, Tat-Seng Chua |
AAAI | 1 |
| 2026 | Synergistic audio-textual cues: A cross-modal framework for weakly-supervised temporal action localization
Linkai Liu 0002, Yuchen Zhou 0002, Zipeng Guo, Chao Gou |
Pattern Recognit. | 2 |
| 2026 | Learning From Individual to Collective: A Unified Framework for Driver-Aware Attention Prediction
Yuchen Zhou 0002, Zipeng Guo, Linkai Liu 0002, Yueyao Lin, Chao Gou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Where, What, Why: Towards Explainable Driver Attention PredictionabstractModeling task-driven attention in driving is a fundamental challenge for both autonomous vehicles and cognitive science. Existing methods primarily predict where drivers look by generating spatial heatmaps, but fail to capture the cognitive motivations behind attention allocation in specific contexts, which limits deeper understanding of attention mechanisms. To bridge this gap, we introduce Explainable Driver Attention Prediction, a novel task paradigm that jointly predicts spatial attention regions (where), parses attended semantics (what), and provides cognitive reasoning for attention allocation (why). To support this, we present W3DA, the first large-scale explainable driver attention dataset. It enriches existing benchmarks with detailed semantic and causal annotations across diverse driving scenarios, including normal conditions, safety-critical situations, and traffic accidents. We further propose LLada, a Large Language model-driven framework for driver attention prediction, which unifies pixel modeling, semantic parsing, and cognitive reasoning within an end-to-end architecture. Extensive experiments demonstrate the effectiveness of LLada, exhibiting robust generalization across datasets and driving conditions. This work serves as a key step toward a deeper understanding of driver attention mechanisms, with significant implications for autonomous driving, intelligent driver training, and human-computer interaction. Yuchen Zhou 0002, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin, Linkai Liu 0002, Zipeng Guo, Hao Fei 0001, Xiaobo Xia, Chao Gou |
ICCV | 1 |
| 2025 | Boosting Road Event Detection with Adaptive Multi-Modal ModelsabstractDespite significant advancements in road event detection (RED), existing approaches encounter critical limitations. These include reliance on single-modal inputs and joint optimization of detection and classification tasks, often leading to conflicting objectives and suboptimal performance. Moreover, their heavy dependence on large-scale annotated datasets restricts generalization in data-scarce scenarios. To address these challenges, we propose AdaRED, a novel framework that decouples agent detection from event classification, thereby mitigating optimization conflicts and enhancing task-specific performance. To comprehensively understand road events, AdaRED uses diverse input modalities, including fine-grained local features, global contextual information, and spatial layout embeddings. Additionally, we introduce the Cross-modal Scene Adaptation Module (CSAM), which integrates lightweight adapters into the multimodal model. This design enables efficient extraction of spatiotemporal features and the integration of visual priors, thereby improving generalization and robustness in challenging scenarios. Extensive experiments on the ROAD-R dataset validate the effectiveness of AdaRED, achieving state-of-the-art performance and addressing the limitations of existing methods. The code can be found on our project page: https://liulinkai.github.io/AdaRED/. Linkai Liu 0002, Xiaoyan Xiao, Yijian Yang, Yuchen Zhou 0002, Zipeng Guo, Chao Gou |
ICME | 4 |
| 2025 | CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG DecodingabstractUnderstanding and decoding human brain activity from electroencephalography (EEG) signals is a fundamental problem in neuroscience and artificial intelligence, with applications ranging from cognition and emotion recognition to clinical diagnosis and brain–computer interfaces. While recent EEG foundation models have made progress in generalized brain decoding by leveraging unified architectures and large-scale pretraining, they inherit a scale-agnostic dense modeling paradigm from NLP and vision. This design overlooks an intrinsic property of neural activity—cross-scale spatiotemporal structure. Different EEG task patterns span a broad range of temporal and spatial scales, from brief neural activations to slow-varying rhythms, and from localized cortical activations to large-scale distributed interactions. Ignoring this diversity may lead to suboptimal representations and weakened generalization ability. To address these limitations, we propose CSBrain, a Cross-scale Spatiotemporal Brain foundation model for generalized EEG decoding. CSBrain introduces two key components: (i) Cross-scale Spatiotemporal Tokenization (CST), which aggregates multi-scale features within localized temporal windows and anatomical brain regions into compact scale-aware token representations; and (ii) Structured Sparse Attention (SSA), which models cross-window and cross-region dependencies for diverse decoding tasks, further enriching scale diversities while eliminating the spurious dependencies. CST and SSA are alternately stacked to progressively integrate cross-scale spatiotemporal dependencies. Extensive experiments across 11 representative EEG tasks and 16 datasets demonstrate that CSBrain consistently outperforms both task-specific models and strong foundation baselines. These results establish cross-scale modeling as a key inductive bias for generalized EEG decoding and highlight CSBrain as a robust backbone for future brain–AI research. Yuchen Zhou 0002, Zichen Ren, Zhouheng Yao, Weiheng Lu, Kunyu Peng, Qihao Zheng, Chunfeng Song, Wanli Ouyang, Chao Gou |
NeurIPS | 1 |
| 2025 | Behavior-Aware Knowledge-Embedded Model for Driver Attention PredictionabstractAccurately predicting driver attention is crucial for enhancing advanced driving assistance systems and autonomous vehicles, attracting increasing research interest. Most existing approaches, rooted in general, task-free saliency detection, adopt data-driven paradigms to correlate bottom-up environmental situations with attention distributions. However, they often overlook the complex top-down task-driven aspects of driver attention that are fundamental for the safe navigation of driving tasks, leading to limitations in handling real-world scenarios. In this paper, we take an initial step to explore and introduce BKnet, a Behavior-aware Knowledge-embedded model that innovatively integrates driving behaviors and empirical knowledge. Specifically, inspired by the human long-term cognitive process, we introduce a novel knowledge memory mechanism. It dynamically associates varied traffic scenarios with consistent driving behaviors, fostering the generation of robust behavior-aware empirical knowledge representations. To this end, BKnet facilitates a nuanced and comprehensive simulation of drivers’ attention mechanisms, driven synergistically by both top-down and bottom-up processes. Additionally, we further contribute to the field by collecting a novel Behavior-Aware Driver Attention (BADA) dataset. To the best of our knowledge, BADA is the first attention dataset explicitly incorporated into real-world driving behavior tasks from multiple drivers. Lastly, comprehensive experiments underscore BKnet’s superiority over existing state-of-the-art approaches and validate the effectiveness and necessity of integrating behavior-aware knowledge into driver attention prediction. Yuchen Zhou 0002, Chao Gou, Zipeng Guo, Yihua Cheng, Hyung Jin Chang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Learning from Observer Gaze: Zero-Shot Attention Prediction Oriented by Human-Object Interaction RecognitionabstractMost existing attention prediction research focuses on salient instances like humans and objects. However, the more complex interaction-oriented attention, arising from the comprehension of interactions between instances by human observers, remains largely unexplored. This is equally crucial for advancing human-machine interaction and human-centered artificial intelligence. To bridge this gap, we first collect a novel gaze fixation dataset named IG, comprising 530,000 fixation points across 740 diverse in-teraction categories, capturing visual attention during human observers' cognitive processes of interactions. Subsequently, we introduce the zero-shot interaction-oriented attention prediction task (ZeroIA), which challenges models to predict visual cues for interactions not encountered during training. Thirdly, we present the Interactive Attention model (IA), designed to emulate human observers' cognitive processes to tackle the ZeroIA problem. Extensive experiments demonstrate that the proposed IA outperforms other state-of-the-art approaches in both ZeroIA and fully supervised settings. Lastly, we endeavor to apply interaction-oriented attention to the interaction recognition task itself. Further experimental results demonstrate the promising potential to enhance the performance and inter-pretability of existing state-of-the-art HOI models by incorporating real human attention data from IG and attention labels generated by IA. Yuchen Zhou 0002, Linkai Liu 0002, Chao Gou |
CVPR | 1 |
| 2024 | Driver Scanpath Prediction Based On Inverse Reinforcement LearningabstractModeling driver attention allocation by scanpath prediction plays a crucial role in advancing autonomous driving capabilities and enhancing accident anticipation. Existing studies primarily predict human scanpath for visual search, visual question answering, and free-viewing. Few studies have focused on predicting scanpath in driving scenarios. To address these limitations, we propose a novel inverse reinforcement learning-based approach through adversarial learning to effectively anticipate human-like scanpaths within different driving tasks. Particularly, we introduce a Transformer-based architecture to construct the generator and discriminator models, while integrating top-down, bottom-up, and historical information for dynamic state updated through State-Encode-with-Attention (SEA). Inspired by the human visual system, SEA adopts a fovea-like movement strategy. Experimental results on the benchmark dataset of BDD-X-diverse validate the effectiveness of our proposed method. Yuchen Zhou 0002, Chao Gou |
ICASSP | 2 |
| 2024 | Hierarchical Home Action Understanding with Implicit and Explicit Prior KnowledgeabstractExisting investigations on action understanding have made noteworthy advancements by treating activities as holistic events occurring in videos. However, these investigations have limited ability to comprehensively extract and represent human experiential knowledge, which hampers various practical applications, such as robotics and human-computer interaction. We argue that human actions can be better understood as hierarchical compositions of multiple interactive objects and atomic actions with spatio-temporal relations. To this end, we propose a hierarchical understanding framework for home actions, which decomposes a single holistic action into multiple quintuples of. Within this framework, we introduce a two-stage network architecture that leverages multiple prior knowledge in both implicit and explicit ways to facilitate mutual learning within quintuples. In particular, we fully exploit statistical knowledge to enhance the inference of data-driven visual model. Experiments validate the effectiveness of our proposed method. This is also the winning solution for HOMAGE Competition @ ActivityNet Challenge in CVPR 2022 and our brief oral representation is available at https://youtu.be/KK3SPK6iueE?si=hrFZzSABNyrrL6jF&t=1727. Yuchen Zhou 0002, Guang Tan, Chao Gou |
ICASSP | 1 |
| 2024 | DrivingGen: Efficient Safety-Critical Driving Video Generation with Latent Diffusion ModelsabstractWith the increasing popularity of autonomous driving, a demand for high-quality safety-critical driving video data is urgently required. However, such large-scale data is hard to obtain due to expensive and risky collection costs. To alleviate the problem, we propose DrivingGen, an efficient approach built upon the T2I diffusion model for safety-critical driving video generation. Our model employs the "Spatio-Temporal-then-Temporal" paradigm, learning motion priors from a local to global perspective. Firstly, we design an innovative Segment Flow Module to achieve local spatio-temporal modeling by capturing the distinctive dynamic features of different video segments. Secondly, a lightweight Directional Consistency Attention is proposed to further enhance temporal consistency from a global perspective. Additionally, we propose an efficient Temporal Shift Adapter to expand the T2I U-Net into the temporal dimension. Empowered with these modules, DrivingGen outperforms the state-of-the-arts in driving video generation for safety-critical scenarios, as determined by both quality and efficiency measures. Video examples are available in our project page: https://gzp6688.github.io/DrivingGen Zipeng Guo, Yuchen Zhou 0002, Chao Gou |
ICME | 2 |
| 2024 | Dynamic Attention-Enhanced Spatio-Temporal Network for Pedestrian Collision Risk Assessment
Benfei Wang, Xinxin Liu 0015, Yuchen Zhou 0002, Chao Gou |
PRCV (10) | 4 |
| 2024 | Task-Oriented Scanpath Prediction with Spatial-Temporal Information in Driving Scenarios
Yuchen Zhou 0002, Chao Gou |
PRCV (10) | 2 |
| 2023 | Learning from Easy to Hard Pairs: Multi-step Reasoning Network for Human-Object Interaction DetectionabstractHuman-object interaction (HOI) detection aims to interpret the interactions of human-object pairs. Existing methods adopt a one-step reasoning paradigm that simultaneously outputs multi-label results for all HOI pairs without distinguishing difficulties. However, there are significant variations among HOI pairs in the same image, making their performance degrade in challenging situations. In this paper, we argue that the model should prioritize hard samples after inferring easy ones, and hard samples can benefit from easy ones. To this end, we propose a novel Multi-step Reasoning Network that progressively learns from easy to hard samples. In particular, an Easy-to-Hard Learning Block is introduced to enhance the representation of hard HOI pairs by prior associations. Additionally, we propose a Multi-step Reasoning Probability Transfer mechanism to enhance multi-label interaction classifications, which leverages cognitive associations and semantic dependencies. Extensive experiments demonstrate that our method outperforms other state-of-the-art on two challenging benchmark datasets. Yuchen Zhou 0002, Guang Tan, Mengtang Li, Chao Gou |
ACM Multimedia | 1 |
| 2023 | PIT: Progressive Interaction Transformer for Pedestrian Crossing Intention PredictionabstractFor autonomous driving, one of the major challenges is to predict pedestrian crossing intention in ego-view. Pedestrian intention depends not only on their intrinsic goals but also on the stimulation of surrounding traffic elements. Considering the influence of other traffic elements on pedestrian intention, recent work introduced more traffic element information into the model to successfully improve performance. However, it is still difficult to effectively capture and fully exploit the potential dynamic spatio-temporal interactions among the target pedestrian and its surrounding traffic elements for accurate reasoning. In this work, inspired by neuroscience that human drivers tend to make continuous sensory-motor driving decisions by progressive visual stimulation, we propose a model termed Progressive Interaction Transformer (PIT) for pedestrian crossing intention prediction. Local pedestrian, global environment, and ego-vehicle motion are considered simultaneously in the proposed PIT. In particular, the temporal fusion block and self-attention mechanism are introduced to jointly and progressively model the dynamic spatio-temporal interactions among the three parties, allowing it to capture richer information and make prediction in a similar way to human drivers. Experimental results demonstrate that PIT achieves higher performance compared with other state-of-the-arts and preserves real-time inference. Yuchen Zhou 0002, Guang Tan, Rui Zhong 0001, Yaokun Li, Chao Gou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | Driver attention prediction based on convolution and transformers
Chao Gou, Yuchen Zhou 0002 |
J. Supercomput. | 2 |