Zipeng Guo

dblp:246/3093 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 UM-Text: A Unified Multimodal Model for Image Understanding and Visual Text Editing
abstract
With the rapid advancement of image generation, visual text editing using natural language instructions has received increasing attention. The main challenge of this task is to fully understand the instruction and reference image, and thus generate visual text that is style-consistent with the image. Previous methods often involve complex steps of specifying the text content and attributes, such as font size, color, and layout, without considering the stylistic consistency with the reference image. To address this, we propose UM-Text, a unified multimodal model for context understanding and visual text editing by natural language instructions. Specifically, we introduce a Visual Language Model (VLM) to process the instruction and reference image, so that the text content and layout can be elaborately designed according to the context information. To generate an accurate and harmonious visual text image, we further propose the UM Encoder to combine the embeddings of various condition information, where the combination is automatically configured by VLM according to the input instruction. During training, we propose a regional consistency loss to offer more effective supervision for glyph generation on both latent and RGB space, and design a tailored three-stage training strategy to further enhance model performance. In addition, we contribute the UM-DATA-200K, a large-scale visual text image dataset on diverse scenes for model training. Extensive qualitative and quantitative results on multiple public benchmarks demonstrate that our method achieves state-of-the-art performance.
Lichen Ma, Xiaolong Fu, Gaojing Zhou, Zipeng Guo, Yichun Liu, Junshi Huang
AAAI4
2026 Synergistic audio-textual cues: A cross-modal framework for weakly-supervised temporal action localization
Linkai Liu 0002, Yuchen Zhou 0002, Zipeng Guo, Chao Gou
Pattern Recognit.3
2026 Learning From Individual to Collective: A Unified Framework for Driver-Aware Attention Prediction
Yuchen Zhou 0002, Zipeng Guo, Linkai Liu 0002, Yueyao Lin, Chao Gou
IEEE Trans. Intell. Transp. Syst.2
2025 Where, What, Why: Towards Explainable Driver Attention Prediction
abstract
Modeling task-driven attention in driving is a fundamental challenge for both autonomous vehicles and cognitive science. Existing methods primarily predict where drivers look by generating spatial heatmaps, but fail to capture the cognitive motivations behind attention allocation in specific contexts, which limits deeper understanding of attention mechanisms. To bridge this gap, we introduce Explainable Driver Attention Prediction, a novel task paradigm that jointly predicts spatial attention regions (where), parses attended semantics (what), and provides cognitive reasoning for attention allocation (why). To support this, we present W3DA, the first large-scale explainable driver attention dataset. It enriches existing benchmarks with detailed semantic and causal annotations across diverse driving scenarios, including normal conditions, safety-critical situations, and traffic accidents. We further propose LLada, a Large Language model-driven framework for driver attention prediction, which unifies pixel modeling, semantic parsing, and cognitive reasoning within an end-to-end architecture. Extensive experiments demonstrate the effectiveness of LLada, exhibiting robust generalization across datasets and driving conditions. This work serves as a key step toward a deeper understanding of driver attention mechanisms, with significant implications for autonomous driving, intelligent driver training, and human-computer interaction.
Yuchen Zhou 0002, Jiayu Tang, Xiaoyan Xiao, Yueyao Lin, Linkai Liu 0002, Zipeng Guo, Hao Fei 0001, Xiaobo Xia, Chao Gou
ICCV6
2025 Boosting Road Event Detection with Adaptive Multi-Modal Models
abstract
Despite significant advancements in road event detection (RED), existing approaches encounter critical limitations. These include reliance on single-modal inputs and joint optimization of detection and classification tasks, often leading to conflicting objectives and suboptimal performance. Moreover, their heavy dependence on large-scale annotated datasets restricts generalization in data-scarce scenarios. To address these challenges, we propose AdaRED, a novel framework that decouples agent detection from event classification, thereby mitigating optimization conflicts and enhancing task-specific performance. To comprehensively understand road events, AdaRED uses diverse input modalities, including fine-grained local features, global contextual information, and spatial layout embeddings. Additionally, we introduce the Cross-modal Scene Adaptation Module (CSAM), which integrates lightweight adapters into the multimodal model. This design enables efficient extraction of spatiotemporal features and the integration of visual priors, thereby improving generalization and robustness in challenging scenarios. Extensive experiments on the ROAD-R dataset validate the effectiveness of AdaRED, achieving state-of-the-art performance and addressing the limitations of existing methods. The code can be found on our project page: https://liulinkai.github.io/AdaRED/.
Linkai Liu 0002, Xiaoyan Xiao, Yijian Yang, Yuchen Zhou 0002, Zipeng Guo, Chao Gou
ICME5
2025 Behavior-Aware Knowledge-Embedded Model for Driver Attention Prediction
abstract
Accurately predicting driver attention is crucial for enhancing advanced driving assistance systems and autonomous vehicles, attracting increasing research interest. Most existing approaches, rooted in general, task-free saliency detection, adopt data-driven paradigms to correlate bottom-up environmental situations with attention distributions. However, they often overlook the complex top-down task-driven aspects of driver attention that are fundamental for the safe navigation of driving tasks, leading to limitations in handling real-world scenarios. In this paper, we take an initial step to explore and introduce BKnet, a Behavior-aware Knowledge-embedded model that innovatively integrates driving behaviors and empirical knowledge. Specifically, inspired by the human long-term cognitive process, we introduce a novel knowledge memory mechanism. It dynamically associates varied traffic scenarios with consistent driving behaviors, fostering the generation of robust behavior-aware empirical knowledge representations. To this end, BKnet facilitates a nuanced and comprehensive simulation of drivers’ attention mechanisms, driven synergistically by both top-down and bottom-up processes. Additionally, we further contribute to the field by collecting a novel Behavior-Aware Driver Attention (BADA) dataset. To the best of our knowledge, BADA is the first attention dataset explicitly incorporated into real-world driving behavior tasks from multiple drivers. Lastly, comprehensive experiments underscore BKnet’s superiority over existing state-of-the-art approaches and validate the effectiveness and necessity of integrating behavior-aware knowledge into driver attention prediction.
Yuchen Zhou 0002, Chao Gou, Zipeng Guo, Yihua Cheng, Hyung Jin Chang
IEEE Trans. Circuits Syst. Video Technol.3
2024 DrivingGen: Efficient Safety-Critical Driving Video Generation with Latent Diffusion Models
abstract
With the increasing popularity of autonomous driving, a demand for high-quality safety-critical driving video data is urgently required. However, such large-scale data is hard to obtain due to expensive and risky collection costs. To alleviate the problem, we propose DrivingGen, an efficient approach built upon the T2I diffusion model for safety-critical driving video generation. Our model employs the "Spatio-Temporal-then-Temporal" paradigm, learning motion priors from a local to global perspective. Firstly, we design an innovative Segment Flow Module to achieve local spatio-temporal modeling by capturing the distinctive dynamic features of different video segments. Secondly, a lightweight Directional Consistency Attention is proposed to further enhance temporal consistency from a global perspective. Additionally, we propose an efficient Temporal Shift Adapter to expand the T2I U-Net into the temporal dimension. Empowered with these modules, DrivingGen outperforms the state-of-the-arts in driving video generation for safety-critical scenarios, as determined by both quality and efficiency measures. Video examples are available in our project page: https://gzp6688.github.io/DrivingGen
Zipeng Guo, Yuchen Zhou 0002, Chao Gou
ICME1
2024 Cascaded learning with transformer for simultaneous eye landmark, eye state and gaze estimation
Chao Gou, Yuezhao Yu, Zipeng Guo, Chen Xiong
Pattern Recognit.3
2023 Controllable Diffusion Models for Safety-Critical Driving Scenario Generation
abstract
Safety-critical driving scenarios are essential to the development and validation of autonomous driving algorithms. Currently, most of the data is acquired in naturalistic scenarios, resulting in a sparsity of the safety-critical cases. Consequently, synthetic scenario generation based on deep models becomes crucial to validate the risk and reduce the cost. However, previous works either fail to generate realistic scenarios with high quality or can hardly follow instructions to generate desired scenarios. In this paper, we propose a controllable diffusion model that operates on a naturalistic driving scenario to generate safety-critical cases with high fidelity and controllability. In particular, our proposed method encompasses a pre-trained text-to-image diffusion model and a bounding box to incorporate the conditions of category and position of generated objects, respectively. To enhance generation quality and mitigate boundary artifacts, we introduce a mask-aware adapter to better integrate the generated objects into the driving scenarios. Moreover, we propose a transformer encoder and region-guided cross attention to fuse the additional coordinate inputs into the diffusion model, while encouraging more interaction between different generated objects. Comparative analysis demonstrates that our method outperforms exciting work by offering more realistic, diverse and controllable synthetic scenarios and allowing for multiple objects generation with complex spatial relationship.
Zipeng Guo, Yuezhao Yu, Chao Gou
ICTAI1
2019 Multiple Partitions Aligned Clustering
abstract
Multi-view clustering is an important yet challenging task due to the difficulty of integrating the information from multiple representations. Most existing multi-view clustering methods explore the heterogeneous information in the space where the data points lie. Such common practice may cause significant information loss because of unavoidable noise or inconsistency among views. Since different views admit the same cluster structure, the natural space should be all partitions. Orthogonal to existing techniques, in this paper, we propose to leverage the multi-view information by fusing partitions. Specifically, we align each partition to form a consensus cluster indicator matrix through a distinct rotation matrix. Moreover, a weight is assigned for each view to account for the clustering capacity differences of views. Finally, the basic partitions, weights, and consensus clustering are jointly learned in a unified framework. We demonstrate the effectiveness of our approach on several real datasets, where significant improvement is found over other state-of-the-art multi-view clustering methods.
Zhao Kang 0001, Zipeng Guo, Shudong Huang, Siying Wang 0002, Wenyu Chen 0001, Yuanzhang Su, Zenglin Xu
IJCAI2