Zhepeng Wang 0002

dblp:242/8456-2 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0002-6088-3517ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 8 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Dynamic Deep Graph Learning for Incomplete Multi-View Clustering with Masked Graph Reconstruction Loss
abstract
The prevalence of real-world multi-view data makes incomplete multi-view clustering (IMVC) a crucial research. The rapid development of Graph Neural Networks (GNNs) has established them as one of the mainstream approaches for multi-view clustering. Despite significant progress in GNNs-based IMVC, some challenges remain: (1) Most methods rely on the K-Nearest Neighbors (KNN) algorithm to construct static graphs from raw data, which introduces noise and diminishes the robustness of the graph topology. (2) Existing methods typically utilize the Mean Squared Error (MSE) loss between the reconstructed graph and the sparse adjacency graph directly as the graph reconstruction loss, leading to substantial gradient noise during optimization. To address these issues, we propose a novel Dynamic Deep Graph Learning for Incomplete Multi-View Clustering with Masked Graph Reconstruction Loss (DGIMVCM). Firstly, we construct a missing-robust global graph from the raw data. A graph convolutional embedding layer is then designed to extract primary features and refined dynamic view-specific graph structures, leveraging the global graph for imputation of missing views. This process is complemented by graph structure contrastive learning, which identifies consistency among view-specific graph structures. Secondly, a graph self-attention encoder is introduced to extract high-level representations based on the imputed primary features and view-specific graphs, and is optimized with a masked graph reconstruction loss to mitigate gradient noise during optimization. Finally, a clustering module is constructed and optimized through a pseudo-label self-supervised training mechanism. Extensive experiments on multiple datasets validate the effectiveness and superiority of DGIMVCM.
Jun Xie 0003, Xingchen Chen, Hongzhu Yi, Kaixin Xu, Yuanxiang Wang, Tianyu Zong, Jiahuan Chen, Guoqing Chao, Feng Chen 0044, Zhepeng Wang 0002, Jungang Xu
AAAI13
2025 Learning Supplementary Information for First-Person Perception Referring Expression Comprehension
Zetao Du, Yan Huang 0008, Liang Wang 0001, Feng Chen 0044, Zhepeng Wang 0002
ICIG (2)6
2025 DVS-Aware Visual Perception for Pose Estimation of Mobile Robots with Neuromorphic Implementation
abstract
The Dynamic Vision Sensor (DVS) is a distinctive visual sensor that exclusively responds to alterations in pixel brightness, enabling the real-time capture of swift and subtle movements with reduced power consumption and data bandwidth requirements. This paper proposes a DVS-aware visual perception method and presents its application for pose estimation of mobile robots. Specifically, a new marker is designed to provide pose reference data that leverages the inherent advantages of DVS more effectively. Moreover, we formulate a pose recognition system incorporating DVS, an algorithm based on Spiking Convolutional Neural Networks (SCNN) and a neuromorphic computing accelerator (Lynxi HS110). Such a formulation can well explore the DVS's advantages, as its event-triggered feature matches the nature of SCNN while the neuromorphic hardware enables efficient, low-power execution, making the system highly suitable for real-time embedded applications. Comparative analysis with traditional ARcode-based pose recognition methods reveals that our innovative approach demonstrates significant advantages in recognition speed and energy efficiency. The whole system is deployed on mobile robots and evaluated in real-world scenarios.
Hanzhong Zhong, Yingjie Jin, Guangbin Li, Zhepeng Wang 0002, Xiang Li 0009
ICRA4
2025 PurifyGen: A Risk-Discrimination and Semantic-Purification Model for Safe Text-to-Image Generation
abstract
Recent advances in diffusion models have notably enhanced text-to-image (T2I) generation quality, but they also raise the risk of generating unsafe content. Traditional safety methods like text blacklisting or harmful content classification have significant drawbacks: they can be easily circumvented or require extensive datasets and extra training. To overcome these challenges, we introduce PurifyGen, a novel, training-free approach for safe T2I generation that retains the model's original weights. PurifyGen introduces a dual-stage strategy for prompt purification. First, we evaluate the safety of each token in a prompt by computing its complementary semantic distance, which measures the semantic proximity between the prompt tokens and concept embeddings from predefined toxic and clean lists. This enables fine-grained prompt classification without explicit keyword matching or retraining. Tokens closer to toxic concepts are flagged as risky. Second, for risky prompts, we apply a dual-space transformation: we project toxic-aligned embeddings into the null space of the toxic concept matrix, effectively removing harmful semantic components, and simultaneously align them into the range space of clean concepts. This dual alignment purifies risky prompts by both subtracting unsafe semantics and reinforcing safe ones, while retaining the original intent and coherence. We further define a token-wise strategy to selectively replace only risky token embeddings, ensuring minimal disruption to safe content. PurifyGen offers a plug-and-play solution with theoretical grounding and strong generalization to unseen prompts and models. Extensive testing shows that PurifyGen surpasses current methods in reducing unsafe content across five datasets and competes well with training-dependent approaches.
Zongsheng Cao, Yangfan He, Jun Xie 0003, Zhepeng Wang 0002, Feng Chen 0044
ACM Multimedia5
2025 CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
abstract
Large Vision-Language Models (LVLMs) have achieved impressive progress in multi-modal understanding and generation. However, they still tend to produce hallucinated content that is inconsistent with the visual input, which limits their reliability in real-world applications. We propose CoFi-Dec, a training-free decoding framework that mitigates hallucinations by integrating generative self-feedback with coarse-to-fine visual conditioning. Inspired by the human visual process from global scene perception to detailed inspection, CoFi-Dec first generates two intermediate textual responses conditioned on coarse- and fine-grained views of the original image. These responses are then transformed into synthetic images using a text-to-image model, forming multi-level visual hypotheses that enrich grounding cues. To unify the predictions from these multiple visual conditions, we introduce a Wasserstein-based fusion mechanism that aligns their predictive distributions into a geometrically consistent decoding trajectory. This principled fusion reconciles high-level semantic consistency with fine-grained visual grounding, leading to more robust and faithful outputs. Extensive experiments on six hallucination-focused benchmarks show that CoFi-Dec substantially reduces both entity-level and semantic-level hallucinations, outperforming existing decoding strategies. The framework is model-agnostic, requires no additional training, and can be seamlessly applied to a wide range of LVLMs.
Zongsheng Cao, Yangfan He, Jun Xie 0003, Zhepeng Wang 0002, Feng Chen 0044
ACM Multimedia5
2025 TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
abstract
Large Video Language Models (LVLMs) have rapidly emerged as the focus of multimedia AI research. Nonetheless, when confronted with lengthy videos, these models struggle: their temporal windows are narrow, and they fail to notice fine-grained semantic shifts that unfold over extended durations. Moreover, mainstream text-based retrieval pipelines, which rely chiefly on surface-level lexical overlap, ignore the rich temporal interdependence among visual, audio, and subtitle channels. To mitigate these limitations, we propose TV-RAG, a training-free architecture that couples temporal alignment with entropy-guided semantics to improve long-video reasoning. The framework contributes two main mechanisms: (i) a time-decay retrieval module that injects explicit temporal offsets into the similarity computation, thereby ranking text queries according to their true multimedia context; and (ii) an entropy-weighted key-frame sampler that selects evenly spaced, information-dense frames, reducing redundancy while preserving representativeness. By weaving these temporal and semantic signals together, TV-RAG realises a dual-level reasoning routine that can be grafted onto any LVLM without re-training or fine-tuning. The resulting system offers a lightweight, budget-friendly upgrade path and consistently surpasses most leading baselines across established long-video benchmarks such as Video-MME, MLVU, and LongVideoBench, confirming the effectiveness of our model.
Zongsheng Cao, Yangfan He, Jun Xie 0003, Feng Chen 0044, Zhepeng Wang 0002
ACM Multimedia6
2025 Personality Prediction via Multimodal Fusion with Sentiment Analysis Enhancement
abstract
Accurately identifying personality traits is of profound significance for gaining in-depth insights into human behavior, facilitating efficient human-computer interaction, and developing personalized intelligent systems. However, existing studies often treat personality trait prediction and emotion recognition as relatively independent tasks, neglecting the inherent correlation between them. This report proposes a multimodal fusion prediction framework through the research topic ''On the Interaction between Personality and Emotion in Human Behavior and Social Interaction''. The core goal of this framework is to explore and verify the positive gain of emotional analysis on the accuracy of personality prediction. We extracted visual and audio features based on large-scale data pre-training and fine-tuning, aggregated video features at the character level to enhance the stability of personality prediction, and then integrated the video-level emotion prediction branch. By jointly optimizing the losses of personality prediction and emotion prediction, the generalization performance of the model is improved. Experimental results show that the multi-task learning method integrating emotional information can improve the prediction performance of personality traits to a certain extent, and achieved the first place in the MER-PR validation set of the MER2025 Challenge. This provides empirical evidence for us to explore the complex interaction between emotion and personality.
Xuerui Cheng, Feng Chen 0044, Jun Xie 0003, Kanokphan Lertniphonphan, Zhepeng Wang 0002
ACM Multimedia6
2024 i-Octree: A Fast, Lightweight, and Dynamic Octree for Proximity Search
abstract
Establishing the correspondences between newly acquired points and historically accumulated data (i.e., the map) through nearest neighbor search is crucial in numerous robotic applications. However, static tree data structures are inadequate to handle large and dynamically growing maps in real-time. To address this issue, we present the i-Octree, a dynamic octree data structure that supports both fast nearest neighbor search and real-time dynamic updates, such as point insertion, deletion, and on-tree down-sampling. The i-Octree is built upon a leaf-based octree and has two key features: a local spatially continuous storing strategy that allows for fast access to points while minimizing memory usage, and local on-tree updates that significantly reduce computation time compared to existing static or dynamic tree structures. The experiments show that the i-Octree outperforms contemporary state-of-the-art approaches by achieving, on average, a 19% reduction in runtime on real-world open datasets.
Zhepeng Wang 0002, Shengjie Wang 0002, Tao Zhang 0006
ICRA3
2024 STADet: Streaming Timing-Aware Video Lane Detection
abstract
Lane detection is a fundamental task in autonomous driving, which lies in the real-time detection of lanes of streaming video during driving. We address the lack of temporal flow understanding of existing video lane detectors, propose a streaming video lane detection training framework, and focus on building a series of inter-frame temporal information conduction structures. Specifically, we propose the Deformable Spatio-Temporal Attention (DSTA) module, which accurately captures the instantaneous changing features and position shifts between frames and incorporates key information under different spatio-temporal conditions. Also, to maintain long-time memory at a very low computational cost, we design instance caches that suggest possible lanes for the current frame and resist short-time lane disappearance based on historical memory. We experimented with the inclusion of background category prediction, which is able to simply filter low-confidence false predictions of lanes, while also conveying a more holistic and uniform relationship between lanes and background to the model. These methods allow our model to achieve a significant lead in the video lane detection dataset VIL-100, reaching an accuracy of 94.9 at a speed of 39 FPS.
Kaijie He, Jun Xie 0003, Xinguang Dai, Kenglun Chang, Feng Chen 0044, Zhepeng Wang 0002
IEEE Trans. Circuits Syst. Video Technol.6
2023 URFormer: Unified Representation LiDAR-Camera 3D Object Detection with Transformer
Jun Xie 0003, Zhepeng Wang 0002, Kuihe Yang, Ziying Song
PRCV (3)4
2023 A Data-Related Patch Proposal for Semantic Segmentation of Aerial Images
abstract
Large-size images cannot be directly put into GPU for training and need to be cropped to patches due to GPU memory limitation. The commonly used cropping methods before are random cropping and sequential cropping, which are crude and fatally inefficient. Firstly, categories of datasets are often imbalanced, and just simple cropping misses an excellent opportunity to make the data distribution balanced. Secondly, the training needs to crop a large number of patches to cover all patterns, which greatly increases the training time. This problem is of great practical hazards but is often overlooked by previous works. The optimal solution is to generate valuable patches. Valuable patches refer to the value to network training, i.e., the value of this patch for the convergence of the network, and the improvement of the accuracy. To this end, we propose a data-related patch proposal strategy to sample high valuable patches. The core idea is to score each patch according to the accuracy of each category, so as to perform balanced sampling. Compared with random cropping or sequential cropping, our method can improve the segmentation accuracy and accelerate the training vastly. Moreover, our method also shows great advantages over the loss-based balanced approaches. Experiments on Deepglobe and Potsdam show the excellent effect of our method.
Lianlei Shan, Guiqin Zhao, Jun Xie 0003, Peirui Cheng, Xiaobin Li 0006, Zhepeng Wang 0002
IEEE Geosci. Remote. Sens. Lett.6
2023 VoxelNextFusion: A Simple, Unified, and Effective Voxel Fusion Framework for Multimodal 3-D Object Detection
abstract
LiDAR-camera fusion can enhance the performance of 3D object detection by utilizing complementary information between depth-aware LiDAR points and semantically rich images. Existing voxel-based methods face significant challenges when fusing sparse voxel features with dense image features in a one-to-one manner, resulting in the loss of the advantages of images, including semantic and continuity information, leading to sub-optimal detection performance, especially at long distances. In this paper, we present VoxelNextFusion, a multi-modal 3D object detection framework specifically designed for voxel-based methods, which effectively bridges the gap between sparse point clouds and dense images. In particular, we propose a voxel-based image pipeline that involves projecting point clouds onto images to obtain both pixel- and patch-level features. These features are then fused using a self-attention to obtain a combined representation. Moreover, to address the issue of background features present in patches, we propose a feature importance module that effectively distinguishes between foreground and background features, thus minimizing the impact of the background features. Extensive experiments were conducted on the widely used KITTI and nuScenes 3D object detection benchmarks. Notably, our VoxelNextFusion achieved around +3.20% in [email protected] improvement for car detection in hard level compared to the Voxel R-CNN baseline on the KITTI test dataset.
Ziying Song, Jun Xie 0003, Caiyan Jia, Shaoqing Xu, Zhepeng Wang 0002
IEEE Trans. Geosci. Remote. Sens.7
2019 A Fast and Accurate Fully Convolutional Network for End-to-End Handwritten Chinese Text Segmentation and Recognition
abstract
Handwritten Chinese Text Recognition (HCTR) is a challenging problem due to its high complexity. Previous methods based on over-segmentation, hidden Markov model (HMM) or long short-term memory recurrent neural network (LSTM-RNN) have achieved great success in recognition results. However, all of them, including over-segmentation based methods, are incompetent in accurate segmentation of single character. To solve this problem, we propose a fast and accurate fully convolutional network for end-to-end segmentation and recognition of handwritten Chinese text. Experiments on CASIA-HWDB datasets and ICDAR 2013 competition dataset show that our method achieves a competitive performance on recognition and produces great character segmentation results. Moreover, our model reaches a real-time speed of 70 fps, which is fast enough for various applications.
Dezhi Peng, Yaqiang Wu, Zhepeng Wang 0002, Mingxiang Cai
ICDAR4
2019 Omnidirectional Scene Text Detection with Sequential-free Box Discretization
abstract
Scene text in the wild is commonly presented with high variant characteristics. Using quadrilateral bounding box to localize the text instance is nearly indispensable for detection methods. However, recent researches reveal that introducing quadrilateral bounding box for scene text detection will bring a label confusion issue which is easily overlooked, and this issue may significantly undermine the detection performance. To address this issue, in this paper, we propose a novel method called Sequential-free Box Discretization (SBD) by discretizing the bounding box into key edges (KE) which can further derive more effective methods to improve detection performance. Experiments showed that the proposed method can outperform state-of-the-art methods in many popular scene text benchmarks, including ICDAR 2015, MLT, and MSRA-TD500. Ablation study also showed that simply integrating the SBD into Mask R-CNN framework, the detection performance can be substantially improved. Furthermore, an experiment on the general object dataset HRSC2016 (multi-oriented ships) showed that our method can outperform recent state-of-the-art methods by a large margin, demonstrating its powerful generalization ability.
Sheng Zhang 0024, Lele Xie, Yaqiang Wu, Zhepeng Wang 0002
IJCAI6