EDBT 2026 Demo / reviewers in the wild / expert
Jun Xie 0003
dblp:33/3881-3
· DBLP profile ↗
11ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-6152-3943ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Deep Graph Learning for Incomplete Multi-View Clustering with Masked Graph Reconstruction LossabstractThe prevalence of real-world multi-view data makes incomplete multi-view clustering (IMVC) a crucial research. The rapid development of Graph Neural Networks (GNNs) has established them as one of the mainstream approaches for multi-view clustering. Despite significant progress in GNNs-based IMVC, some challenges remain: (1) Most methods rely on the K-Nearest Neighbors (KNN) algorithm to construct static graphs from raw data, which introduces noise and diminishes the robustness of the graph topology. (2) Existing methods typically utilize the Mean Squared Error (MSE) loss between the reconstructed graph and the sparse adjacency graph directly as the graph reconstruction loss, leading to substantial gradient noise during optimization. To address these issues, we propose a novel Dynamic Deep Graph Learning for Incomplete Multi-View Clustering with Masked Graph Reconstruction Loss (DGIMVCM). Firstly, we construct a missing-robust global graph from the raw data. A graph convolutional embedding layer is then designed to extract primary features and refined dynamic view-specific graph structures, leveraging the global graph for imputation of missing views. This process is complemented by graph structure contrastive learning, which identifies consistency among view-specific graph structures. Secondly, a graph self-attention encoder is introduced to extract high-level representations based on the imputed primary features and view-specific graphs, and is optimized with a masked graph reconstruction loss to mitigate gradient noise during optimization. Finally, a clustering module is constructed and optimized through a pseudo-label self-supervised training mechanism. Extensive experiments on multiple datasets validate the effectiveness and superiority of DGIMVCM. Jun Xie 0003, Xingchen Chen, Hongzhu Yi, Kaixin Xu, Yuanxiang Wang, Tianyu Zong, Jiahuan Chen, Guoqing Chao, Feng Chen 0044, Zhepeng Wang 0002, Jungang Xu |
AAAI | 2 |
| 2025 | PurifyGen: A Risk-Discrimination and Semantic-Purification Model for Safe Text-to-Image GenerationabstractRecent advances in diffusion models have notably enhanced text-to-image (T2I) generation quality, but they also raise the risk of generating unsafe content. Traditional safety methods like text blacklisting or harmful content classification have significant drawbacks: they can be easily circumvented or require extensive datasets and extra training. To overcome these challenges, we introduce PurifyGen, a novel, training-free approach for safe T2I generation that retains the model's original weights. PurifyGen introduces a dual-stage strategy for prompt purification. First, we evaluate the safety of each token in a prompt by computing its complementary semantic distance, which measures the semantic proximity between the prompt tokens and concept embeddings from predefined toxic and clean lists. This enables fine-grained prompt classification without explicit keyword matching or retraining. Tokens closer to toxic concepts are flagged as risky. Second, for risky prompts, we apply a dual-space transformation: we project toxic-aligned embeddings into the null space of the toxic concept matrix, effectively removing harmful semantic components, and simultaneously align them into the range space of clean concepts. This dual alignment purifies risky prompts by both subtracting unsafe semantics and reinforcing safe ones, while retaining the original intent and coherence. We further define a token-wise strategy to selectively replace only risky token embeddings, ensuring minimal disruption to safe content. PurifyGen offers a plug-and-play solution with theoretical grounding and strong generalization to unseen prompts and models. Extensive testing shows that PurifyGen surpasses current methods in reducing unsafe content across five datasets and competes well with training-dependent approaches. Zongsheng Cao, Yangfan He, Jun Xie 0003, Zhepeng Wang 0002, Feng Chen 0044 |
ACM Multimedia | 4 |
| 2025 | CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language ModelsabstractLarge Vision-Language Models (LVLMs) have achieved impressive progress in multi-modal understanding and generation. However, they still tend to produce hallucinated content that is inconsistent with the visual input, which limits their reliability in real-world applications. We propose CoFi-Dec, a training-free decoding framework that mitigates hallucinations by integrating generative self-feedback with coarse-to-fine visual conditioning. Inspired by the human visual process from global scene perception to detailed inspection, CoFi-Dec first generates two intermediate textual responses conditioned on coarse- and fine-grained views of the original image. These responses are then transformed into synthetic images using a text-to-image model, forming multi-level visual hypotheses that enrich grounding cues. To unify the predictions from these multiple visual conditions, we introduce a Wasserstein-based fusion mechanism that aligns their predictive distributions into a geometrically consistent decoding trajectory. This principled fusion reconciles high-level semantic consistency with fine-grained visual grounding, leading to more robust and faithful outputs. Extensive experiments on six hallucination-focused benchmarks show that CoFi-Dec substantially reduces both entity-level and semantic-level hallucinations, outperforming existing decoding strategies. The framework is model-agnostic, requires no additional training, and can be seamlessly applied to a wide range of LVLMs. Zongsheng Cao, Yangfan He, Jun Xie 0003, Zhepeng Wang 0002, Feng Chen 0044 |
ACM Multimedia | 4 |
| 2025 | TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and UnderstandingabstractLarge Video Language Models (LVLMs) have rapidly emerged as the focus of multimedia AI research. Nonetheless, when confronted with lengthy videos, these models struggle: their temporal windows are narrow, and they fail to notice fine-grained semantic shifts that unfold over extended durations. Moreover, mainstream text-based retrieval pipelines, which rely chiefly on surface-level lexical overlap, ignore the rich temporal interdependence among visual, audio, and subtitle channels. To mitigate these limitations, we propose TV-RAG, a training-free architecture that couples temporal alignment with entropy-guided semantics to improve long-video reasoning. The framework contributes two main mechanisms: (i) a time-decay retrieval module that injects explicit temporal offsets into the similarity computation, thereby ranking text queries according to their true multimedia context; and (ii) an entropy-weighted key-frame sampler that selects evenly spaced, information-dense frames, reducing redundancy while preserving representativeness. By weaving these temporal and semantic signals together, TV-RAG realises a dual-level reasoning routine that can be grafted onto any LVLM without re-training or fine-tuning. The resulting system offers a lightweight, budget-friendly upgrade path and consistently surpasses most leading baselines across established long-video benchmarks such as Video-MME, MLVU, and LongVideoBench, confirming the effectiveness of our model. Zongsheng Cao, Yangfan He, Jun Xie 0003, Feng Chen 0044, Zhepeng Wang 0002 |
ACM Multimedia | 4 |
| 2025 | Personality Prediction via Multimodal Fusion with Sentiment Analysis EnhancementabstractAccurately identifying personality traits is of profound significance for gaining in-depth insights into human behavior, facilitating efficient human-computer interaction, and developing personalized intelligent systems. However, existing studies often treat personality trait prediction and emotion recognition as relatively independent tasks, neglecting the inherent correlation between them. This report proposes a multimodal fusion prediction framework through the research topic ''On the Interaction between Personality and Emotion in Human Behavior and Social Interaction''. The core goal of this framework is to explore and verify the positive gain of emotional analysis on the accuracy of personality prediction. We extracted visual and audio features based on large-scale data pre-training and fine-tuning, aggregated video features at the character level to enhance the stability of personality prediction, and then integrated the video-level emotion prediction branch. By jointly optimizing the losses of personality prediction and emotion prediction, the generalization performance of the model is improved. Experimental results show that the multi-task learning method integrating emotional information can improve the prediction performance of personality traits to a certain extent, and achieved the first place in the MER-PR validation set of the MER2025 Challenge. This provides empirical evidence for us to explore the complex interaction between emotion and personality. Xuerui Cheng, Feng Chen 0044, Jun Xie 0003, Kanokphan Lertniphonphan, Zhepeng Wang 0002 |
ACM Multimedia | 3 |
| 2024 | STADet: Streaming Timing-Aware Video Lane DetectionabstractLane detection is a fundamental task in autonomous driving, which lies in the real-time detection of lanes of streaming video during driving. We address the lack of temporal flow understanding of existing video lane detectors, propose a streaming video lane detection training framework, and focus on building a series of inter-frame temporal information conduction structures. Specifically, we propose the Deformable Spatio-Temporal Attention (DSTA) module, which accurately captures the instantaneous changing features and position shifts between frames and incorporates key information under different spatio-temporal conditions. Also, to maintain long-time memory at a very low computational cost, we design instance caches that suggest possible lanes for the current frame and resist short-time lane disappearance based on historical memory. We experimented with the inclusion of background category prediction, which is able to simply filter low-confidence false predictions of lanes, while also conveying a more holistic and uniform relationship between lanes and background to the model. These methods allow our model to achieve a significant lead in the video lane detection dataset VIL-100, reaching an accuracy of 94.9 at a speed of 39 FPS. Kaijie He, Jun Xie 0003, Xinguang Dai, Kenglun Chang, Feng Chen 0044, Zhepeng Wang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | CaT: Cyclic-Accumulation Transformer for Lane DetectionabstractLane detection is a special task in autonomous driving. Its most prominent inherent feature is to learn the imagination of severely occluded objects. Traditional CNN-based networks learning the imagination tend to perform poorly. In this work, we propose a novel architecture, called Cycle_accumulation-Transformer (CaT), which is the first structure to handle the lane detection by fusing CNN and Transformer. In particular, Cycle_accumulation structure and Transformer structure complement each other, and they adopt the four-direction cyclic accumulation process of “up to down”, “down to up”, “left to right” and “right to left” in the convolutional mode and the self-attention mechanism of “QKV” to fuse global information respectively. Our method is based on pixel-level semantic segmentation with high detection accuracy while meeting real-time requirements. Moreover, our proposed method achieves state-of-the-art results on the Tusimple and also achieves competitive results on the CULane. Dezhen Qi, Jun Xie 0003, Guoyu Yang, Ye Qiu, Yuer Lu, Xiaoming Jiang, Jianwei Shuai |
IJCNN | 2 |
| 2023 | URFormer: Unified Representation LiDAR-Camera 3D Object Detection with Transformer
Jun Xie 0003, Zhepeng Wang 0002, Kuihe Yang, Ziying Song |
PRCV (3) | 2 |
| 2023 | A Data-Related Patch Proposal for Semantic Segmentation of Aerial ImagesabstractLarge-size images cannot be directly put into GPU for training and need to be cropped to patches due to GPU memory limitation. The commonly used cropping methods before are random cropping and sequential cropping, which are crude and fatally inefficient. Firstly, categories of datasets are often imbalanced, and just simple cropping misses an excellent opportunity to make the data distribution balanced. Secondly, the training needs to crop a large number of patches to cover all patterns, which greatly increases the training time. This problem is of great practical hazards but is often overlooked by previous works. The optimal solution is to generate valuable patches. Valuable patches refer to the value to network training, i.e., the value of this patch for the convergence of the network, and the improvement of the accuracy. To this end, we propose a data-related patch proposal strategy to sample high valuable patches. The core idea is to score each patch according to the accuracy of each category, so as to perform balanced sampling. Compared with random cropping or sequential cropping, our method can improve the segmentation accuracy and accelerate the training vastly. Moreover, our method also shows great advantages over the loss-based balanced approaches. Experiments on Deepglobe and Potsdam show the excellent effect of our method. Lianlei Shan, Guiqin Zhao, Jun Xie 0003, Peirui Cheng, Xiaobin Li 0006, Zhepeng Wang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | VoxelNextFusion: A Simple, Unified, and Effective Voxel Fusion Framework for Multimodal 3-D Object DetectionabstractLiDAR-camera fusion can enhance the performance of 3D object detection by utilizing complementary information between depth-aware LiDAR points and semantically rich images. Existing voxel-based methods face significant challenges when fusing sparse voxel features with dense image features in a one-to-one manner, resulting in the loss of the advantages of images, including semantic and continuity information, leading to sub-optimal detection performance, especially at long distances. In this paper, we present VoxelNextFusion, a multi-modal 3D object detection framework specifically designed for voxel-based methods, which effectively bridges the gap between sparse point clouds and dense images. In particular, we propose a voxel-based image pipeline that involves projecting point clouds onto images to obtain both pixel- and patch-level features. These features are then fused using a self-attention to obtain a combined representation. Moreover, to address the issue of background features present in patches, we propose a feature importance module that effectively distinguishes between foreground and background features, thus minimizing the impact of the background features. Extensive experiments were conducted on the widely used KITTI and nuScenes 3D object detection benchmarks. Notably, our VoxelNextFusion achieved around +3.20% in [email protected] improvement for car detection in hard level compared to the Voxel R-CNN baseline on the KITTI test dataset. Ziying Song, Jun Xie 0003, Caiyan Jia, Shaoqing Xu, Zhepeng Wang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Improving Deep Image Matting Via Local Smoothness AssumptionabstractNatural image matting is a fundamental and challenging computer vision task. Conventionally, the problem is formulated as an underconstrained problem. Since the problem is ill-posed, further assumptions on the data distribution are required to make the problem well-posed. For classical matting methods, a commonly adopted assumption is the local smoothness assumption on foreground and background colors. However, the use of such assumptions was not systematically considered for deep learning based matting methods. In this work, we consider two local smoothness assumptions which can help improving deep image matting models. Based on the local smoothness assumptions, we propose three techniques, i.e., training set refinement, color augmentation and backpropagating refinement, which can improve the performance of the deep image matting model significantly. We conduct experiments to examine the effectiveness of the proposed algorithm. The experimental results show that the proposed method has favorable performance compared with existing matting methods. Jun Xie 0003, Jiacheng Han, Dezhen Qi |
ICME | 2 |