EDBT 2026 Demo / reviewers in the wild / expert
Xianghao Zang
dblp:184/6472
· DBLP profile ↗
13ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0001-8421-7167ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive Evidential Learning for Temporal-Semantic Robustness in Moment RetrievalabstractIn the domain of moment retrieval, accurately identifying temporal segments within videos based on natural language queries remains challenging. Traditional methods often employ pre-trained models that struggle with fine-grained information and deterministic reasoning, leading to difficulties in aligning with complex or ambiguous moments. To overcome these limitations, we explore Deep Evidential Regression (DER) to construct a vanilla Evidential baseline. However, this approach encounters two major issues: the inability to effectively handle modality imbalance and the structural differences in DER's heuristic uncertainty regularizer, which adversely affect uncertainty estimation. This misalignment results in high uncertainty being incorrectly associated with accurate samples rather than challenging ones. Our observations indicate that existing methods lack the adaptability required for complex video scenarios. In response, we propose Debiased Evidential Learning for Moment Retrieval (DEMR), a novel framework that incorporates a Reflective Flipped Fusion (RFF) block for cross-modal alignment and a query reconstruction task to enhance text sensitivity, thereby reducing bias in uncertainty estimation. Additionally, we introduce a Geom-regularizer to refine uncertainty predictions, enabling adaptive alignment with difficult moments and improving retrieval accuracy. Extensive testing on standard datasets and debiased datasets ActivityNet-CD and Charades-CD demonstrates significant enhancements in effectiveness, robustness, and interpretability, positioning our approach as a promising solution for temporal-semantic robustness in moment retrieval. Haojian Huang, Kaijing Ma, Xianghao Zang, Han Fang 0002, Chao Ban, Hao Sun 0038, Mulin Chen, Zhongjiang He |
AAAI | 6 |
| 2025 | ViCo: A Multitask Video-enhanced and Cognition-preserving Modality Alignment Training FrameworkabstractThe rapid development of multimodal large language models (MLLMs) has brought significant breakthroughs to this field. However, current MLLMs typically rely on vision instruction tuning based on large language models (LLMs) to endow them with multimodal capabilities, which may lead to low video utilization and catastrophic forgetting due to the task-specific nature of vision instruction tuning. Take case of multiple-choice Video Question Answering (VideoQA) as an example, requiring only selecting correct answers makes it easier for LLMs to shortcut rather than fully comprehend video content. Moreover, the fixed task format of ’’predicting the correct option" risks catastrophic forgetting, which can cause LLMs to lose their origin cognitive abilities. To address this, we propose ViCo, a training framework that enables LLMs to perform a specific auxiliary tasks during vision instruction tuning, which helps the model leverage video information more thoroughly while mitigating catastrophic forgetting, thereby improving modality alignment and enhancing performance. We evaluate our method on several strong VideoQA benchmarks, where it outperforms all state-of-the-art methods, demonstrating its effectiveness and generality. Additionally, we also provide extensive analysis on the video information utilization and catastrophic forgetting resulting. The code will be made available at https://github.com/dunknsabsw/ViCo. Zhenda Yu, Jiayu Shen, Lanxiang Zhou, Han Fang 0002, Xianghao Zang, Chao Ban, Jingfeng Chen, Zhongjiang He, Hao Sun 0038, Yanmei Kang |
ICASSP | 6 |
| 2024 | ProTA: Probabilistic Token Aggregation for Text-Video RetrievalabstractText-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (ProTA) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, ProTA achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%). Han Fang 0002, Xianghao Zang, Chao Ban, Zerun Feng, Lanxiang Zhou, Zhongjiang He, Hao Sun 0038 |
ICME | 2 |
| 2024 | GOAL: Grounded text-to-image Synthesis with Joint Layout Alignment TuningabstractRecent text-to-image (T2I) synthesis models have demonstrated intriguing abilities to produce high-quality images based on text prompts. However, current models still face Text-Image Misalignment problem (e.g., attribute errors and relation mistakes) for compositional generation. Existing models attempted to condition T2I models on grounding inputs to improve controllability while ignoring the explicit supervision from the layout conditions. To tackle this issue, we propose Grounded jOint lAyout aLignment (GOAL), an effective framework for T2I synthesis. Two novel modules, discriminative semantic alignment (DSAlign) and masked attention alignment (MAAlign), are proposed and incorporated in this framework to improve the text-image alignment. DSAlign leverages discriminative tasks at the region-wise level to ensure low-level semantic alignment. MAAlign provides high-level attention alignment by guiding the model to focus on the target object. We also build a dataset GOAL2K for model fine-tuning, which composes 2000 semantically accurate image-text pairs and their layout annotations. Comprehensive evaluations on T2I-Compbench, NSR-1K, and Drawbench demonstrate the superior generation performance of our method. Especially, there are improvements of 19%, 13%, and 12% in color, shape, and texture metrics for T2I-Compbench. Additionally, Q-Align metrics demonstrate that our method can generate images of higher quality. Han Fang 0002, Zerun Feng, Kaijing Ma, Chao Ban, Xianghao Zang, Lanxiang Zhou, Zhongjiang He, Jingyan Chen, Jiani Hu, Hao Sun 0038 |
ACM Multimedia | 6 |
| 2023 | Mask to Reconstruct: Cooperative Semantics Completion for Video-text RetrievalabstractRecently, masked video modeling has been widely explored and improved the model's understanding ability of visual regions at a local level. However, existing methods usually adopt random masking and follow the same reconstruction paradigm to complete the masked regions, which do not leverage the correlations between cross-modal content. In this paper, we present MAsk for Semantics COmpleTion (MASCOT) based on semantic-based masked modeling. Specifically, after applying attention-based video masking to generate high-informed and low-informed masks, we propose Informed Semantics Completion to recover masked semantics information. The recovery mechanism is achieved by aligning the masked content with the unmasked visual regions and corresponding textual context, which makes the model capture more text-related details at a patch level. Additionally, we shift the emphasis of reconstruction from irrelevant backgrounds to discriminative parts to ignore regions with low-informed masks. Furthermore, we design co-learning to incorporate video cues under different masks and learn more aligned representation. Our MASCOT performs state-of-the-art performance on four text-video retrieval benchmarks, including MSR-VTT, LSMDC, ActivityNet, and DiDeMo. Han Fang 0002, Zhifei Yang 0004, Xianghao Zang, Chao Ban, Zhongjiang He, Hao Sun 0038, Lanxiang Zhou |
ACM Multimedia | 3 |
| 2023 | A Baseline Investigation: Transformer-based Cross-view Baseline for Text-based Person SearchabstractThis paper investigates a baseline approach for text-based person search by using a transformer-based framework. Existing methods usually treat the visual and textual features as independent entities for speeding up the model inference process. However, the attention to the same images should be changed according to different texts. In this paper, we use a commonly employed framework with a fused feature as the baseline, which overcomes the misalignment problem introduced by fixed features. A thorough investigation is conducted in this paper. Moreover, we propose Cross-View Matching (CVM) to provide challenging, positive text-image pairs that enable the model to learn cross-view meta-information. Furthermore, we suggest a novel evaluation process to reduce the inference time and GPU memory demand. The experiments are conducted on CUHK-PEDES, ICFG-PEDES, and RSTPReid benchmarks. Through extensive parameter analysis, the potentials of a transformer-based framework are fully explored. Although the proposed scheme is a simple framework, it achieves significant performance improvements compared with other state-of-the-art methods. Xianghao Zang, Wei Gao 0003, Ge Li 0002, Han Fang 0002, Chao Ban, Zhongjiang He, Hao Sun 0038 |
ACM Multimedia | 1 |
| 2022 | Exploiting robust unsupervised video person re-identificationabstractAbstract Unsupervised video person re‐identification (reID) methods usually depend on global‐level features. Many supervised reID methods employed local‐level features and achieved significant performance improvements. However, applying local‐level features to unsupervised methods may introduce an unstable performance. To improve the performance stability for unsupervised video reID, this paper introduces a general scheme fusing part models and unsupervised learning. In this scheme, the global‐level feature is divided into equal local‐level feature. A local‐aware module is employed to explore the potentials of local‐level feature for unsupervised learning. A global‐aware module is proposed to overcome the disadvantages of local‐level features. Features from these two modules are fused to form a robust feature representation for each input image. This feature representation has the advantages of local‐level feature without suffering from its disadvantages. Comprehensive experiments are conducted on three benchmarks, including PRID2011, iLIDS‐VID, and DukeMTMC‐VideoReID, and the results demonstrate that the proposed approach achieves state‐of‐the‐art performance. Extensive ablation studies demonstrate the effectiveness and robustness of proposed scheme, local‐aware module and global‐aware module. The code and generated features are available at https://github.com/deropty/uPMnet . Xianghao Zang, Ge Li 0002, Wei Gao 0003, Xiujun Shu |
IET Image Process. | 1 |
| 2022 | Large-Scale Spatio-Temporal Person Re-Identification: Algorithms and BenchmarkabstractPerson re-identification (re-ID) in the scenario with large spatial and temporal spans has not been fully explored. This fact partially occurs because existing benchmark datasets were mainly collected with limited spatial and temporal ranges,e.g.,using videos recorded in a few days by cameras in a specific region of the campus. Such limited spatial and temporal ranges make it hard to simulate the difficulties of person re-ID in real scenarios. In this work, we contribute a novel Large-scale Spatio-Temporal (LaST) person re-ID dataset, including 10,862 identities with more than 228k images. Compared with existing datasets, LaST presents more challenging and high-diversity re-ID settings and significantly larger spatial and temporal ranges. For instance, each person can appear in different cities or countries, and in various time slots from day to evening, and in different seasons from spring to winter. To our best knowledge, LaST is a novel person re-ID dataset with the largest spatio-temporal ranges. Based on LaST, we verified its challenge by conducting a comprehensive performance evaluation of 14 re-ID algorithms. We further propose an easy-to-implement baseline that works well in such challenging re-ID settings. We also verified that models pre-trained on LaST can generalize well on existing datasets with short-term and cloth-changing scenarios. We expect LaST to inspire future works toward more realistic and challenging re-ID tasks. More information about the dataset is available athttps://github.com/shuxjweb/last.git. Xiujun Shu, Xiao Wang 0014, Xianghao Zang, Shiliang Zhang, Yuanqi Chen, Ge Li 0002, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Multidirection and Multiscale Pyramid in Transformer for Video-Based Pedestrian RetrievalabstractIn video surveillance, pedestrian retrieval (also called person reidentification) is a critical task. This task aims to retrieve the pedestrian of interest from nonoverlapping cameras. Recently, transformer-based models have achieved significant progress for this task. However, these models still suffer from ignoring fine-grained, part-informed information. This article proposes a multidirection and multiscale Pyramid in Transformer (PiT) to solve this problem. In transformer-based architecture, each pedestrian image is split into many patches. Then, these patches are fed to transformer layers to obtain the feature representation of this image. To explore the fine-grained information, this article proposes to apply vertical division and horizontal division on these patches to generate different-direction human parts. These parts provide more fine-grained information. To fuse multiscale feature representation, this article presents a pyramid structure containing global-level information and many pieces of local-level information from different scales. The feature pyramids of all the pedestrian images from the same video are fused to form the final multidirection and multiscale feature representation. Experimental results on two challenging video-based benchmarks, MARS and iLIDS-VID, show the proposed PiT achieves state-of-the-art performance. Extensive ablation studies demonstrate the superiority of the proposed pyramid structure. Data is available on-line athttps://git.openi.org.cn/zangxh/PiT.git. Xianghao Zang, Ge Li 0002, Wei Gao 0003 |
IEEE Trans. Ind. Informatics | 1 |
| 2021 | Learning to disentangle scenes for person re-identification
Xianghao Zang, Ge Li 0002, Wei Gao 0003, Xiujun Shu |
Image Vis. Comput. | 1 |
| 2021 | Diverse part attentive network for video-based person re-identification
Xiujun Shu, Ge Li 0002, Longhui Wei, Jia-Xing Zhong, Xianghao Zang, Shiliang Zhang, Yaowei Wang 0001, Yongsheng Liang 0001, Qi Tian 0001 |
Pattern Recognit. Lett. | 5 |
| 2017 | Adaptive difference modelling for background subtractionabstractBackground subtraction plays a very important role in video analysis, especially in surveillance systems. While being straightforward, the performance based on frame differencing is unsatisfied due to its sensitiveness to issues such as camera shake and swinging objects. To address its limitations, in this paper we propose a complete adaptive difference modelling framework. First, we introduce two difference discriminators to model the evolution process of pixels. Second, we use Gaussian Mixture Models to adaptively learn the difference threshold to distinguish foreground from background. Third, three heuristics are employed to further improve the model adaptability. Experiments on real-world videos of the Background Models Challenge (BMC) demonstrate that our method performs better on global quality metric (FSD) than other state-of-the-art methods. Xianghao Zang, Ge Li 0002, Jun Yang 0033, Wenmin Wang 0001 |
VCIP | 1 |
| 2016 | A Novel Shadow-Free Feature Extractor for Real-Time Road DetectionabstractRoad detection is one of the most important research areas in driver assistance and automated driving field. However, the performance of existing methods is still unsatisfactory, especially in severe shadow conditions. To overcome those difficulties, first we propose a novel shadow-free feature extractor based on the color distribution of road surface pixels. Then we present a road detection framework based on the extractor, whose performance is more accurate and robust than that of existing extractors. Also, the proposed framework has much low-complexity, which is suitable for usage in practical systems. Zhenqiang Ying, Ge Li 0002, Xianghao Zang, Ronggang Wang, Wenmin Wang 0001 |
ACM Multimedia | 3 |