VLDB 2026 Research / reviewers in the wild / expert
Han Jiang 0012
dblp:80/10002-12
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0009-0007-2014-7906ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action LocalizationabstractCurrent Zero-Shot Temporal Action Localization (ZSTAL) methods, whether training-based or training-free ones, still predominantly rely on a single, unified query to localize an entire action. This unified representation is fundamentally ill-suited for complex real-world activities, as it fails to capture their internal compositional structure and adapt to dynamic, multi-stage variations across videos. To address this, we regard ZSTAL as a compositional reasoning task and introduce CASCADE, a Context-Aware Staged Action DEcomposition framework. Inspired by the human cognitive process of perceiving context, decomposing events, and reconstructing instances, CASCADE follows a training-free pipeline. It first perceives the video's context by leveraging a Multimodal Large Language Model (MLLM) to both filter out irrelevant actions and then generate a rich, video-specific caption for each action present in the video. An LLM then decomposes this caption into multiple, temporally ordered stages, which serve as fine-grained queries to guide the MLLM in estimating frame-level confidence scores. Recognizing that this decomposition can fragment a single action, a novel hierarchical merging logic then reconstructs complete instances by intelligently fusing these preliminary temporal segments based on their semantic progression and coherence. Extensive experiments and ablation studies on THUMOS14 and ActivityNet-1.3 show that CASCADE not only sets a new state-of-the-art among training-free methods but, most notably, significantly outperforms all prior training-based approaches on ActivityNet-1.3. Haoyu Tang 0002, Tianyuan Liang, Han Jiang 0012, Qinghai Zheng, Yupeng Hu 0003 |
AAAI | 3 |
| 2026 | Visual Self-paced Iterative Learning for Unsupervised Temporal Action LocalizationabstractRecently, Temporal Action Localization (TAL) has garnered significant interest in information retrieval community. However, existing supervised/weakly supervised methods are heavily dependent on extensive labeled temporal boundaries and action categories, which is labor-intensive and time-consuming. Although some unsupervised methods have utilized the “iteratively clustering and localization” paradigm for TAL, they still suffer from two pivotal impediments: (1) unsatisfactory video clustering confidence, and (2) unreliable video pseudolabels for model training. To address these limitations, we present a novel self-paced iterative learning model to enhance clustering and localization training simultaneously, thereby facilitating more effective unsupervised TAL. Concretely, we improve the clustering confidence through exploring the contextual feature-robust visual information. Thereafter, we design two (constant- and variable-speed) incremental instance learning strategies for easy-to-hard model training, thus ensuring the reliability of these video pseudolabels and further improving overall localization performance. Extensive experiments on two public datasets demonstrate the superiority of our model over several state-of-the-art competitors. Yupeng Hu 0003, Han Jiang 0012, Hao Liu 0072, Kun Wang 0039, Haoyu Tang 0002, Liqiang Nie |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video LocalizationabstractZero-shot Natural Language Video Localization (NLVL) aims to automatically generate moments and corresponding pseudo queries from raw videos for the training of the localization model without any manual annotations. Existing approaches typically produce pseudo queries as simple words, which overlook the complexity of queries in real-world scenarios. Considering the powerful text modeling capabilities of large language models (LLMs), leveraging LLMs to generate complete queries that are closer to human descriptions is a potential solution. However, directly integrating LLMs into existing approaches introduces several issues, including insensitivity, isolation, and lack of regulation, which prevent the full exploitation of LLMs to enhance zero-shot NLVL performance. To address these issues, we propose BTDP, an innovative framework for Boundary-aware Temporal Dynamic Pseudo-supervision pairs generation. Our method contains two crucial operations: 1) Boundary Segmentation that identifies both visual boundaries and semantic boundaries to generate the atomic segments and activity descriptions, tackling the issue of insensitivity. 2) Context Aggregation that employs the LLMs with a self-evaluation process to aggregate and summarize global video information for optimized pseudo moment-query pairs, tackling the issue of isolation and lack of regulation. Comprehensive experimental results on the Charades-STA and ActivityNet Captions datasets demonstrate the effectiveness of our BTDP method. Xiongwen Deng, Haoyu Tang 0002, Han Jiang 0012, Qinghai Zheng, Jihua Zhu |
AAAI | 3 |
| 2025 | FACE: A Dual-Template and Adaptive Curriculum Framework for Unsupervised Text-Based Person SearchabstractText-Based Person Search, which aims to retrieve target pedestrian images using natural language descriptions, has garnered significant attention in multimedia research due to its potential in suspect retrieval and missing person identification. While supervised and weakly supervised methods rely on costly annotated training data, unsupervised TBPS eliminates the need for textual descriptions or identity annotations, presenting a more practical paradigm. Current unsupervised TBPS approaches face two primary challenges: 1) Predefined attribute templates for caption generation limit linguistic diversity and real-world adaptability, and 2) Threshold-based sample selection using pre-trained vision-language models (VLMs) introduces noisy pairs due to inadequate pedestrian-specific representation. To address these limitations, we propose FACE, a unified framework featuring Dual-template Caption Generation (DCG) and Adaptive Curriculum Training (ACT). The DCG module generates high-quality captions through complementary flexible-style (natural language) and fixed-style (attribute-enumerated) templates, enhanced by LLM-based noise filtering. The ACT framework progressively refines training through a self-improving loop: initial high-confidence sample selection using VLMs bootstraps the model, while evolving feature representations enable dynamic incorporation of harder samples through curriculum learning. This dual strategy achieves mutual reinforcement between caption quality and model discriminability. Extensive experiments on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets under unsupervised settings demonstrate that our framework achieves the state-of-the-art performance. Xiaoxuan Mu, Haoyu Tang 0002, Han Jiang 0012, Tianyuan Liang, Qinghai Zheng, Jihua Zhu |
ACM Multimedia | 3 |
| 2024 | Revisiting Unsupervised Temporal Action Localization: The Primacy of High-Quality Actionness and PseudolabelsabstractRecently, temporal action localization (TAL) methods, especially the weakly-supervised and unsupervised ones, have become a hot research topic. Existing unsupervised methods follow an iterative ''clustering and training'' strategy with diverse model designs during training stage, while they often overlook maintaining consistency between these stages, which is crucial: more accurate clustering results can reduce the noises of pseudolabels and thus enhance model training, while more robust training can in turn enrich clustering feature representation. We identify two critical challenges in unsupervised scenarios: 1. What features should the model generate for clustering? 2. Which pseudolabeled instances from clustering should be chosen for model training? After extensive explorations, we proposed a novel yet simple framework called Consistency-Oriented Progressive high actionness Learning to address these issues. For feature generation, our framework adopts a High Actionness snippet Selection (HAS) module to generate more discriminative global video features for clustering from the enhanced actionness features obtained from a designed Inner-Outer Consistency Network (IOCNet). For pseudolabel selection, we introduces a Progressive Learning With Representative Instances (PLRI) strategy to identify the most reliable and informative instances within each cluster for model training. These three modules, HAS, IOCNet, and PLRI, synergistically improve consistency in model training and clustering performance. Extensive experiments on THUMOS'14 and ActivityNet v1.2 datasets under both unsupervised and weakly-supervised settings demonstrate that our framework achieves the state-of-the-art results. Han Jiang 0012, Haoyu Tang 0002, Ming Yan 0008, Ji Zhang 0011, Yupeng Hu 0003, Jihua Zhu, Liqiang Nie |
ACM Multimedia | 1 |
| 2024 | FGCL: Fine-Grained Contrastive Learning For Mandarin Stuttering Event DetectionabstractThis paper presents the T031 team’s approach to the StutteringSpeech Challenge in SLT2024. Mandarin Stuttering Event Detection (MSED) aims to detect instances of stuttering events in Mandarin speech. We propose a detailed acoustic analysis method to improve the accuracy of stutter detection by capturing subtle nuances that previous Stuttering Event Detection (SED) techniques have overlooked. To this end, we introduce the Fine-Grained Contrastive Learning (FGCL) framework for MSED. Specifically, we model the frame-level probabilities of stuttering events and introduce a mining algorithm to identify both easy and confusing frames. Then, we propose a stutter contrast loss to enhance the distinction between stuttered and fluent speech frames, thereby improving the discriminative capability of stuttered feature embeddings. Extensive evaluations on English and Mandarin datasets demonstrate the effectiveness of FGCL, achieving a significant increase of over 5.0% in F1 score on Mandarin data1.1FGCL won 3rd place in Mandarin stuttering event detection and automatic speech recognition in SLT2024 Han Jiang 0012, Yiquan Zhou, Hongwu Ding, Jiacheng Xu 0007, Jihua Zhu |
SLT | 1 |