EDBT 2026 Demo / reviewers in the wild / expert
Kun Wang 0039
dblp:05/1958-39
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0008-4856-8806ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Interference to Stability: Adversarial Reliability Correction for Video Moment Retrieval with Relevance FeedbackabstractVideo Moment Retrieval (VMR) aims to retrieve target video moments that correspond to natural language queries. Most existing methods rely on a positive-only assumption that the queried moment always exists within the video, which limits their reliability in practical scenarios. Departing from this restrictive setting, we study Video Moment Retrieval with Relevance Feedback (VMR-RF), which requires models to both retrieve relevant moments and reject irrelevant queries. This task remains challenging due to the following issues: 1) Intrinsic Semantic Interference caused by visually similar but irrelevant moments, and 2) Propagative Decision Irreversibility induced by unidirectional relevance prediction. In light of these, we introduce AdversaRial Reliability cOrrection netWork (ARROW) for VMR-RF. ARROW adopts an active discriminative strategy through two synergetic components: a Gradient-induced Semantic Adversary (GSA) that probes model vulnerabilities by actively amplifying semantic interference, and an Adversarial Reliability Predictor (ARP) that quantifies prediction stability under such interference to effectively suppress unreliable decisions. Extensive experiments validate the effectiveness of ARROW. Hao Liu 0072, Yupeng Hu 0003, Kun Wang 0039, Ruping Cao, Yutao Yao, Zilu Cai |
SIGIR | 3 |
| 2026 | Redundancy Mitigation: Toward Accurate and Efficient Image-Text RetrievalabstractImage-text retrieval (ITR) is a pivotal task in cross-modal research. However, existing methods often suffer from a fundamental yet overlooked challenge: redundancy. This issue manifests as both semantic redundancy within unimodal representations and relationship redundancy in cross-modal alignments. This not only inflates computational costs but also degrades retrieval accuracy by masking salient features and reinforcing spurious correlations. In this work, we are the first to explicitly analyze and address the ITR problem from a redundancy perspective by proposing the iMage-text rEtrieval rEdundancy miTigation (MEET) framework. MEET employs a cascaded, two-stage process to systematically mitigate both forms of redundancy. First, for Semantic Redundancy Mitigation, it repurposes deep hashing and quantization as synergistic tools, producing compact yet highly discriminative representations. Second, for Relationship Redundancy Mitigation, it progressively refines the cross-modal alignment space by filtering misleading negative samples and adaptively reweighting informative pairs. The structural integration of these modules under a unified optimization objective provides a clear and interpretable pathway to retrieval. Extensive experiments on multiple benchmarks demonstrate that MEET consistently surpasses state-of-the-art methods, validating its effectiveness and generalizability. Kun Wang 0039, Yupeng Hu 0003, Hao Liu 0072, Lirong Jie, Liqiang Nie |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Cross-modal Representation Shift Refinement for Point-supervised Video Moment RetrievalabstractVideo Moment Retrieval (VMR) aims to retrieve temporal moments in videos that align with natural language queries, a task requiring cross-modal reasoning between video and text. Among various supervision paradigms, point-supervised VMR has emerged as a practical solution, leveraging single-frame annotations to significantly reduce annotation costs while maintaining competitive retrieval performance. However, this sparse supervision approach induces cross-modal representation shift. This shift complicates the model’s ability to accurately capture action sequences and associate text with visual content. To tackle this, we propose a novel framework called pseuDo fRame-based tempOral and semaNtic rEfinement (DRONE) with two key modules: (1) Pseudo-Frame Temporal Alignment (PTA), which embeds textual queries as pseudo-frames to enhance temporal coherence, and (2) Curriculum-Guided Semantic Refinement (CSR), which uses a progressive contrastive learning strategy to refine semantic representations from easy to hard cases. Extensive experiments show that DRONE achieves effective retrieval performance while keeping annotation costs low. Kun Wang 0039, Yupeng Hu 0003, Hao Liu 0072, Liqiang Nie |
ACM Trans. Inf. Syst. | 1 |
| 2026 | Visual Self-paced Iterative Learning for Unsupervised Temporal Action LocalizationabstractRecently, Temporal Action Localization (TAL) has garnered significant interest in information retrieval community. However, existing supervised/weakly supervised methods are heavily dependent on extensive labeled temporal boundaries and action categories, which is labor-intensive and time-consuming. Although some unsupervised methods have utilized the “iteratively clustering and localization” paradigm for TAL, they still suffer from two pivotal impediments: (1) unsatisfactory video clustering confidence, and (2) unreliable video pseudolabels for model training. To address these limitations, we present a novel self-paced iterative learning model to enhance clustering and localization training simultaneously, thereby facilitating more effective unsupervised TAL. Concretely, we improve the clustering confidence through exploring the contextual feature-robust visual information. Thereafter, we design two (constant- and variable-speed) incremental instance learning strategies for easy-to-hard model training, thus ensuring the reliability of these video pseudolabels and further improving overall localization performance. Extensive experiments on two public datasets demonstrate the superiority of our model over several state-of-the-art competitors. Yupeng Hu 0003, Han Jiang 0012, Hao Liu 0072, Kun Wang 0039, Haoyu Tang 0002, Liqiang Nie |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | CurMIM: Curriculum Masked Image ModelingabstractMasked Image Modeling (MIM), following “mask-andreconstruct” scheme, is a promising self-supervised method to learn scalable visual representation. Studies indicate that selecting an effective mask strategy is vital for MIM. However, existing approaches often rely on static pre-defined priors, which limit their ability to adapt mask strategies dynamically for network optimization. In this paper, we focus on the learning process of the network and introduce human-like curriculum into MIM for dynamic representation refinement, and propose an end-to-end framework Curriculum Masked Image Modeling (CurMIM). CurMIM consists of two components: Mask Priority Measurer, which acts as a curriculum learner to determine mask priority values using the network’s intrinsic state information, and Dual Adaptive Selector, which serves as a curriculum scheduler to create effective masks based on these values. With negligible extra parameters, our curriculum-based method consistently establishes noticeable improvements across varying model sizes and benchmarks, showing effectiveness and generalization. Hao Liu 0072, Kun Wang 0039, Haocong Wang, Yupeng Hu 0003, Liqiang Nie |
ICASSP | 2 |
| 2025 | Gaming for Boundary: Elastic Localization for Frame-Supervised Video Moment RetrievalabstractVideo moment retrieval aims to determine the temporal boundaries of moments within a video that are most relevant to textual queries. Unlike fully-supervised and weakly-supervised methods, frame-supervised methods use a single annotated frame to model the similarities between the target moment and queries. This task is still in its infancy due to the following challenges: 1) indiscernible intra-modal information and 2) inflexible inter-modal information interaction. In light of these challenges, we introduce the Gaming fOr elAstic Localization (GOAL) method for frame-supervised video moment retrieval. It enables target moment boundary localization from a novel strategic game perspective. GOAL encompasses two core components: a game-based paradigm to find the most reliable moment and a Dynamic Updating Technique (DUT) to continuously optimize moment retrieval through dynamic gradients, thereby refining boundary predictions with different feedback. Extensive experiments on Charades-STA, ActivityNet Captions, and TACoS have validated the effectiveness of GOAL. Hao Liu 0072, Yupeng Hu 0003, Kun Wang 0039, Yinwei Wei, Liqiang Nie |
SIGIR | 3 |
| 2024 | Prometheus: Out-of-distribution Fluid Dynamics Modeling with Disentangled Graph ODEabstractFluid dynamics modeling has received extensive attention in the machine learning community. Although numerous graph neural network (GNN) approaches have been proposed for this problem, the problem of out-of-distribution (OOD) generalization remains underexplored. In this work, we propose a new large-scale dataset Prometheus which simulates tunnel and pool fires across various environmental conditions and builds an extensive benchmark of 12 baselines, which demonstrates that the OOD generalization performance is far from satisfactory. To tackle this, this paper introduces a new approach named Disentangled Graph ODE (DGODE), which learns disentangled representations for continuous interacting dynamics modeling. In particular, we utilize a temporal GNN and a frequency network to extract semantics from historical trajectories into node representations and environment representations respectively. To mitigate the potential distribution shift, we minimize the mutual information between invariant node representations and the discretized environment features using adversarial learning. Then, they are fed into a coupled graph ODE framework, which models the evolution using neighboring nodes and dynamical environmental context. In addition, we enhance the stability of the framework by perturbing the environment features to enhance robustness. Extensive experiments validate the effectiveness of DGODE compared with state-of-the-art approaches. Hao Wu 0094, Huiyuan Wang, Kun Wang 0039, Weiyan Wang, Changan Ye, Yangyu Tao, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
ICML | 3 |
| 2024 | Explicit Granularity and Implicit Scale Correspondence Learning for Point-Supervised Video Moment LocalizationabstractVideo moment localization (VML) aims to identify the temporal boundary semantically matching the given query. Point-supervised VML balances localization accuracy and annotation cost but is still immature due to granularity alignment and scale perception issues. To this end, we propose a Semantic Granularity and Scale Correspondence Integration (SG-SCI) framework aimed at leveraging limited single-frame annotation for correspondence learning. It explicitly models semantic relations of different feature granularities and adaptively mines the implicit semantic scale, thereby enhancing feature representations of varying granularities and scales. SG-SCI uses granularity correspondence alignment to align semantics via latent prior knowledge and a scale correspondence learning to identify and address semantic scale differences. Extensive experiments on benchmark datasets have demonstrated the promising performance of our model over several state-of-the-art competitors. Kun Wang 0039, Hao Liu 0072, Lirong Jie, Zixu Li 0001, Yupeng Hu 0003, Liqiang Nie |
ACM Multimedia | 1 |
| 2024 | Semantic Collaborative Learning for Cross-Modal Moment LocalizationabstractLocalizing a desired moment within an untrimmed video via a given natural language query, i.e., cross-modal moment localization, has attracted widespread research attention recently. However, it is a challenging task because it requires not only accurately understanding intra-modal semantic information, but also explicitly capturing inter-modal semantic correlations (consistency and complementarity). Existing efforts mainly focus on intra-modal semantic understanding and inter-modal semantic alignment, while ignoring necessary semantic supplement. Consequently, we present a cross-modal semantic perception network for more effective intra-modal semantic understanding and inter-modal semantic collaboration. Concretely, we design a dual-path representation network for intra-modal semantic modeling. Meanwhile, we develop a semantic collaborative network to achieve multi-granularity semantic alignment and hierarchical semantic supplement. Thereby, effective moment localization can be achieved based on sufficient semantic collaborative learning. Extensive comparison experiments demonstrate the promising performance of our model compared with existing state-of-the-art competitors. Yupeng Hu 0003, Kun Wang 0039, Meng Liu 0006, Haoyu Tang 0002, Liqiang Nie |
ACM Trans. Inf. Syst. | 2 |
| 2021 | Coarse-to-Fine Semantic Alignment for Cross-Modal Moment LocalizationabstractVideo moment localization, as an important branch of video content analysis, has attracted extensive attention in recent years. However, it is still in its infancy due to the following challenges: cross-modal semantic alignment and localization efficiency. To address these impediments, we present a cross-modal semantic alignment network. To be specific, we first design a video encoder to generate moment candidates, learn their representations, as well as model their semantic relevance. Meanwhile, we design a query encoder for diverse query intention understanding. Thereafter, we introduce a multi-granularity interaction module to deeply explore the semantic correlation between multi-modalities. Thereby, we can effectively complete target moment localization via sufficient cross-modal semantic understanding. Moreover, we introduce a semantic pruning strategy to reduce cross-modal retrieval overhead, improving localization efficiency. Experimental results on two benchmark datasets have justified the superiority of our model over several state-of-the-art competitors. Yupeng Hu 0003, Liqiang Nie, Meng Liu 0006, Kun Wang 0039, Yinglong Wang 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Image Process. | 4 |