Hao Liu 0072

dblp:09/3214-72 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0003-1026-4499ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 From Interference to Stability: Adversarial Reliability Correction for Video Moment Retrieval with Relevance Feedback
abstract
Video Moment Retrieval (VMR) aims to retrieve target video moments that correspond to natural language queries. Most existing methods rely on a positive-only assumption that the queried moment always exists within the video, which limits their reliability in practical scenarios. Departing from this restrictive setting, we study Video Moment Retrieval with Relevance Feedback (VMR-RF), which requires models to both retrieve relevant moments and reject irrelevant queries. This task remains challenging due to the following issues: 1) Intrinsic Semantic Interference caused by visually similar but irrelevant moments, and 2) Propagative Decision Irreversibility induced by unidirectional relevance prediction. In light of these, we introduce AdversaRial Reliability cOrrection netWork (ARROW) for VMR-RF. ARROW adopts an active discriminative strategy through two synergetic components: a Gradient-induced Semantic Adversary (GSA) that probes model vulnerabilities by actively amplifying semantic interference, and an Adversarial Reliability Predictor (ARP) that quantifies prediction stability under such interference to effectively suppress unreliable decisions. Extensive experiments validate the effectiveness of ARROW.
Hao Liu 0072, Yupeng Hu 0003, Kun Wang 0039, Ruping Cao, Yutao Yao, Zilu Cai
SIGIR1
2026 METRON: Metabolic Dynamic Perception Kolmogorov-Arnold Network for Biological Age Estimation
abstract
Biological age is a more direct reflection of physiological status than chronological age, serving as a vital measure to evaluate health risks and aging interventions. While steroid metabolomics offers rich information for exploring aging mechanisms, the complex and nonlinear interactions within metabolic networks remain challenging in modeling. Here, we propose and describe METRON as a deep learning framework to predict biological ages from steroid metabolomics. Specifically, a Metabolite Interaction Perception Module (MIPM) is proposed to capture the interactions. Subsequently, a Group-Rational Kolmogorov-Arnold Network is also integrated to capture intricate dependencies and enhance the representation capability. We demonstrate that METRON achieves promising performance as compared to other machine learning and deep learning methods. Beyond performance, METRON offers interpretability by recovering the established markers such as Dehydroepiandrosterone (DHEA) and identifying 17-hydroxyprogesterone (17-OH-P4) as the key signature linked to hypothalamic-pituitary-adrenal axis dynamics. These results support the capacity of METRON not only to estimate biological age but also to uncover underappreciated metabolic drivers behind aging.
Zhongshen Li, Jixiang Yu, Shen You, Hao Liu 0072, Luyang Cai, Yuxuan Deng, Leyi Wei, Junkai Ji, Qiuzhen Lin, Xiangtao Li, Ka-Chun Wong
IEEE Trans. Comput. Biol. Bioinform.4
2026 Redundancy Mitigation: Toward Accurate and Efficient Image-Text Retrieval
abstract
Image-text retrieval (ITR) is a pivotal task in cross-modal research. However, existing methods often suffer from a fundamental yet overlooked challenge: redundancy. This issue manifests as both semantic redundancy within unimodal representations and relationship redundancy in cross-modal alignments. This not only inflates computational costs but also degrades retrieval accuracy by masking salient features and reinforcing spurious correlations. In this work, we are the first to explicitly analyze and address the ITR problem from a redundancy perspective by proposing the iMage-text rEtrieval rEdundancy miTigation (MEET) framework. MEET employs a cascaded, two-stage process to systematically mitigate both forms of redundancy. First, for Semantic Redundancy Mitigation, it repurposes deep hashing and quantization as synergistic tools, producing compact yet highly discriminative representations. Second, for Relationship Redundancy Mitigation, it progressively refines the cross-modal alignment space by filtering misleading negative samples and adaptively reweighting informative pairs. The structural integration of these modules under a unified optimization objective provides a clear and interpretable pathway to retrieval. Extensive experiments on multiple benchmarks demonstrate that MEET consistently surpasses state-of-the-art methods, validating its effectiveness and generalizability.
Kun Wang 0039, Yupeng Hu 0003, Hao Liu 0072, Lirong Jie, Liqiang Nie
IEEE Trans. Circuits Syst. Video Technol.3
2026 Cross-modal Representation Shift Refinement for Point-supervised Video Moment Retrieval
abstract
Video Moment Retrieval (VMR) aims to retrieve temporal moments in videos that align with natural language queries, a task requiring cross-modal reasoning between video and text. Among various supervision paradigms, point-supervised VMR has emerged as a practical solution, leveraging single-frame annotations to significantly reduce annotation costs while maintaining competitive retrieval performance. However, this sparse supervision approach induces cross-modal representation shift. This shift complicates the model’s ability to accurately capture action sequences and associate text with visual content. To tackle this, we propose a novel framework called pseuDo fRame-based tempOral and semaNtic rEfinement (DRONE) with two key modules: (1) Pseudo-Frame Temporal Alignment (PTA), which embeds textual queries as pseudo-frames to enhance temporal coherence, and (2) Curriculum-Guided Semantic Refinement (CSR), which uses a progressive contrastive learning strategy to refine semantic representations from easy to hard cases. Extensive experiments show that DRONE achieves effective retrieval performance while keeping annotation costs low.
Kun Wang 0039, Yupeng Hu 0003, Hao Liu 0072, Liqiang Nie
ACM Trans. Inf. Syst.3
2026 Visual Self-paced Iterative Learning for Unsupervised Temporal Action Localization
abstract
Recently, Temporal Action Localization (TAL) has garnered significant interest in information retrieval community. However, existing supervised/weakly supervised methods are heavily dependent on extensive labeled temporal boundaries and action categories, which is labor-intensive and time-consuming. Although some unsupervised methods have utilized the “iteratively clustering and localization” paradigm for TAL, they still suffer from two pivotal impediments: (1) unsatisfactory video clustering confidence, and (2) unreliable video pseudolabels for model training. To address these limitations, we present a novel self-paced iterative learning model to enhance clustering and localization training simultaneously, thereby facilitating more effective unsupervised TAL. Concretely, we improve the clustering confidence through exploring the contextual feature-robust visual information. Thereafter, we design two (constant- and variable-speed) incremental instance learning strategies for easy-to-hard model training, thus ensuring the reliability of these video pseudolabels and further improving overall localization performance. Extensive experiments on two public datasets demonstrate the superiority of our model over several state-of-the-art competitors.
Yupeng Hu 0003, Han Jiang 0012, Hao Liu 0072, Kun Wang 0039, Haoyu Tang 0002, Liqiang Nie
ACM Trans. Multim. Comput. Commun. Appl.3
2025 CurMIM: Curriculum Masked Image Modeling
abstract
Masked Image Modeling (MIM), following “mask-andreconstruct” scheme, is a promising self-supervised method to learn scalable visual representation. Studies indicate that selecting an effective mask strategy is vital for MIM. However, existing approaches often rely on static pre-defined priors, which limit their ability to adapt mask strategies dynamically for network optimization. In this paper, we focus on the learning process of the network and introduce human-like curriculum into MIM for dynamic representation refinement, and propose an end-to-end framework Curriculum Masked Image Modeling (CurMIM). CurMIM consists of two components: Mask Priority Measurer, which acts as a curriculum learner to determine mask priority values using the network’s intrinsic state information, and Dual Adaptive Selector, which serves as a curriculum scheduler to create effective masks based on these values. With negligible extra parameters, our curriculum-based method consistently establishes noticeable improvements across varying model sizes and benchmarks, showing effectiveness and generalization.
Hao Liu 0072, Kun Wang 0039, Haocong Wang, Yupeng Hu 0003, Liqiang Nie
ICASSP1
2025 DCount: Decoupled Spatial Perception and Attribute Discrimination for Referring Expression Counting
abstract
Referring Expression Counting (REC) is an emerging task that aims to count specific objects in images based on textual phrases describing their attributes and categories. While current REC baselines inherit architectures from pre-trained open-vocabulary object detectors and demonstrate promising counting and localization capabilities, they overlook critical limitations in the original single-decoder design with shared object queries. This architectural constraint entangles the semantic and localization perception processes, hindering fine-grained understanding of attribute-aware visual features. To address these challenges, we propose DCount, a decoupled counting framework comprising two innovative components: a Decoupled Dual-Decoder (DDD) module and an Attribute Semantic Discriminator (ASD) module. The DDD module separates spatial perception tasks by employing distinct semantic and localization decoders with task-specific object queries, thereby enhancing the capture of discriminative visual features. Building upon the positional and semantic feedback from DDD, the ASD module introduces a two-stage filtering strategy to explicitly mine challenging hard negative attribute samples in the visual domain, while synergistically refining attribute discrimination across both modalities through contrastive learning in the textual domain. Our method achieves state-of-the-art results on both the REC and Zero-Shot Object Counting (ZSOC) benchmarks.
Ming Li 0083, Yupeng Hu 0003, Yinwei Wei, Hao Liu 0072, Haocong Wang, Weili Guan
ACM Multimedia4
2025 Gaming for Boundary: Elastic Localization for Frame-Supervised Video Moment Retrieval
abstract
Video moment retrieval aims to determine the temporal boundaries of moments within a video that are most relevant to textual queries. Unlike fully-supervised and weakly-supervised methods, frame-supervised methods use a single annotated frame to model the similarities between the target moment and queries. This task is still in its infancy due to the following challenges: 1) indiscernible intra-modal information and 2) inflexible inter-modal information interaction. In light of these challenges, we introduce the Gaming fOr elAstic Localization (GOAL) method for frame-supervised video moment retrieval. It enables target moment boundary localization from a novel strategic game perspective. GOAL encompasses two core components: a game-based paradigm to find the most reliable moment and a Dynamic Updating Technique (DUT) to continuously optimize moment retrieval through dynamic gradients, thereby refining boundary predictions with different feedback. Extensive experiments on Charades-STA, ActivityNet Captions, and TACoS have validated the effectiveness of GOAL.
Hao Liu 0072, Yupeng Hu 0003, Kun Wang 0039, Yinwei Wei, Liqiang Nie
SIGIR1
2024 Explicit Granularity and Implicit Scale Correspondence Learning for Point-Supervised Video Moment Localization
abstract
Video moment localization (VML) aims to identify the temporal boundary semantically matching the given query. Point-supervised VML balances localization accuracy and annotation cost but is still immature due to granularity alignment and scale perception issues. To this end, we propose a Semantic Granularity and Scale Correspondence Integration (SG-SCI) framework aimed at leveraging limited single-frame annotation for correspondence learning. It explicitly models semantic relations of different feature granularities and adaptively mines the implicit semantic scale, thereby enhancing feature representations of varying granularities and scales. SG-SCI uses granularity correspondence alignment to align semantics via latent prior knowledge and a scale correspondence learning to identify and address semantic scale differences. Extensive experiments on benchmark datasets have demonstrated the promising performance of our model over several state-of-the-art competitors.
Kun Wang 0039, Hao Liu 0072, Lirong Jie, Zixu Li 0001, Yupeng Hu 0003, Liqiang Nie
ACM Multimedia2
2024 Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs
abstract
Although Large Language Models (LLMs) have demonstrated significant capabilities in executing complex tasks in a zero-shot manner, they are susceptible to jailbreak attacks and can be manipulated to produce harmful outputs. Recently, a growing body of research has categorized jailbreak attacks into token-level and prompt-level attacks. However, previous work primarily overlooks the diverse key factors of jailbreak attacks, with most studies concentrating on LLM vulnerabilities and lacking exploration of defense-enhanced LLMs. To address these issues, we introduced JailTrickBench to evaluate the impact of various attack settings on LLM performance and provide a baseline for jailbreak attacks, encouraging the adoption of a standardized evaluation framework. Specifically, we evaluate the eight key factors of implementing jailbreak attacks on LLMs from both target-level and attack-level perspectives. We further conduct seven representative jailbreak attacks on six defense methods across two widely used datasets, encompassing approximately 354 experiments with about 55,000 GPU hours on A800-80G. Our experimental results highlight the need for standardized benchmarking to evaluate these attacks on defense-enhanced LLMs. Our code is available at https://github.com/usail-hkust/JailTrickBench.
Fan Liu 0008, Hao Liu 0072
NeurIPS3
2023 CoGCN: co-occurring item-aware GCN for recommendation
Xinxiao Zhao, Fan Liu 0008, Hao Liu 0072, Haoyu Tang 0002, Yupeng Hu 0003
Neural Comput. Appl.3