Haoyu Tang 0002

dblp:254/9341-2 · DBLP profile ↗
← Back
28ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0003-1344-2513ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 5 first-author · 16 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Pseudo Multi-view K-means Clustering
abstract
Clustering with k-means is well-established and efficient, but often struggles with complex data distributions because the clustering performance hinges on how well the centroids capture the data distribution, and conventional k-means usually fails to produce representative centroids under such conditions. To address this limitation, we propose Pseudo Multi-view K-means Clustering (PMKC), a novel framework that simulates a multi-view learning paradigm within a single-view setting by generating multiple soft k-means decompositions. Each decomposition can be treated as an individual view and investigates a distinct perspective of the data. Specifically, to encourage complementary structure, we impose an independence constraint among cluster centers, and to integrate these diverse clusterings, we model the soft assignment matrices as a third-order tensor and apply low-rank regularization to extract a shared latent structure. This design not only enhances clustering robustness but also improves the stability and consistency of the final results. Experimental results on several benchmark datasets demonstrate that PMKC achieves superior clustering performance compared to state-of-the-art methods.
Jinqian Chen, Jihua Zhu, Haoyu Tang 0002, Qinghai Zheng
AAAI3
2026 Decompose and Conquer: Compositional Reasoning for Zero-Shot Temporal Action Localization
abstract
Current Zero-Shot Temporal Action Localization (ZSTAL) methods, whether training-based or training-free ones, still predominantly rely on a single, unified query to localize an entire action. This unified representation is fundamentally ill-suited for complex real-world activities, as it fails to capture their internal compositional structure and adapt to dynamic, multi-stage variations across videos. To address this, we regard ZSTAL as a compositional reasoning task and introduce CASCADE, a Context-Aware Staged Action DEcomposition framework. Inspired by the human cognitive process of perceiving context, decomposing events, and reconstructing instances, CASCADE follows a training-free pipeline. It first perceives the video's context by leveraging a Multimodal Large Language Model (MLLM) to both filter out irrelevant actions and then generate a rich, video-specific caption for each action present in the video. An LLM then decomposes this caption into multiple, temporally ordered stages, which serve as fine-grained queries to guide the MLLM in estimating frame-level confidence scores. Recognizing that this decomposition can fragment a single action, a novel hierarchical merging logic then reconstructs complete instances by intelligently fusing these preliminary temporal segments based on their semantic progression and coherence. Extensive experiments and ablation studies on THUMOS14 and ActivityNet-1.3 show that CASCADE not only sets a new state-of-the-art among training-free methods but, most notably, significantly outperforms all prior training-based approaches on ActivityNet-1.3.
Haoyu Tang 0002, Tianyuan Liang, Han Jiang 0012, Qinghai Zheng, Yupeng Hu 0003
AAAI1
2026 Resonating with RoPE: Spectral Quantization for High-Fidelity Key Cache Compression
abstract
The linear growth of KV cache bottlenecks long-context LLMs, yet RoPE-induced oscillations complicate Key cache quantization.To address this issue, we propose SpectrumQuant, a frequency-domain framework that utilizes the Discrete Cosine Transform (DCT) to convert these oscillations into sparse spectral representations.Specifically, our pipeline integrates dominant frequency extraction, hybrid bit-width allocation, and high-frequency preemphasis to maximize fidelity while minimizing memory footprint.To eliminate computational overhead, we develop fused Triton kernels featuring deferred inverse transformation and on-chip sparse accumulation.Extensive experiments on several benchmarks confirm SpectrumQuant achieves efficient compression with performance and latency comparable to FP16 baselines.
Haoyu Tang 0002, Tianyuan Liang, Yupeng Hu 0003, Weili Guan
ACL (1)2
2026 From One Comes Two: A Tensorized Graph Learning Framework for Clustering
Qinghai Zheng, Jihua Zhu, Yuanlong Yu 0001, Haoyu Tang 0002
IEEE Trans. Knowl. Data Eng.4
2026 Visual Self-paced Iterative Learning for Unsupervised Temporal Action Localization
abstract
Recently, Temporal Action Localization (TAL) has garnered significant interest in information retrieval community. However, existing supervised/weakly supervised methods are heavily dependent on extensive labeled temporal boundaries and action categories, which is labor-intensive and time-consuming. Although some unsupervised methods have utilized the “iteratively clustering and localization” paradigm for TAL, they still suffer from two pivotal impediments: (1) unsatisfactory video clustering confidence, and (2) unreliable video pseudolabels for model training. To address these limitations, we present a novel self-paced iterative learning model to enhance clustering and localization training simultaneously, thereby facilitating more effective unsupervised TAL. Concretely, we improve the clustering confidence through exploring the contextual feature-robust visual information. Thereafter, we design two (constant- and variable-speed) incremental instance learning strategies for easy-to-hard model training, thus ensuring the reliability of these video pseudolabels and further improving overall localization performance. Extensive experiments on two public datasets demonstrate the superiority of our model over several state-of-the-art competitors.
Yupeng Hu 0003, Han Jiang 0012, Hao Liu 0072, Kun Wang 0039, Haoyu Tang 0002, Liqiang Nie
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Boundary-Aware Temporal Dynamic Pseudo-Supervision Pairs Generation for Zero-Shot Natural Language Video Localization
abstract
Zero-shot Natural Language Video Localization (NLVL) aims to automatically generate moments and corresponding pseudo queries from raw videos for the training of the localization model without any manual annotations. Existing approaches typically produce pseudo queries as simple words, which overlook the complexity of queries in real-world scenarios. Considering the powerful text modeling capabilities of large language models (LLMs), leveraging LLMs to generate complete queries that are closer to human descriptions is a potential solution. However, directly integrating LLMs into existing approaches introduces several issues, including insensitivity, isolation, and lack of regulation, which prevent the full exploitation of LLMs to enhance zero-shot NLVL performance. To address these issues, we propose BTDP, an innovative framework for Boundary-aware Temporal Dynamic Pseudo-supervision pairs generation. Our method contains two crucial operations: 1) Boundary Segmentation that identifies both visual boundaries and semantic boundaries to generate the atomic segments and activity descriptions, tackling the issue of insensitivity. 2) Context Aggregation that employs the LLMs with a self-evaluation process to aggregate and summarize global video information for optimized pseudo moment-query pairs, tackling the issue of isolation and lack of regulation. Comprehensive experimental results on the Charades-STA and ActivityNet Captions datasets demonstrate the effectiveness of our BTDP method.
Xiongwen Deng, Haoyu Tang 0002, Han Jiang 0012, Qinghai Zheng, Jihua Zhu
AAAI2
2025 Towards Stable and Storage-efficient Dataset Distillation: Matching Convexified Trajectory
abstract
The rapid evolution of deep learning and large language models has led to an exponential growth in the demand for training data, prompting the development of Dataset Distillation methods to address the challenges of managing large datasets. Among these, Matching Training Trajectories (MTT) has been a prominent approach, which replicates the training trajectory of an expert network on real data with a synthetic dataset. However, our investigation found that this method suffers from three significant limitations: 1. Instability of expert trajectory generated by Stochastic Gradient Descent (SGD); 2. Low convergence speed of the distillation process; 3. High storage consumption of the expert trajectory. To address these issues, we offer a new perspective on understanding the essence of Dataset Distillation and MTT through a simple transformation of the objective function, and introduce a novel method called Matching Convexified Trajectory (MCT), which aims to provide better guidance for the student trajectory. MCT creates convex combinations of expert trajectories by selecting a few expert models, guiding student networks to converge quickly and stably. This trajectory is not only easier to store, but also enables continuous sampling strategies during the distillation process, ensuring thorough learning and fitting of the entire expert trajectory. The comprehensive experiment of three public datasets verified that MCT is superior to the traditional MTT method.
Leon Wenliang Zhong, Haoyu Tang 0002, Qinghai Zheng, Yupeng Hu 0003, Weili Guan
CVPR2
2025 FACE: A Dual-Template and Adaptive Curriculum Framework for Unsupervised Text-Based Person Search
abstract
Text-Based Person Search, which aims to retrieve target pedestrian images using natural language descriptions, has garnered significant attention in multimedia research due to its potential in suspect retrieval and missing person identification. While supervised and weakly supervised methods rely on costly annotated training data, unsupervised TBPS eliminates the need for textual descriptions or identity annotations, presenting a more practical paradigm. Current unsupervised TBPS approaches face two primary challenges: 1) Predefined attribute templates for caption generation limit linguistic diversity and real-world adaptability, and 2) Threshold-based sample selection using pre-trained vision-language models (VLMs) introduces noisy pairs due to inadequate pedestrian-specific representation. To address these limitations, we propose FACE, a unified framework featuring Dual-template Caption Generation (DCG) and Adaptive Curriculum Training (ACT). The DCG module generates high-quality captions through complementary flexible-style (natural language) and fixed-style (attribute-enumerated) templates, enhanced by LLM-based noise filtering. The ACT framework progressively refines training through a self-improving loop: initial high-confidence sample selection using VLMs bootstraps the model, while evolving feature representations enable dynamic incorporation of harder samples through curriculum learning. This dual strategy achieves mutual reinforcement between caption quality and model discriminability. Extensive experiments on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets under unsupervised settings demonstrate that our framework achieves the state-of-the-art performance.
Xiaoxuan Mu, Haoyu Tang 0002, Han Jiang 0012, Tianyuan Liang, Qinghai Zheng, Jihua Zhu
ACM Multimedia2
2025 Superpixel Segmentation With Edge Guided Local-Global Attention Network
abstract
Superpixel segmentation aims to automatically group visually similar pixels within an image into compact regions. This approach provides an efficient low-level representation of image data, effectively reducing the complexity of image primitives for subsequent vision tasks. Recent deep convolutional networks have shown their advantages in superpixel segmentation task. However, many existing deep learning methods still struggle to preserve object edges and accurately perceive similar pixels. This limitation can be attributed to their inadequate ability to model edge information and capture effective context within the image. To address these issues, we propose an Edge guided Local-Global Attention Network (ELGANet) for superpixel segmentation. Specifically, we first devise an Edge Enhancement Module (EeEM), which integrates multiple edge features into the superpixel-friendly features. Then, we develop a Local-Global Attention Module (LGAM) to analyze the relationship between pixels and local or global region patches, expecting to obtain effective context information for grouping similar pixels. The edge features and deep global semantic features are subsequently fused to generate the superpixel-friendly features. The final superpixel-friendly features are then mapped into final superpixels. Extensive experiments on four benchmark datasets demonstrate the effectiveness and superiority of our ELGANet compared with ten state-of-the-art models.
Zhengyu Sun, Haoyu Tang 0002, Yupeng Hu 0003, Xuemeng Song, Liqiang Nie
IEEE Trans. Circuits Syst. Video Technol.4
2025 CMIRNet: Cross-Modal Interactive Reasoning Network for Referring Image Segmentation
abstract
Referring Image Segmentation (RIS) aims to semantically segment the target object (referent) in alignment with the provided natural language query. Existing works still suffer from that the non-referent was segmented mistakenly, which can be attributed to the insufficient comprehension of vision and language. To tackle this problem, we propose a Cross-Modal Interactive Reasoning Network (CMIRNet) to explore semantic information that consistently existed between vision and language. Specifically, we first devise a novel Text-Guided Multi-Modality Joint Encoder (TGMM-JE), where the key expression can be extracted and the important visual features will be encoded under the continuous guidance of language expression. Then, we design a Cross-Graph Interactive Positioning (CGIP) module to locate the key pixels of the referent object in deepest layer. The multi-modality graph data is constructed between visual and linguistic features, and the important pixels can be positioned from cross-graph interaction and intra-graph reasoning. Finally, a novel Cross-Modal Attention Enhanced DEcoder (CMAE-DE) is dedicated to refine the referent object mask from coarse to fine progressively, where hybrid cross modal attentions are explored to enhance the representation of referent object. Extensive ablation studies validate the efficacy of our key modules and comprehensive experimental results show the superiority of our proposed model over 22 state-of-the-art (SOTA) models.
Tianxiang Xiao, Yutong Liu 0002, Haoyu Tang 0002, Yupeng Hu 0003, Liqiang Nie
IEEE Trans. Circuits Syst. Video Technol.4
2025 Cross-Model Nested Fusion Network for Salient Object Detection in Optical Remote Sensing Images
abstract
Recently, salient object detection (SOD) in optical remote sensing images, dubbed ORSI-SOD, has attracted increasing research interest. Although deep-based models have achieved impressive performance, several limitations remain: a single image contains multiple objects with varying scales, complex topological structures, and background interference. These unresolved issues render ORSI-SOD a challenging task. To address these challenges, we introduce a distinctive cross-model nested fusion network (CMNFNet), which leverages heterogeneous features to increase the performance of ORSI-SOD. Specifically, the proposed model comprises two heterogeneous encoders, a conventional CNN-based encoder that can model local features, and a specially designed graph convolutional network (GCN)-based encoder with local and global receptive fields that can model local and global features simultaneously. To effectively differentiate between multiple salient objects of different sizes or complex topological structures within an image, we project the image into two different graphs with different receptive fields and conduct message passing through two parallel graph convolutions. Finally, the heterogeneous features extracted from the two encoders are fused in the well-designed attention enhanced cross model nested fusion module (AECMNFM). This module is meticulously crafted to integrate features progressively, allowing the model to adaptively eliminate background interference while simultaneously refining the feature representations. We conducted comprehensive experimental analyzes on benchmark datasets. The results demonstrate the superiority of our CMNFNet over 16 state-of-the-art (SOTA) models.
Yupeng Hu 0003, Haoyu Tang 0002, Runmin Cong, Liqiang Nie
IEEE Trans. Cybern.4
2025 HDNet: A Hybrid Domain Network With Multiscale High-Frequency Information Enhancement for Infrared Small-Target Detection
abstract
The InfraRed Small Target Detection (IRSTD) task involves identifying and separating small targets from complex backgrounds. However, these targets pose significant challenges due to their small, variable sizes and dim appearance with a low signal-to-noise ratio, often obscured by cluttered backgrounds. Standard spatial-domain Convolutional Neural Networks (CNNs) act as low-pass filters, hindering their ability to detect small, variably sized, low-contrast targets against complex backgrounds. Infrared images also exhibit diverse spectral energy distributions, yet CNNs lack a global spectral view to discern these patterns, making them susceptible to background clutter. To address these shortcomings, we propose a novel Hybrid Domain Network (HDNet), which fuses frequency-domain features with conventional spatial-domain CNN features to markedly enhance target-background contrast and explicitly suppress background interference. Specifically, HDNet comprises two main branches: the spatial domain branch and the frequency domain branch. In the spatial domain, we innovatively introduce a Multi-scale Atrous Contrast convolution (MAC) module, utilizing multiple parallel atrous contrast convolutions with varying kernel sizes to enhance perception of small, variably sized targets. In the frequency domain, we have specifically designed the Dynamic High-Pass Filter (DHPF) module, hierarchically calculating low-frequency signal energy and dynamically removing specific low-frequency information to preserve high-frequency image details. This effectively filters out slowly varying low-frequency backgrounds, highlighting small targets. Comprehensive ablation studies and experimental analysis on three datasets (IRSTD-1K, NUAA-SIRST, NUDT-SIRST) validate HD-Net’s effectiveness and superiority compared to 26 state-of-the-art (SOTA) methods. The source code is available at: https://github.com/xumingzhu989/HDNet-TGRS.
Haoyu Tang 0002, Yupeng Hu 0003, Liqiang Nie
IEEE Trans. Geosci. Remote. Sens.4
2024 Exploiting the Social-Like Prior in Transformer for Visual Reasoning
abstract
Benefiting from instrumental global dependency modeling of self-attention (SA), transformer-based approaches have become the pivotal choices for numerous downstream visual reasoning tasks, such as visual question answering (VQA) and referring expression comprehension (REC). However, some studies have recently suggested that SA tends to suffer from rank collapse thereby inevitably leads to representation degradation as the transformer layer goes deeper. Inspired by social network theory, we attempt to make an analogy between social behavior and regional information interaction in SA, and harness two crucial notions of structural hole and degree centrality in social network to explore the possible optimization towards SA learning, which naturally deduces two plug-and-play social-like modules. Based on structural hole, the former module allows to make information interaction in SA more structured, which effectively avoids redundant information aggregation and global feature homogenization for better rank remedy, followed by latter module to comprehensively characterize and refine the representation discrimination via considering degree centrality of regions and transitivity of relations. Without bells and whistles, our model outperforms a bunch of baselines by a noticeable margin when considering our social-like prior on five benchmarks in VQA and REC tasks, and a series of explanatory results are showcased to sufficiently reveal the social-like behaviors in SA.
Yupeng Hu 0003, Xuemeng Song, Haoyu Tang 0002, Liqiang Nie
AAAI4
2024 Two-Stage Information Bottleneck For Temporal Language Grounding
abstract
Existing cross-modal fusion methods for temporal language grounding suffer from issues where the generated cross-modal embeddings are affected by noise in their unimodal representations, thus hindering the expression of interactions between language query and the target video segment. Furthermore, the cross-modal representations contain many irrelevant redundancies, which compromises the quality of cross-modal features and thus interferes with accurate moment localization. To address these drawbacks, we propose a novel CrOss-modaL information-constraineD (COLD) model for temporal language grounding, aiming at learning a robust cross-modal embedding representation devoid of irrelevant redundancies and maximizing the interaction between language query and target video moment. Specifically, our model is built upon the principles of the information bottleneck and features two information-constrained modules from different perspectives: 1) the Cross-modal Highlight Information Bottleneck module is designed to maximize the mutual information between language query and target video moment; 2) the Fusion Information Bottleneck module is introduced to constrain the correlations between the cross-modal representations and the localization labels. Comprehensive experimental results on two public benchmark datasets demonstrate the superiority of the proposed model.
Haoyu Tang 0002, Shuaike Zhang, Ming Yan 0008, Ji Zhang 0011, Yupeng Hu 0003, Liqiang Nie
ICME1
2024 Breaking Barriers of System Heterogeneity: Straggler-Tolerant Multimodal Federated Learning via Knowledge Distillation
Jinqian Chen, Haoyu Tang 0002, Ming Yan 0008, Ji Zhang 0011, Yupeng Hu 0003, Liqiang Nie
IJCAI2
2024 Revisiting Unsupervised Temporal Action Localization: The Primacy of High-Quality Actionness and Pseudolabels
abstract
Recently, temporal action localization (TAL) methods, especially the weakly-supervised and unsupervised ones, have become a hot research topic. Existing unsupervised methods follow an iterative ''clustering and training'' strategy with diverse model designs during training stage, while they often overlook maintaining consistency between these stages, which is crucial: more accurate clustering results can reduce the noises of pseudolabels and thus enhance model training, while more robust training can in turn enrich clustering feature representation. We identify two critical challenges in unsupervised scenarios: 1. What features should the model generate for clustering? 2. Which pseudolabeled instances from clustering should be chosen for model training? After extensive explorations, we proposed a novel yet simple framework called Consistency-Oriented Progressive high actionness Learning to address these issues. For feature generation, our framework adopts a High Actionness snippet Selection (HAS) module to generate more discriminative global video features for clustering from the enhanced actionness features obtained from a designed Inner-Outer Consistency Network (IOCNet). For pseudolabel selection, we introduces a Progressive Learning With Representative Instances (PLRI) strategy to identify the most reliable and informative instances within each cluster for model training. These three modules, HAS, IOCNet, and PLRI, synergistically improve consistency in model training and clustering performance. Extensive experiments on THUMOS'14 and ActivityNet v1.2 datasets under both unsupervised and weakly-supervised settings demonstrate that our framework achieves the state-of-the-art results.
Han Jiang 0012, Haoyu Tang 0002, Ming Yan 0008, Ji Zhang 0011, Yupeng Hu 0003, Jihua Zhu, Liqiang Nie
ACM Multimedia2
2024 Twin Reciprocal Completion for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering is an important and challenging task, which has attracted significant attention in recent years. The key objective of incomplete multi-view clustering is to excavate the underlying avaliable consistency of multi-view data, so as to enable the effective reconstruction of missing views for clustering. In this paper, we introduce a completion framework that deeply explores the underlying consistency and effectively completes the missing views. Following that, we propose a novel Twin Reciprocal Completion for Incomplete multi-view clustering, termed TRC-IMC for short. To be specific, TRC-IMC jointly conducts the Completion in Feature space (CF) and the Completion in Subspace (CS) to reciprocally complete the data with missing views. The underlying high-order consistency of multi-view data can be fully explored in both the feature space and subspace to guide the completion process of missing views. Extensive experiments are conducted on eight real-world multi-view datasets, and experimental results indicate the promising performance of our method, compared to several state-of-the-arts.
Qinghai Zheng, Haoyu Tang 0002
IEEE Trans. Circuits Syst. Video Technol.2
2024 Heterogeneous Feature Collaboration Network for Salient Object Detection in Optical Remote Sensing Images
abstract
In recent years, the task of salient object detection in optical remote sensing images (RSI-SOD) has gained increasing interest. Despite some advancement in current methods, challenges such as the irregular topology of salient objects and cluttered backgrounds in optical RSI remain. To tackle these issues, we propose a novel heterogeneous feature collaboration network (HFCNet). Specifically, we design a new hybrid heterogeneous encoder that combines CNN and transformer to extract a set of heterogeneous features, famous in modeling local and global information, respectively. Subsequently, the adaptive global-local integration (AGLI) module is devised to integrate these complementary heterogeneous features through our feature alignment methods at global and local levels, so the global irregular topology structure and local details can be well-modeled. Furthermore, the proposed saliency-guided attention enhanced decoder (SGAED) leverages deep salient cues to guide the shallow decoders to pay more attention to important areas and suppress irrelevant areas, reducing the interference of cluttered backgrounds. Extensive experiments on three benchmark datasets have confirmed the significant superiority of our method compared with 18 state-of-the-art methods. All codes and results of our method are available athttps://github.com/xumingzhu989/HFCNet-TGRS.
Yutong Liu 0002, Tianxiang Xiao, Haoyu Tang 0002, Yupeng Hu 0003, Liqiang Nie
IEEE Trans. Geosci. Remote. Sens.4
2024 Semantic Collaborative Learning for Cross-Modal Moment Localization
abstract
Localizing a desired moment within an untrimmed video via a given natural language query, i.e., cross-modal moment localization, has attracted widespread research attention recently. However, it is a challenging task because it requires not only accurately understanding intra-modal semantic information, but also explicitly capturing inter-modal semantic correlations (consistency and complementarity). Existing efforts mainly focus on intra-modal semantic understanding and inter-modal semantic alignment, while ignoring necessary semantic supplement. Consequently, we present a cross-modal semantic perception network for more effective intra-modal semantic understanding and inter-modal semantic collaboration. Concretely, we design a dual-path representation network for intra-modal semantic modeling. Meanwhile, we develop a semantic collaborative network to achieve multi-granularity semantic alignment and hierarchical semantic supplement. Thereby, effective moment localization can be achieved based on sufficient semantic collaborative learning. Extensive comparison experiments demonstrate the promising performance of our model compared with existing state-of-the-art competitors.
Yupeng Hu 0003, Kun Wang 0039, Meng Liu 0006, Haoyu Tang 0002, Liqiang Nie
ACM Trans. Inf. Syst.4
2023 Label Information Bottleneck for Label Enhancement
abstract
In this work, we focus on the challenging problem of Label Enhancement (LE), which aims to exactly recover label distributions from logical labels, and present a novel Label Information Bottleneck (LIB) method for LE. For the recovery process of label distributions, the label irrelevant information contained in the dataset may lead to unsatisfactory recovery performance. To address this limitation, we make efforts to excavate the essential label relevant information to improve the recovery performance. Our method formulates the LE problem as the following two joint processes: 1) learning the representation with the essential label relevant information, 2) recovering label distributions based on the learned representation. The label relevant information can be excavated based on the “bottleneck” formed by the learned representation. Significantly, both the label relevant information about the label assignments and the label relevant information about the label gaps can be explored in our method. Evaluation experiments conducted on several benchmark label distribution learning datasets verify the effectiveness and competitiveness of LIB. Our source codes are available at https://github.com/qinghai-zheng/LIBLE.
Qinghai Zheng, Jihua Zhu, Haoyu Tang 0002
CVPR3
2023 CoGCN: co-occurring item-aware GCN for recommendation
Xinxiao Zhao, Fan Liu 0008, Hao Liu 0072, Haoyu Tang 0002, Yupeng Hu 0003
Neural Comput. Appl.5
2023 Graph-Guided Unsupervised Multiview Representation Learning
abstract
Without the valuable label information to guide the learning process, it is demanding to fully excavate and integrate the underlying information from different views to learn the unified multi-view representation. This paper focuses on this challenge and presents a novel method, termed Graph-guided Unsupervised Multi-view Representation Learning (GUMRL), taking full advantage of multi-view graph information during the learning process. To be specific, GUMRL jointly conducts the view-specific feature representation learning, which is under the guidance of graph information, and the unified feature representation learning, which fuses the underlying graph information of different views to learn the desired unified multi-view feature representation. Regarding downstream tasks, such as clustering and classification, the classic single-view algorithms can be directly performed on the learned unified multi-view representation. The designed objective function is effectively optimized based on an alternating direction minimization method, and experiments conducted on six real-world multi-view datasets show the effectiveness and competitiveness of our GUMRL, compared to several state-of-the-art methods.
Qinghai Zheng, Jihua Zhu, Zhongyu Li 0002, Haoyu Tang 0002
IEEE Trans. Circuits Syst. Video Technol.4
2023 Adaptive Edge-Aware Semantic Interaction Network for Salient Object Detection in Optical Remote Sensing Images
abstract
In recent years, the task of salient object detection in optical remote sensing images (RSI-SOD) has received extensive attention. Benefiting from the development of deep learning, much progress has been made in RSI-SOD field. However, existing methods still face challenges in addressing various issues present in optical RSI, including uncertain numbers of salient objects, cluttered backgrounds, and interference from shadows. To address these challenges, we propose a novel approach, Adaptive Edge-aware Semantic Interaction Network (AESINet) for efficient salient object detection. Specifically, to improve the extraction of complex edge information, we design a Local Detail Aggregation Module (LDAM). This module can adaptively enhance the edge information of salient objects by leveraging our proposed difference perception mechanism. Notably, our difference perception mechanism is a novel edge enhancement method without the supervision of edge groundtruth. Additionally, to accurately locate salient objects of varying numbers and scales, we design a Multi-scale Feature Enhancement Module (MFEM), which effectively captures and utilizes multi-scale information. Moreover, we design the Deep Semantic Interaction Module (DSIM) to identify salient objects amidst cluttered backgrounds and effectively mitigate the interference of shadows. We conduct extensive experiments on three well-established optical RSI datasets and the results demonstrate that our proposed model outperforms 14 state-of-the-art methods. All codes and detection results are available at https://github.com/xumingzhu989/AESINet-TGRS.
Xiangyu Zeng 0004, Haoyu Tang 0002, Yupeng Hu 0003, Liqiang Nie
IEEE Trans. Geosci. Remote. Sens.4
2023 Generalized Label Enhancement With Sample Correlations
abstract
Recently, label distribution learning (LDL) has drawn much attention in machine learning, where LDL model is learned from labelel instances. Different from single-label and multi-label annotations, label distributions describe the instance by multiple labels with different intensities and accommodate to more general scenes. Since most existing machine learning datasets merely provide logical labels, label distributions are unavailable in many real-world applications. To handle this problem, we propose two novel label enhancement methods, i.e., Label Enhancement with Sample Correlations (LESC) and generalized Label Enhancement with Sample Correlations (gLESC). More specifically, LESC employs a low-rank representation of samples in the feature space, and gLESC leverages a tensor multi-rank minimization to further investigate the sample correlations in both the feature space and label space. Benefitting from the sample correlations, the proposed methods can boost the performance of label enhancement. Extensive experiments on 14 benchmark datasets demonstrate the effectiveness and superiority of our methods.
Qinghai Zheng, Jihua Zhu, Haoyu Tang 0002, Xinyuan Liu 0001, Zhongyu Li 0002, Huimin Lu 0001
IEEE Trans. Knowl. Data Eng.3
2022 Multi-Level Query Interaction for Temporal Language Grounding
abstract
Understanding what is happening in the surveillance video is important for human-machine interface in transportation systems, where temporal language grounding is one of the key tasks, targeting at localizing the desired moment in an untrimmed video with a given sentence query that is relevant to the moment. This task is challenging due to the following reasons: 1) the requirement of understanding the video contents and query semantics comprehensively, and 2) building the bridge between the cross-modal semantics. To tackle these problems, early methods first sample video clips and then match them with the sentence to find the most relevant one. To reduce the computational complexity associated with video clip sampling, recent methods directly predict the temporal boundaries of the desired moment on the fused features of the sentence and the video frames. However, all the previous methods often learn the word-level or phrase-level features of the sentence, or directly generates the global sentence representation by attention mechanisms or graph network. However, we argue that applying only word-level or phrase-level semantic information and cross-modal interactions is not enough to fully capture the correspondence between the video and the query. To this end, we proposed a novel Multi-level Query Exploration and Interaction (MQEI) model, which explores the semantics in both the word- and phrase-level and captures the multi-level interactions between the video and the query through an attention module. Extensive experiments on two public benchmark datasets ActivityNet Captions and Charades-STA demonstrate that the proposed model can outperform all the state-of-the-art methods consistently.
Haoyu Tang 0002, Jihua Zhu, Lin Wang 0026, Qinghai Zheng
IEEE Trans. Intell. Transp. Syst.1
2022 Frame-Wise Cross-Modal Matching for Video Moment Retrieval
abstract
Video moment retrieval targets at retrieving a golden moment in a video for a given natural language query. The main challenges of this task include 1) the requirement of accurately localizing (i.e., the start time and the end time of) the relevant moment in an untrimmed video stream, and 2) bridging the semantic gap between textual query and video contents. To tackle those problems, early approaches adopt the sliding window or uniform sampling to collect video clips first and then match each clip with the query to identify relevant clips. Obviously, these strategies are time-consuming and often lead to unsatisfied accuracy in localization due to the unpredictable length of the golden moment. To avoid the limitations, researchers recently attempt to directly predict the relevant moment boundaries without the requirement to generate video clips first. One mainstream approach is to generate a multimodal feature vector for the target query and video frames (e.g., concatenation) and then use a regression approach upon the multimodal feature vector for boundary detection. Although some progress has been achieved by this approach, we argue that those methods have not well captured the cross-modal interactions between the query and video frames. In this paper, we propose an Attentive Cross-modal Relevance Matching (ACRM) model which predicts the temporal boundaries based on an interaction modeling between two modalities. In addition, an attention module is introduced to automatically assign higher weights to query words with richer semantic cues, which are considered to be more important for finding relevant video contents. Another contribution is that we propose an additional predictor to utilize the internal frames in the model training to improve the localization accuracy. Extensive experiments on two public datasets TACoS and Charades-STA demonstrate the superiority of our method over several state-of-the-art methods. Ablation studies have been also conducted to examine the effectiveness of different modules in our ACRM model.
Haoyu Tang 0002, Jihua Zhu, Meng Liu 0006, Zan Gao 0001, Zhiyong Cheng 0001
IEEE Trans. Multim.1
2020 Label Enhancement with Sample Correlations via Low-Rank Representation
abstract
Compared with single-label and multi-label annotations, label distribution describes the instance by multiple labels with different intensities and accommodates to more-general conditions. Nevertheless, label distribution learning is unavailable in many real-world applications because most existing datasets merely provide logical labels. To handle this problem, a novel label enhancement method, Label Enhancement with Sample Correlations via low-rank representation, is proposed in this paper. Unlike most existing methods, a low-rank representation method is employed so as to capture the global relationships of samples and predict implicit label correlation to achieve label enhancement. Extensive experiments on 14 datasets demonstrate that the algorithm accomplishes state-of-the-art results as compared to previous label enhancement baselines.
Haoyu Tang 0002, Jihua Zhu, Qinghai Zheng, Jun Wang 0024, Shanmin Pang, Zhongyu Li 0002
AAAI1
2020 Attention feature matching for weakly-supervised video relocalization
abstract
Localizing the desired video clip for a given query in an untrimmed video has been a hot research topic for multimedia understanding. Recently, a new task named video relocalization, in which the query is a video clip, has been raised. Some methods have been developed for this task, however, these methods often require dense annotations of the temporal boundaries inside long videos for training. A more practical solution is the weakly-supervised approach, which only needs the matching information between the query and video.
Haoyu Tang 0002, Jihua Zhu, Zan Gao 0001, Tao Zhuo, Zhiyong Cheng 0001
MMAsia1