EDBT 2026 Demo / reviewers in the wild / expert
Yang Yang 0121
dblp:48/450-121
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0003-2690-7180ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hierarchical Semantics Interaction for Compressed Video Action RecognitionabstractDirectly recognizing action based on compressed video shows significant advantages, such as low storage demands, efficient decoding, and fast inference speeds. Existing methods on compressed video achieve promising performance by separately modeling spatial and motion cues and directly fusing recognition results of I-frames and P-frames. However, these approaches overlook the following inherent attributes of compressed videos: 1) Temporal misalignment between I-frames and P-frames impairs the accuracy of action recognition. 2) Spatiotemporal sparsity of compressed video frames severely hinders the semantic modeling of complex actions. 3) Semantic discrepancy between I-frames and P-frames, which capture appearance and motion information respectively, leads to suboptimal performance when they are fused directly. To address these challenges, we propose a Hierarchical Semantics Interaction Network (HSINet) that ensures refined semantic modeling of compressed video through alignment, interaction, and calibration within a hierarchical fusion framework. Specifically, to resolve temporal misalignment, we propose an efficient cross-modal temporal alignment module that fully combines the spatiotemporal information of P-frames and the spatial information of I-frames, and includes both pre-alignment and fine alignment stages. To mitigate semantic degradation caused by sparsity, we propose a cross-modal semantics interaction module to provide the multi-scale semantics interaction between spatial and temporal representation learning and enhance representations’ spatiotemporal awareness. To calibrate the semantic imbalance between I-frames and P-frames, we propose a modal imbalance calibration module that optimizing directional differences via cosine similarity in hyperspherical space. Experiments on HMDB-51, UCF-101, and Kinetics-400 benchmarks, demonstrate the effectiveness of hierarchical semantic interaction for compressed video action recognition. Jinxin Guo, Yang Yang 0121, Huaiwen Zhang, Shengsheng Qian, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | VCSA: Video Copy Localization Via Single Frame AnnotationabstractVideo Copy Localization (VCL) aims to identify all copied segments within untrimmed video pairs. Fully supervised methods, which require annotating the boundaries of copied segments, are labor-intensive and susceptible to distraction from non-copied segments. To address this problem, we propose a more efficient annotation paradigm called "single frame supervision", which only requires two randomly selected timestamps within the copied segments. With single frame supervision, we introduce a method called VCSA. This method combines spatial-temporal perception and fusion modules to extract and fuse features, which are then segmented for contrastive learning. Additionally, we propose grid alignment loss to enhance the ability of model to distinguish copied from non-copied segments. Extensive experiments demonstrate that VCSA outperforms other methods within our paradigm, sometimes even reaching performance levels comparable to fully supervised methods. Shuguo Hu, Yang Yang 0121, Huaiwen Zhang |
ICASSP | 3 |
| 2025 | Active Supervised Cross-Modal RetrievalabstractSupervised Cross-Modal Retrieval (SCMR) achieves significant performance with the supervision provided by substantial label annotations of multi-modal data. However, the requirement for large annotated multi-modal datasets restricts the use of supervised cross-modal retrieval in many practical scenarios. Active Learning (AL) has been proposed to reduce labeling costs while improving performance in various label-dependent tasks, in which the most informative unlabeled samples are selected for labeling and training. Directly exploiting the existing AL methods for supervised cross-modal retrieval may not be a good idea since they only focus on the uncertainty within each modality, ignoring the inter-modality relationship within the text-image pairs. Furthermore, existing methods focus exclusively on the informativeness of data during sample selection, leading to a biased, homogenized set where selected samples often contain nearly identical semantics and are densely distributed in a region of the feature space. Persistent training with such biased data selections can disturb multi-modal representation learning and substantially degrade the retrieval performance of SCMR. In this work, we propose an Active Supervised Cross-Modal Retrieval (ASCMR) framework, which effectively identifies informative multi-modal samples and generates unbiased sample selections. In particular, we propose a probabilistic multi-modal informativeness estimation that captures both the intra-modality and inter-modality uncertainty of multi-modal pairs within a unified representation. To ensure unbiased sample selection, we introduce a density-aware budget allocation strategy that constrains the active learning objective of maximizing the informativeness of selection with a novel semantic density regularization term. The proposed methods are evaluated on three widely used benchmark datasets, MS-COCO, NUS-WIDE, and MIRFlickr, demonstrating our effectiveness in significantly reducing the annotation cost while outperforming other baselines of active learning strategies. We could achieve over 95% of the fully supervised model's performance by only utilizing 6%, 3%, and 4% active selected samples for MS-COCO, NUS-WIDE, and MIRFlickr, respectively. Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | 3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource ApplicationsabstractTarget speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-time applications, espetially on resource-constrained platforms. In this paper, we address the TSE task using microphone array and introduce a novel three-stage solution that systematically decouples the process: First, a neural network is trained to estimate the direction of the target speaker. Second, with the direction determined, the Generalized Sidelobe Canceller (GSC) is used to extract the target speech. Third, an Inplace Convolutional Recurrent Neural Network (ICRN) acts as a denoising post-processor, refining the GSC output to yield the final separated speech. Our approach delivers superior performance while drastically reducing computational load, setting a new standard for efficient real-time target speaker extraction. Shulin He, Hao Li 0046, Yang Yang 0121, Fei Chen 0011, Xueliang Zhang 0001 |
ICASSP | 4 |
| 2024 | Hierarchical Speaker Representation for Target Speaker ExtractionabstractTarget speaker extraction aims to isolate a specific speaker’s voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and integrate them into the separation network to separate the voice of the target speaker. However, the representation of the speaker embedding is too simplistic, often being merely a 1×1024 vector. This dense information makes it difficult for the separation network to harness effectively. To address this limitation, we introduce a pioneering methodology called Hierarchical Representation (HR) that seamlessly fuses anchor data across granular and overarching 5 layers of the separation network, enhancing the precision of target extraction. HR amplifies the efficacy of anchors to improve target speaker isolation. On the Libri-2talker dataset, HR substantially outperforms state-of-the-art time-frequency domain techniques. Further demonstrating HR’s capabilities, we achieved first place in the prestigious ICASSP 2023 Deep Noise Suppression Challenge. The proposed HR methodology shows great promise for advancing target speaker extraction through enhanced anchor utilization. Shulin He, Huaiwen Zhang, Wei Rao 0002, Kanghao Zhang, Yukai Jv, Yang Yang 0121, Xueliang Zhang 0001 |
ICASSP | 6 |
| 2024 | Multi-Instance Multi-Label Learning for Text-motion RetrievalabstractText-motion retrieval (TMR) is a significant cross-modal task that retrieves motion sequences semantically similar to a given query text. Existing TMR methods primarily utilize single embeddings to represent and align text and motion sequences. However, real-world motion sequences typically contain multiple atomic motions with complex semantics, which is hard to precisely capture by single embeddings. Additionally, the common co-occurring and coupling of atomic motions further post significant challenges in effective modeling and aligning text and motion sequences. In this paper, we regard TMR as a Multi-Instance Multi-Label (MIML) learning problem, where the motion sequence is viewed as a bag of atomic motions and the text is the bag of corresponding phrases. To address the MIML problem, we propose a novel Multi-Granularity Semantics Interaction (MGSI) approach, which effectively captures and aligns the semantics of text and motion sequences across various levels. Specifically, the MGSI approach initially decomposes both the query and motion sequences into three hierarchical levels: token, instance, and bag. Then, we utilize graph neural networks to explicitly model their semantics correlation and perform semantics interaction at these respective levels, precisely capturing the semantics at multiple granularities. To identify and model co-occurring atomic motions, we measure the frame-wise semantic consistency between motions and then fuse and interact the accordant ones to refine their representations. Finally, we exploit token, instance, and bag-wise semantics interaction to comprehensively align text and motion sequence. We evaluated our methods on two widely-used benchmark datasets, HumanML3D and KIT-ML. The proposed method achieves significant improvements, outperforming the state-of-the-art with a 23.09% increase in Rsum on HumanML3D and a 21.84% increase on KIT-ML. Yang Yang 0121, Haoyu Shi, Huaiwen Zhang |
ACM Multimedia | 1 |
| 2024 | Hierarchical Semantics Alignment for 3D Human Motion RetrievalabstractText to 3D human Motion Retrieval (TMR) is a challenging task in information retrieval, aiming to query relevant motion sequences with the natural language description. The conventional approach for TMR is to represent the data instances as point embeddings for alignment. However, in real-world scenarios, multiple motions often co-occur and superimpose on a single avatar. Simply aggregating text and motion sequences into a single global embedding may be inadequate for capturing the intricate semantics of superimposing motions. In addition, most of the motion variations occur locally and subtly, which further presents considerable challenges in precisely aligning motion sequences with their corresponding text. To address the aforementioned challenges, we propose a novel Hierarchical Semantics Alignment (HSA) framework for text-to-3D human motion retrieval. Beyond global alignment, we propose the Probabilistic-based Distribution Alignment (PDA) and a Descriptors-based Fine-grained Alignment (DFA) to achieve precise semantic matching. Specifically, the PDA encodes the text and motion sequences into multidimensional probabilistic distributions, effectively capturing the semantics of superimposing motions. By optimizing the problem of probabilistic distribution alignment, PDA achieves a precise match between superimposing motions and their corresponding text. The DFA first adopts a fine-grained feature gating by selectively filtering to the significant and representative local representations and meanwhile excluding the interferences of meaningless features. Then we adaptively assign local representations from text and motion into a set of cross-modal local aggregated descriptors, enabling local comparison and interaction between fine-grained text and motion features. Extensive experiments on two widely used benchmark datasets, HumanML3D and KIT-ML, demonstrate the effectiveness of the proposed method. It significantly outperforms existing state-of-the-art retrieval methods, achieving Rsum improvements of 24.74% on HumanML3D and 23.08% on KIT-ML. Yang Yang 0121, Haoyu Shi, Huaiwen Zhang |
SIGIR | 1 |
| 2024 | T3RD: Test-Time Training for Rumor Detection on Social MediaabstractWith the increasing number of news uploaded to the internet daily, rumor detection has garnered significant attention in recent years. Existing rumor detection methods excel on familiar topics with sufficient training data (high resource) collected from the same domain. However, when facing emergent events or rumors propagated in different languages, the performance of these models is significantly degraded, due to the lack of training data and prior knowledge (low resource). To tackle this challenge, we introduce the Test-Time Training for Rumor Detection (T^3RD) to enhance the performance of rumor detection models on low-resource datasets. Specifically, we introduce self-supervised learning (SSL) as an auxiliary task in the test-time training. It consists of global and local contrastive learning, in which the global contrastive learning focuses on obtaining invariant graph representations and the local one focuses on acquiring invariant node representations. We employ the auxiliary SSL tasks for both the training and test-time training phases to mine the intrinsic traits of test samples and further calibrate the trained model for these test samples. To mitigate the risk of distribution distortion in test-time training, we introduce feature alignment constraints aimed at achieving a balanced synergy between the knowledge derived from the training set and the test samples. The experiments conducted on the two widely used cross-domain datasets demonstrate that the proposed model achieves a new state-of-the-art in performance. Our code is available at https://github.com/social-rumors/T3RD. Huaiwen Zhang, Xinxin Liu 0016, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu |
WWW | 4 |
| 2023 | C2ST: Cross-modal Contextualized Sequence Transduction for Continuous Sign Language RecognitionabstractContinuous Sign Language Recognition (CSLR) aims to transcribe the signs of an untrimmed video into written words or glosses. The mainstream framework for CSLR consists of a spatial module for visual representation learning, a temporal module aggregating the local and global temporal information of frame sequence, and the connectionist temporal classification (CTC) loss, which aligns video features with gloss sequence. Unfortunately, the language prior implicit in the gloss sequence is ignored throughout the modeling process. Furthermore, the contextualization of glosses is further ignored in alignment learning, as CTC makes an independence assumption between glosses. In this paper, we propose a Cross-modal Contextualized Sequence Transduction (C2ST) for CSLR, which effectively incorporates the knowledge of gloss sequence into the process of video representation learning and sequence transduction. Specifically, we introduce a cross-modal context learning framework for CSLR, in which the linguistic features of gloss sequences are extracted by a language model, and recurrently integrate with visual features for video modelling. Moreover, we introduce the contextualized sequence transduction loss that incorporates the contextual information of gloss sequences in label prediction, without making any independence assumptions between the glosses. Our method sets the new state of the art on three widely used large-scale sign language recognition datasets: Phoenix-2014, Phoenix-2014-T, and CSL-Daily. On CSL-Daily, our approach achieves an absolute gain of 4.9% WER compared to the best published results. Huaiwen Zhang, Zihang Guo, Yang Yang 0121, De Hu |
ICCV | 3 |
| 2023 | C2MR: Continual Cross-Modal Retrieval for Streaming Multi-modal DataabstractMassive numbers of new images are uploaded to the internet every day. However, existing cross-modal retrieval (CMR) approaches struggle to accommodate this continuously growing data. The prevalent practice involves periodically retraining or fine-tuning a new model based on the accumulated data, which in turn invalidates billions of indexed features extracted by the previous model and incurs another substantial computational cost to extract new features for the entire data archive. Is it possible to develop a retrieval model that effectively captures the knowledge of upcoming sessions while preserving the discriminative power of features extracted in previous sessions? In this paper, we propose an online continual learning setup, OC-CMR, to formalize the data-incremental growth challenge faced by cross-modal retrieval systems. It consists of two key settings: 1) Similar to the real-world scenarios, the streaming multi-modal data arrives once per session; 2) Consider the computational costs, each instance of archived data has its feature extracted only once and by its corresponding model in its session. Based on our OC-CMR, we perform in-depth evaluations of state-of-the-art cross-modal retrieval methods and observe that they suffer from representational shift and collapse due to the catastrophic forgetting. To address this issue, we propose the Continual Cross-Modal Retrieval (C2MR) approach, which learns a shared common space not only across modalities but also sessions and maintains relationships between samples from distinct sessions via cross-modal relational coherence and semantic representation coordination. We construct two new benchmarks by adapting MS-COCO and Flickr30K datasets to the OC-CMR setting, providing a more challenging evaluation framework for CMR tasks. Experimental results demonstrate that our method effectively alleviates forgetting and significantly outperforms combinations of previous arts in cross-modal retrieval and continual learning. Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu |
ACM Multimedia | 2 |
| 2023 | Debiased Video-Text Retrieval via Soft Positive Sample CalibrationabstractWith the emergence of enormous videos on various video apps, semantic video-text retrieval has become a critical task for improving the user experience. The primary paradigm for video-text retrieval learns the semantic video-text representations in a common space by pulling the positive samples close to the query and pushing the negative samples away. However, in practice, the video-text datasets contain only the annotations of positive samples. The negative samples are randomly drawn from the entire dataset. There may exist soft positive samples, which are sampled as negatives but share the same semantics as positive samples. Indiscriminately enforcing the model to push all the negative samples away from the query leads to inaccurate supervision and then misleads the video-text feature representation learning. In this paper, we introduce debiased video-text retrieval objectives that calibrate the punishment of soft positive samples. In particular, we propose a novel uncertainty measure framework to estimate the credibility of negative samples for each instance. Then, the reliability of negative samples is used to find the soft positive samples and rescale their contribution within video-text retrieval losses, including triplet loss and contrastive loss. Experimental results on five widely used datasets demonstrate that our debiased video-text retrieval objectives achieve significant performance improvements and establish a new state-of-the-art. Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Robust Video-Text Retrieval Via Noisy Pair CalibrationabstractVideo-text retrieval is a fundamental task in managing the emerging massive amounts of video data. The main challenge focuses on learning a common representation space for videos and queries where the similarity measurement can reflect the semantic closeness. However, existing video-text retrieval models may suffer from the following noise in the common space learning procedure: First, the video-text correspondences in positive pairs may not be exact matches. The crowdsourcing annotation for existing datasets leads to inevitable tagging noise for non-expert annotators. Second, the learning of video-text representation is based on the negative samples randomly sampled. Instances that are semantically similar to the query may be incorrectly categorized as negative samples. To alleviate the adverse impact of these noisy pairs, we propose a novel robust video-text retrieval method that protects the model from noisy positive and negative pairs by identifying and calibrating noisy pairs with their uncertainty score. In particular, we propose a noisy pair identifier, which divides the training dataset into noisy and clean subsets based on the estimated uncertainty of each pair. Then, with the help of uncertainties, we calibrate the two types of noisy pairs with an adaptive margin triplet loss and a weighted triplet loss function, respectively. To verify the effectiveness of our methods, we conduct extensive experiments on three widely used datasets. Experimental results show that the proposed robust video-text retrieval methods successfully identify and calibrate the noisy pairs and improve retrieval performance. Huaiwen Zhang, Yang Yang 0121, Fan Qi, Shengsheng Qian, Changsheng Xu |
IEEE Trans. Multim. | 2 |
| 2022 | Alleviating the Loss-Metric Mismatch in Supervised Single-Channel Speech EnhancementabstractIn this paper, we study the loss-metric mismatch problem of supervised single-channel speech enhancement system. Most of the existing speech enhancement systems achieve unsatisfying performance since their empirically selected loss functions have semantic gaps with the non-differentiable evaluation metrics, a.k.a., the loss-metric mismatch problem. In this work, we propose a simple yet efficient method to generate suitable loss functions for the real front-end speech enhancement scenarios to alleviate the loss-metric mismatch problem. Specifically, we adopt the function smoothing technique and approximate the non-differentiable evaluation metrics by a set of basis functions and their linear combination. Experimental results demonstrate that the loss function generated by our method helps the speech enhancement system achieve remarkable performance in most evaluation metrics than the traditional empirically selected ones. Yang Yang 0121, Hui Zhang 0031, Xueliang Zhang 0001, Huaiwen Zhang |
ICASSP | 1 |