Shuo Wang 0008

dblp:63/1591-8 · DBLP profile ↗
← Back
48ranked-venue papers
13as first author
42since 2021 · last 2026
0000-0002-4881-9344ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 35 · 11 first-author · 30 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 17 since 2021Computer networks · 5 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Accelerating Controllable Generation via Hybrid-grained Cache
abstract
Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation computational requirements, resulting in generally low generation efficiency. To address this issue, we propose a Hybrid-Grained Cache (HGC) approach that reduces computational overhead by adopting cache strategies with different granularities at different computational stages. Specifically, (1) we use a coarse-grained cache (block-level) based on feature reuse to dynamically bypass redundant computations in encoder-decoder blocks between each step of model reasoning. (2) We design a fine-grained cache (prompt-level) that acts within a module, where the fine-grained cache reuses cross-attention maps within consecutive reasoning steps and extends them to the corresponding module computations of adjacent steps. These caches of different granularities can be seamlessly integrated into each computational link of the controllable generation process. We verify the effectiveness of HGC on four benchmark datasets, especially its advantages in balancing generation efficiency and visual quality. For example, on the COCO-Stuff segmentation benchmark, our HGC significantly reduces the computational cost (MACs) by 63% (from 18.22T → 6.70T↓), while keeping the loss of semantic fidelity (quantized performance degradation) within 1.5%.
Huixia Ben, Shuo Wang 0008, Jinda Lu, Junxiang Qiu, Shengeng Tang, Yanbin Hao
AAAI3
2026 Hierarchical Semantic Alignment for Image Clustering
abstract
Image clustering is a classic problem in computer vision, which categorizes images into different groups. Recent studies utilize nouns as external semantic knowledge to improve clustering performance. However, these methods often overlook the inherent ambiguity of nouns, which can distort semantic representations and degrade clustering quality. To address this issue, we propose a hierarChical semAntic alignmEnt method for image clustering, dubbed CAE, which improves clustering performance in a training-free manner. In our approach, we incorporate two complementary types of textual semantics: caption-level descriptions, which convey fine-grained attributes of image content, and noun-level concepts, which represent high-level object categories. We first select relevant nouns from WordNet and descriptions from caption datasets to construct a semantic space aligned with image features. Then, we design a residual attention mechanism to further enhance the discriminability of this space. Finally, we combine the enhanced semantic and image features to perform clustering. Extensive experiments across 8 datasets demonstrate the effectiveness of our method, notably surpassing the state-of-the-art training-free approach with a 4.2% improvement in accuracy and a 2.9% improvement in adjusted rand index (ARI) on the ImageNet-1K dataset.
Beier Zhu, Junfeng Fang, Shuo Wang 0008, Kesen Zhao, Hanwang Zhang
AAAI5
2026 Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
abstract
Xingyu Zhu, Junfeng Fang, Shuo Wang, Beier Zhu, Zhicai Wang, Yonghui Yang, Xiangnan He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junfeng Fang, Shuo Wang 0008, Beier Zhu, Zhicai Wang, Yonghui Yang 0001, Xiangnan He 0001
ACL (1)3
2026 Hybrid Granularity Distribution Estimation for Few-Shot Learning: Statistics Transfer From Categories and Instances
abstract
Distribution estimation is a pivotal strategy in few-shot learning (FSL) to mitigate data scarcity by sampling from estimated distributions, utilizing statistical properties (mean and variance) transferred from related base categories. However, category-level estimation alone often fails to generate representative samples due to significant dissimilarities between base and novel categories, leading to suboptimal performance. To address this limitation, we propose Hybrid Granularity Distribution Estimation (HGDE), which integrates both coarse-grained category-level statistics and fine-grained instance-level statistics. By leveraging instance statistics from the nearest base samples, HGDE enhances the characterization of novel categories, capturing subtle features that category-level estimation overlooks. These statistics are fused through linear interpolation to form a robust distribution for novel categories, ensuring both diversity and representativeness in generated samples. Additionally, HGDE employs refined estimation techniques, such as weighted summation for mean calculation and principal component retention for covariance, to further improve accuracy. Empirical evaluations on four FSL benchmarks, including Mini-ImageNet, Tiered-ImageNet, CUB and CIFAR-FS, demonstrate that HGDE offers effective distribution estimation capabilities and leads to notable accuracy gains, with improvements of more than 1.8% in 1-shot tasks on CUB. These results highlight HGDE's ability to balance mean precision and variance diversity, making it a versatile and effective solution for FSL.
Shuo Wang 0008, Tianyu Qi, Yanbin Hao, Beier Zhu, Hanwang Zhang, Meng Wang 0001
IEEE Trans. Image Process.1
2025 Linguistics-Vision Monotonic Consistent Network for Sign Language Production
abstract
Sign Language Production (SLP) aims to generate sign videos corresponding to spoken language sentences, where the conversion of sign Glosses to Poses (G2P) is the key step. Due to the cross-modal semantic gap and the lack of word-action correspondence labels for strong supervision alignment, the SLP suffers huge challenges in linguistics-vision consistency. In this work, we propose a Transformer-based Linguistics-Vision Monotonic Consistent Network (LVMCN) for SLP, which constrains fine-grained cross-modal monotonic alignment and coarse-grained multimodal semantic consistency in language-visual cues through Cross-modal Semantic Aligner (CSA) and Multimodal Semantic Comparator (MSC). In the CSA, we constrain the implicit alignment between corresponding gloss and pose sequences by computing the cosine similarity association matrix between cross-modal feature sequences (i.e., the order consistency of fine-grained sign glosses and actions). As for MSC, we construct multimodal triplets based on paired and unpaired samples in batch data. By pulling closer the corresponding text-visual pairs and pushing apart the non-corresponding text-visual pairs, we constrain the semantic co-occurrence degree between corresponding gloss and pose sequences (i.e., the semantic consistency of coarse-grained textual sentences and sign videos). Extensive experiments on the popular PHOENIX14T benchmark show that the LVMCN outperforms the state-of-the-art.
Shengeng Tang, Peipei Song, Shuo Wang 0008, Dan Guo 0001, Richang Hong
ICASSP4
2025 Accelerating Diffusion Transformer via Gradient-Optimized Cache
Junxiang Qiu, Shuo Wang 0008, Jinda Lu, Kezhou Chen, Yanbin Hao
ICCV3
2025 Dynamic Multimodal Prototype Learning in Vision-Language Models
abstract
With the increasing attention to pre-trained vision-language models (VLMs), \eg, CLIP, substantial efforts have been devoted to many downstream tasks, especially in test-time adaptation (TTA). However, previous works focus on learning prototypes only in the textual modality while overlooking the ambiguous semantics in class names. These ambiguities lead to textual prototypes that are insufficient to capture visual concepts, resulting in limited performance. To address this issue, we introduce \textbf{ProtoMM}, a training-free framework that constructs multimodal prototypes to adapt VLMs during the test time. By viewing the prototype as a discrete distribution over the textual descriptions and visual particles, ProtoMM has the ability to combine the multimodal features for comprehensive prototype learning. More importantly, the visual particles are dynamically updated as the testing stream flows. This allows our multimodal prototypes to continually learn from the data, enhancing their generalizability in unseen scenarios. In addition, we quantify the importance of the prototypes and test images by formulating their semantic distance as an optimal transport problem. Extensive experiments on 15 zero-shot benchmarks demonstrate the effectiveness of our method, achieving a 1.03\% average accuracy improvement over state-of-the-art methods on ImageNet and its variant datasets.
Shuo Wang 0008, Beier Zhu, Miaoge Li, Junfeng Fang, Zhicai Wang, Dongsheng Wang 0003, Hanwang Zhang
ICCV2
2025 DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
abstract
Direct Preference Optimization (DPO) has shown effectiveness in aligning multi-modal large language models (MLLM) with human preferences. However, existing methods exhibit an imbalanced responsiveness to the data of varying hardness, tending to overfit on the easy-to-distinguish data while underfitting on the hard-to-distinguish data. In this paper, we propose Data- and Model-aware DPO (DAMA) to dynamically adjust the optimization process from two key aspects: (1) a data-aware strategy that incorporates data hardness, and (2) a model-aware strategy that integrates real-time model responses. By combining the two strategies, DAMA enables the model to effectively adapt to data with varying levels of hardness. Extensive experiments on five benchmarks demonstrate that DAMA not only significantly enhances the trustworthiness, but also improves the effectiveness over general tasks. For instance, on the Object HalBench, our DAMA-7B reduces response-level and mentioned-level hallucination by 90.0% and 95.3%, respectively, surpassing the performance of GPT-4V.
Jinda Lu, Junkang Wu, Jinghan Li, Xiaojun Jia, Shuo Wang 0008, Yifan Zhang 0004, Junfeng Fang, Xiang Wang 0010, Xiangnan He 0001
ICML5
2025 Accelerating Diffusion Transformer via Error-Optimized Cache
abstract
Diffusion Transformer (DiT) is a crucial method for content generation. However, it needs a lot of time to sample. Many studies have attempted to use caching to reduce the time consumption of sampling. Existing caching methods accelerate generation by reusing DiT features from the previous time step and skipping calculations in the next, but they tend to locate and cache low-error modules without focusing on reducing caching-induced errors, resulting in a sharp decline in generated content quality when increasing caching intensity. To solve this problem, we propose the Error-Optimized Cache (EOC). This method introduces three key improvements: (1) Prior knowledge extraction: Extract and process the caching differences; (2) A judgment method for cache optimization: Determine whether certain caching steps need to be optimized; (3) Cache optimization: reduce caching errors. Experiments show that this algorithm significantly reduces the error accumulation caused by caching, especially excessive caching. On the ImageNet dataset, without substantially increasing the computational load, this method improves the FID↓ of the generated images when the rule-based model FORA has a caching level of 75%, 50%, and 25%, and the training-based model Learning-to-cache has a caching level of 22%. Specifically, the FID↓ values change from 30.454 to 21.690 (28.8%), from 6.857 to 5.821 (15.1%), from 3.870 to 3.692 (4.6%), and from 3.539 to 3.451 (2.5%) respectively. Code is available at https://github.com/qiujx0520/EOC_MM2025.git.
Junxiang Qiu, Shuo Wang 0008, Jinda Lu, Houcheng Jiang, Yanbin Hao
ACM Multimedia2
2025 Mixture of Multimodal Adapters for Sentiment Analysis
abstract
Kezhou Chen, Shuo Wang, Huixia Ben, Shengeng Tang, Yanbin Hao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kezhou Chen, Shuo Wang 0008, Huixia Ben, Shengeng Tang, Yanbin Hao
NAACL (Long Papers)2
2025 Enhancing CLIP Robustness via Cross-Modality Alignment
abstract
Vision-language models (VLMs) such as CLIP demonstrate strong generalization in zero-shot classification but remain highly vulnerable to adversarial perturbations. Existing methods primarily focus on adversarial fine-tuning or prompt optimization, they often overlook the gaps in CLIP’s encoded features, which is shown as the text and image features lie far apart from each other. This misalignment is significantly amplified under adversarial perturbations, leading to severe degradation in classification performance. To address this problem, we propose **C**r**O**ss-moda**L**ity **A**lignment, dubbed **COLA**, an optimal transport-based framework that explicitly addresses adversarial misalignment by restoring both global image-text alignment and local structural consistency in the feature space. (1) COLA first projects adversarial image embeddings onto a subspace spanned by class text features, effectively filtering out non-semantic distortions while preserving discriminative information. (2) It then models images and texts as discrete distributions over multiple augmented views and refines their alignment via OT, with the subspace projection seamlessly integrated into the cost computation. This design ensures stable cross-modal alignment even under adversarial conditions. COLA is training-free and compatible with existing fine-tuned models. Extensive evaluations across 14 zero-shot classification benchmarks demonstrate the effectiveness of COLA, especially with an average improvement of 6.7% on ImageNet and its variants under PGD adversarial attacks, while maintaining high accuracy on clean samples.
Beier Zhu, Shuo Wang 0008, Kesen Zhao, Hanwang Zhang
NeurIPS3
2025 Inversed Pyramid Network with Spatial-adapted and Task-oriented Tuning for few-shot learning
Duorui Wang, Shihao Bai, Shuo Wang 0008, Yajun Gao, Yuqing Ma, Xianglong Liu 0001
Pattern Recognit.4
2025 Video Corpus Moment Retrieval With Query-Specific Context Learning and Progressive Localization
abstract
Video corpus moment retrieval (VCMR) aims to retrieve a moment from a large corpus of untrimmed videos corresponding to a given language query. However, existing methods often fall short due to their reliance on simple cross-modal attention mechanisms and one-stop localization, which fail to handle the complex multimodal information and large search space effectively. To address these challenges, we propose a novel VCMR method with Query-specific Context Learning and Progressive Localization (QCLPL). First, we construct query-specific multimodal contexts that capture complementary and consistent semantics across subtitles and frames, ensuring informative and efficient context building. We further introduce a semantic contrastive loss to refine these multimodal contexts, filtering out query-irrelevant information. Additionally, we introduce a progressive localization strategy that transforms the moment localization task into a two-stage process. By classifying frames into foreground and background regions, we present a simplified binary classification problem before boundary prediction, constrained by a region-aware loss. This progressive approach leverages region priors to improve subsequent moment localization. Extensive experiments on the TVR and DiDeMo datasets demonstrate that our method significantly outperforms existing approaches, setting a new state of the art for VCMR.
Peipei Song, Zhangling Duan, Shuo Wang 0008, Xiaojun Chang, Xun Yang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2025 Symmetric Hallucination With Knowledge Transfer for Few-Shot Learning
abstract
Data hallucination or augmentation is a straightforward solution for few-shot learning (FSL), where FSL is proposed to classify a novel object under limited training samples. Common hallucination strategies use visual or textual knowledge to simulate the distribution of a given novel category and generate more samples for training. However, the diversity and capacity of generated samples through these techniques can be insufficient when the knowledge domain of the novel category is narrow. Therefore, the performance improvement of the classifier is limited. To address this issue, we propose a Symmetric data hallucination strategy with Knowledge Transfer (SHKT) that interacts with multi-modal knowledge in both visual and textual spaces. Specifically, we first calculate the relations based on semantic knowledge and select the most related categories of a given novel category for hallucination. Second, we design two parameter-free data hallucination strategies to enrich the training samples by mixing the given and selected samples in both visual and textual spaces. The generated visual and textual samples improve the visual representation and enrich the textual supervision, respectively. Finally, we connect the visual and textual knowledge through transfer calculation, which not only exchanges content from different modalities but also constrains the distribution of the generated samples during the training. We apply our method to four benchmark datasets and achieve state-of-the-art performance in all experiments. Specifically, compared to the baseline on the Mini-ImageNet dataset, it achieves 12.84% and 3.46% accuracy improvements for 1 and 5 support training samples, respectively.
Shuo Wang 0008, Xinyu Zhang 0017, Meng Wang 0001, Xiangnan He 0001
IEEE Trans. Multim.1
2025 CVLP-NaVD: Contrastive Visual-language Pre-training Models for Non-annotated Visual Description
abstract
Non-annotated visual description (NaVD) aims to describe generic visuals without human-annotated pairwise data. The generic visuals refer to images and videos. Existing works mainly focus on one specific visual modality, i.e., image or video. In this article, we propose a new framework for this task, which can directly be applied to both image and video with the pipeline unchanged. Essentially, it is a unified framework that flexibly adapts to images and videos. Recently, contrastive visual-language pre-training models (CVLPs) have experienced rapid development, demonstrating powerful abilities to align vision and language. To continuously leverage advanced CVLPs, our framework is designed to work well with general CVLPs. It can easily use image-language CVLPs for image input and switch to video-language CVLPs for video input. Specifically, we propose a CVLP-based framework for NaVD, named CVLP-NaVD. It follows the paradigm of adversarial learning, containing a generator and a discriminator. The generator takes an image or a video as input and produces a corresponding language description, while the discriminator evaluates the generated sentence for its naturalness in human-like language. Apart from the naturalness, CVLPs play a crucial role in enhancing the alignment between visual and language signals during generation. Particularly, we explore three rewarding strategies to compute the alignment score, including directly calculating cosine similarity (i.e., VL-cross), projecting visual embeddings into the textual domain (i.e., VL-project), and their combination (i.e., VL-mix). The three strategies are fully examined in different scenarios. Finally, we conduct extensive experiments with various unpaired and unsupervised setups in both image and video captioning tasks. The experimental results demonstrate that our CVLP-NaVD outperforms the state-of-the-art methods significantly.
Yanbin Hao, Jiarui Yu, Bin Zhu 0006, Shuo Wang 0008, Tong Xu 0001
ACM Trans. Multim. Comput. Commun. Appl.5
2025 Gloss-driven Conditional Diffusion Models for Sign Language Production
abstract
Sign Language Production (SLP) aims to convert text or audio sentences into sign language videos corresponding to their semantics, which is challenging due to the diversity and complexity of sign languages, and cross-modal semantic mapping issues. In this work, we propose a Gloss-driven Conditional Diffusion Model (GCDM) for SLP. The core of the GCDM is a diffusion model architecture, in which the sign gloss sequence is encoded by a Transformer-based encoder and input into the diffusion model as a semantic prior condition. In the process of sign pose generation, the textual semantic priors carried in the encoded gloss features are integrated into the embedded Gaussian noise via cross-attention. Subsequently, the model converts the fused features into sign language pose sequences through T-round denoising steps. During the training process, the model uses the ground-truth labels of sign poses as the starting point, generates Gaussian noise through T rounds of noise, and then performs T rounds of denoising to approximate the real sign language gestures. The entire process is constrained by the MAE loss function to ensure that the generated sign language gestures are as close as possible to the real labels. In the inference phase, the model directly randomly samples a set of Gaussian noise, generates multiple sign language gesture sequence hypotheses under the guidance of the gloss sequence, and outputs a high-confidence sign language gesture video by averaging multiple hypotheses. Experimental results on the Phoenix2014T dataset show that the proposed GCDM method achieves competitiveness in both quantitative performance and qualitative visualization.
Shengeng Tang, Feng Xue 0002, Jingjing Wu 0001, Shuo Wang 0008, Richang Hong
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Interventional Feature Generation for Few-shot Learning
abstract
Few-shot learning (FSL) aims to classify a novel object into a specific category under limited training samples. This is a challenging task since (1) the features expressed by pre-trained knowledge introduce perceived bias and then constrain the classification space, and (2) the use of general hallucination techniques based on global features fails to escape the limited classification space, resulting in sub-optimal improvements. To solve these issues, this article proposes an interventional feature generation (IFG) method. Specifically, we first use the relations of the categories or instances as interventional operations to implicitly constrain the feature representations (pre-trained knowledge) into different classification subsets. Then, we employ a parameter-free feature generation strategy to enrich each subset’s training samples of the support category. In other words, IFG provides a multi-subsets learning strategy to reduce the influence of perceived bias, enrich the diversity of generated features, and improve the robustness of the few-shot classifier. We apply our method to four benchmark datasets and observe state-of-the-art performance across all experiments. Specifically, compared to the baseline on the Mini-ImageNet dataset, our approach yields accuracy improvements of 6.03% and 3.46% for 1 and 5 support training samples, respectively. Furthermore, the proposed interventional feature generation technique can improve classifier performance in other FSL methods, demonstrating its versatility and potential for broader applications. The code is available at https://github.com/ShuoWangCS/IFG-FSL/ .
Shuo Wang 0008, Jinda Lu, Huixia Ben, Yanbin Hao, Xingyu Gao 0001, Meng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Boosting Few-Shot Learning via Attentive Feature Regularization
abstract
Few-shot learning (FSL) based on manifold regularization aims to improve the recognition capacity of novel objects with limited training samples by mixing two samples from different categories with a blending factor. However, this mixing operation weakens the feature representation due to the linear interpolation and the overlooking of the importance of specific channels. To solve these issues, this paper proposes attentive feature regularization (AFR) which aims to improve the feature representativeness and discriminability. In our approach, we first calculate the relations between different categories of semantic labels to pick out the related features used for regularization. Then, we design two attention-based calculations at both the instance and channel levels. These calculations enable the regularization procedure to focus on two crucial aspects: the feature complementarity through adaptive interpolation in related categories and the emphasis on specific feature channels. Finally, we combine these regularization strategies to significantly improve the classifier performance. Empirical studies on several popular FSL benchmarks demonstrate the effectiveness of AFR, which improves the recognition accuracy of novel categories without the need to retrain any feature extractor, especially in the 1-shot setting. Furthermore, the proposed AFR can seamlessly integrate into other FSL methods to improve classification performance.
Shuo Wang 0008, Jinda Lu, Yanbin Hao, Xiangnan He 0001
AAAI2
2024 GLCM-Adapter: Global-Local Content Matching for Few-shot CLIP Adaptation
Shuo Wang 0008, Xieenlong, Jinda Lu, Jinghan Li, Yanbin Hao
BMVC1
2024 Enhancing Recipe Retrieval with Foundation Models: A Data Augmentation Perspective
Fangzhou Song, Bin Zhu 0006, Yanbin Hao, Shuo Wang 0008
ECCV (51)4
2024 Pseudo Content Hallucination for Unpaired Image Captioning
abstract
Unpaired Image Captioning (UIC) is designed to describe an image without relying on matched vision-language training data. It is a challenging task since (1) the implicit and unpaired vision-language data nature of the training task limits the captioning model's ability to represent diverse scene representations, and (2) it is difficult for the captioning model to discern the intrinsic relationships among objects, potentially leading to misinterpretation of the image con- tent. To solve these issues, we propose pseudo content hallucination (PCH) to help the captioning model enlarge the perception of the ob- jects and capture the relations between the objects. Specifically, we select similar objects from different images as pseudo content and then hallucinate new visual content for training. This hallucinated content contains a similar scene but with a different representation, thus enriching the diversity of the training samples. Meanwhile, we utilize the relationships among these objects to improve the generated captions as a textual content hallucination and construct pseudo image-sentence pairs to refine the captioning model. These hallucinated sentences are beneficial for the captioning model as they enable the capture of additional semantics from the image, ultimately enhancing the sentence generation ability. Extensive experiments on the two benchmarks, i.e., MSCOCO, and Flickr30k, show the effectiveness of our method. The results show a significant improvement compared to the baseline in the MSCOCO dataset, with 1.5 increase in the CIDEr score.
Huixia Ben, Shuo Wang 0008, Meng Wang 0001, Richang Hong
ICMR2
2024 Selective Vision-Language Subspace Projection for Few-shot CLIP
abstract
Vision-language models such as CLIP are capable of mapping the different modality data into a unified feature space, enabling zero/few-shot inference by measuring the similarity of given images and texts. However, most existing methods overlook modality gaps in CLIP's encoded features, which is shown as the text and image features lie far apart from each other, resulting in limited classification performance. To tackle this issue, we introduce a method called Selective Vision-Language Subspace Projection (SSP), which incorporates local image features and utilizes them as a bridge to enhance the alignment between image-text pairs. Specifically, our SSP framework comprises two parallel modules: a vision projector and a language projector. Both projectors utilize local image features to span the respective subspaces for image and texts, thereby projecting the image and text features into their respective subspaces to achieve alignment. Moreover, our approach entails only training-free matrix calculations and can be seamlessly integrated into advanced CLIP-based few-shot learning frameworks. Extensive experiments on 11 datasets have demonstrated SSP's superior text-image alignment capabilities, outperforming the state-of-the-art alignment methods. The code is available at https://github.com/zhuhsingyuu/SSP
Beier Zhu, Yi Tan 0001, Shuo Wang 0008, Yanbin Hao, Hanwang Zhang
ACM Multimedia4
2024 Hierarchical Supervised Contrastive Learning for Multimodal Sentiment Analysis
Kezhou Chen, Shuo Wang 0008, Yanbin Hao
MMM (2)2
2024 Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting
abstract
Vision-language models, such as CLIP, have shown impressive generalization capacities when using appropriate text descriptions. While optimizing prompts on downstream labeled data has proven effective in improving performance, these methods entail labor costs for annotations and are limited by their quality. Additionally, since CLIP is pre-trained on highly imbalanced Web-scale data, it suffers from inherent label bias that leads to suboptimal performance. To tackle the above challenges, we propose a label-**F**ree p**ro**mpt distribution **l**earning and b**i**as **c**orrection framework, dubbed as **Frolic**, which boosts zero-shot performance without the need for labeled data. Specifically, our Frolic learns distributions over prompt prototypes to capture diverse visual representations and adaptively fuses these with the original CLIP through confidence matching. This fused model is further enhanced by correcting label bias via a label-free logit adjustment. Notably, our method is not only training-free but also circumvents the necessity for hyper-parameter tuning. Extensive experimental results across 16 datasets demonstrate the efficacy of our approach, particularly outperforming the state-of-the-art by an average of $2.6\%$ on 10 datasets with CLIP ViT-B/16 and achieving an average margin of $1.5\%$ on ImageNet and its five distribution shifts with CLIP ViT-B/16. Codes are available in [https://github.com/zhuhsingyuu/Frolic](https://github.com/zhuhsingyuu/Frolic).
Beier Zhu, Yi Tan 0001, Shuo Wang 0008, Yanbin Hao, Hanwang Zhang
NeurIPS4
2024 JPA: A Joint-Part Attention for Mitigating Overfocusing on 3D Human Pose Estimation
Dengqing Yang, Zhenhua Tang 0001, Jinmeng Wu, Shuo Wang 0008, Lechao Cheng, Yanbin Hao
PRCV (6)4
2024 MCAS-GP: Deep Learning-Empowered Middle Cerebral Artery Segmentation and Gate Proposition
abstract
With the fast development of AI technologies, deep learning is widely applied for biomedical data analytics and digital healthcare. However, there remain gaps between AI-aided diagnosis and real-world healthcare demands. For example, hemodynamic parameters of the middle cerebral artery (MCA) have significant clinical value for diagnosing adverse perinatal results. Nevertheless, the current measurement procedure is tedious for sonographers. To reduce the workload of sonographers, we propose MCAS-GP, a deep learning-empowered framework that tackles the Middle Cerebral Artery Segmentation and Gate Proposition. MCAS-GP can automatically segment the region of the MCA and detect the corresponding position of the gate in the procedure of fetal MCA Doppler assessment. In MCAS-GP, a novel learnable atrous spatial pyramid pooling (LASPP) module is designed to adaptively learn multi-scale features. We also propose a novel evaluation metric, Affiliation Index, for measuring the effectiveness of the position of the output gate. To evaluate our proposed MCAS-GP, we build a large-scale MCA dataset, collaborating with the International Peace Maternity and Child Health Hospital of China welfare institute (IPMCH). Extensive experiments on the MCA dataset and two other public surgical datasets demonstrate that MCAS-GP can achieve considerable performance improvement in both accuracy and inference time.
Rui Zhang 0087, Shuo Wang 0008, Ruhui Ma, Yang Hua 0001, Tao Song 0003, Yunyun Cao, Haibing Guan
IEEE Trans. Comput. Biol. Bioinform.2
2024 Feature Mixture on Pre-Trained Model for Few-Shot Learning
abstract
Few-shot learning (FSL) aims at recognizing a novel object under limited training samples. A robust feature extractor (backbone) can significantly improve the recognition performance of the FSL model. However, training an effective backbone is a challenging issue since 1) designing and validating structures of backbones are time-consuming and expensive processes, and 2) a backbone trained on the known (base) categories is more inclined to focus on the textures of the objects it learns, which is hard to describe the novel samples. To solve these problems, we propose a feature mixture operation on the pre-trained (fixed) features: 1) We replace a part of the values of the feature map from a novel category with the content of other feature maps to increase the generalizability and diversity of training samples, which avoids retraining a complex backbone with high computational costs. 2) We use the similarities between the features to constrain the mixture operation, which helps the classifier focus on the representations of the novel object where these representations are hidden in the features from the pre-trained backbone with biased training. Experimental studies on five benchmark datasets in both inductive and transductive settings demonstrate the effectiveness of our feature mixture (FM). Specifically, compared with the baseline on the Mini-ImageNet dataset, it achieves 3.8% and 4.2% accuracy improvements for 1 and 5 training samples, respectively. Additionally, the proposed mixture operation can be used to improve other existing FSL methods based on backbone training.
Shuo Wang 0008, Jinda Lu, Haiyang Xu 0002, Yanbin Hao, Xiangnan He 0001
IEEE Trans. Image Process.1
2023 How Can Contrastive Pre-training Benefit Audio-Visual Segmentation? A Study from Supervised and Zero-shot Perspectives
Jiarui Yu, Yanbin Hao, Jinmeng Wu, Tong Xu 0001, Shuo Wang 0008, Xiangnan He 0001
BMVC6
2023 Bi-Directional Distribution Alignment for Transductive Zero-Shot Learning
abstract
Zero-shot learning (ZSL) suffers intensely from the domain shift issue, i.e., the mismatch (or misalignment) between the true and learned data distributions for classes without training data (unseen classes). By learning additionally from unlabelled data collected for the unseen classes, transductive ZSL (TZSL) could reduce the shift but only to a certain extent. To improve TZSL, we propose a novel approach Bi-VAEGAN which strengthens the distribution alignment between the visual space and an auxiliary space. As a result, it can reduce largely the domain shift. The proposed key designs include (1) a bi-directional distribution alignment, (2) a simple but effective L2-norm based feature normalization approach, and (3) a more sophisticated unseen class prior estimation. Evaluated by four benchmark datasets, Bi-VAEGAN11Code is available at https://github.com/Zhicaiwww/Bi-VAEGAN achieves the new state of the art under both the standard and generalized TZSL settings.
Zhicai Wang, Yanbin Hao, Tingting Mu, Ouxiang Li, Shuo Wang 0008, Xiangnan He 0001
CVPR5
2023 Cross-Modal Contrastive Learning for Event Extraction
Shuo Wang 0008, Meizhi Ju, Yunyan Zhang, Yefeng Zheng 0001, Meng Wang 0001, Guilin Qi
DASFAA (3)1
2023 Semantic-based Selection, Synthesis, and Supervision for Few-shot Learning
abstract
Few-shot learning (FSL) is designed to explore the distribution of novel categories from a few samples. It is a challenging task since the classifier is usually susceptible to over-fitting when learning from limited training samples. To alleviate this phenomenon, a common solution is to achieve more training samples using a generic generation strategy in visual space. However, there are some limitations to this solution. It is because a feature extractor trained on base samples (known knowledge) tends to focus on the textures and structures of the objects it learns, which is inadequate for describing novel samples. To solve these issues, we introduce semantics and propose a Semantic-based Selection, Synthesis, and S upervision (4S) method, where semantics provide more diverse and informative supervision for recognizing novel objects. Specifically, we first utilize semantic knowledge to explore the correlation of categories in the textual space and select base categories related to the given novel category. This process can improve the efficiency of subsequent operations (synthesis and supervision). Then, we analyze the semantic knowledge to hallucinate the training samples by selectively synthesizing the contents from base and support samples. This operation not only increases the number of training samples but also takes advantage of the contents of the base categories to enhance the description of support samples. Finally, we also employ semantic knowledge as both soft and hard supervision to enrich the supervision for the fine-tuning procedure. Empirical studies on four FSL benchmarks demonstrate the effectiveness of 4S.
Jinda Lu, Shuo Wang 0008, Xinyu Zhang 0022, Yanbin Hao, Xiangnan He 0001
ACM Multimedia2
2023 A meaningful learning method for zero-shot semantic segmentation
Xianglong Liu 0001, Shihao Bai, Shan An, Shuo Wang 0008, Wei Liu 0005, Yuqing Ma
Sci. China Inf. Sci.4
2023 Perceptual Data Augmentation for Biomedical Coronary Vessel Segmentation
abstract
Sufficient annotated data is critical to the success of deep learning methods. Annotating for vessel segmentation in X-ray coronary angiograms is extremely difficult because of the small and complex structures to be processed. Although unsupervised domain adaptation methods can be utilized to alleviate the annotation burden by using data in other domains, e.g., eye fundus images, these methods cannot perform well due to the characteristic of medical images. Data augmentation can help improve the similarity of source domain and target domain in unsupervised domain adaptation tasks. Existing data augmentation methods play a limited role in improving domain adaptation performance, especially for special medical image segmentation tasks. In this paper, we propose an effective perceptual data augmentation method to improve the similarity between eye fundus images and coronary angiograms by synthesizing virtual samples. Auto Foreground Augment method is designed to search for geometric transformations that improve the similarity between foreground vessels of eye fundus images and coronary angiograms. The Haar Wavelet-Based Perceptual Similarity Index is utilized to guide the synthesis of virtual samples in foreground and background mixup. Extensive experiments show that our data augmentation method can synthesize high-quality virtual samples and thus improve the domain adaptation performance. To our best knowledge, this is the first work to apply perceptual data augmentation to vessel segmentation in coronary angiograms.
Shuo Wang 0008, Yang Hua 0001, Ruhui Ma, Tao Song 0003, Zhengui Xue, Haibing Guan
IEEE ACM Trans. Comput. Biol. Bioinform.2
2023 Boosting Hyperspectral Image Classification with Dual Hierarchical Learning
abstract
Hyperspectral image (HSI) classification aims at predicting the pixel-wise labels in an image, where there are only a few labeled pixel samples (hard labels) for training. It is a challenging task since the classification process is susceptible to over-fitting under training with limited samples. To relieve this problem, we propose a method based on dual hierarchical learning. First, we employ a connectionist hyperspectral convolution (HC) network to capture the representations of the pixels from different receptive fields. Specifically, an HC is designed to learn the correlation among adjacent pixels and is further extended to a connectionist hierarchical structure. These operations use the correlation to enhance one-pixel learning from multiple receptive fields. Second, we analyze the properties in the hyperspectral image and introduce a hierarchical pseudo label generation algorithm to enrich the supervision of the label information. Finally, we design a dual hierarchical learning strategy to help all HC layers learn from both the hard labels and the hierarchical pseudo labels. In other words, it addresses the HSI classification problem from different views. For inference, we employ two fusion strategies to find a better prediction. The experimental results on four popular HSI benchmarks, i.e., Salinas-A, IndianPines, PaviaU, and PaviaC, demonstrate the effectiveness of the proposed method. Our code is publicly available on GitHub: https://github.com/ShuoWangCS/HSI-DHL.
Shuo Wang 0008, Huixia Ben, Yanbin Hao, Xiangnan He 0001, Meng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Multi-directional Knowledge Transfer for Few-Shot Learning
abstract
Knowledge transfer-based few-shot learning (FSL) aims at improving the recognition ability of a novel object under limited training samples by transferring relevant potential knowledge from other data. Most related methods calculate such knowledge to refine the representation of a novel sample or enrich the supervision to a classifier during a transfer procedure. However, it is easy to introduce new noise during the transfer calculations since: (1) the unbalanced quantity of samples between the known (base) and the novel categories biases the contents capturing of the novel objects, and (2) the semantic gaps existing in different modalities weakens the knowledge interaction during the training.
Shuo Wang 0008, Xinyu Zhang 0022, Yanbin Hao, Chengbing Wang, Xiangnan He 0001
ACM Multimedia1
2022 Hierarchical Hourglass Convolutional Network for Efficient Video Classification
abstract
Videos naturally contain dynamic variation over the temporal axis, which will result in the same visual clues (e.g., semantics, objects) changing their scale, position, and perspective patterns between adjacent frames. A primary trend in video CNN is adopting spatial-2D convolution for spatial semantics and temporal-1D convolution for temporal dynamics. Though the direction achieves a favorable balance between efficiency and efficacy, it suffers from misalignment of visual clues with large displacements. Particularly, rigid temporal convolution would fail to capture correct motions when a specific target moves out of the reception field of temporal convolution between adjacent frames.
Yi Tan 0001, Yanbin Hao, Hao Zhang 0047, Shuo Wang 0008, Xiangnan He 0001
ACM Multimedia4
2022 Parameterization of Cross-token Relations with Relative Positional Encoding for Vision MLP
abstract
Vision multi-layer perceptrons (MLPs) have shown promising performance in computer vision tasks, and become the main competitor of CNNs and vision Transformers. They use token-mixing layers to capture cross-token interactions, as opposed to the multi-head self-attention mechanism used by Transformers. However, the heavily parameterized token-mixing layers naturally lack mechanisms to capture local information and multi-granular non-local relations, thus their discriminative power is restrained. To tackle this issue, we propose a new positional spacial gating unit (PoSGU). It exploits the attention formulations used in the classical relative positional encoding (RPE), to efficiently encode the cross-token relations for token mixing. It can successfully reduce the current quadratic parameter complexity O(N2) of vision MLPs to $O(N)$ and O(1). We experiment with two RPE mechanisms, and further propose a group-wise extension to improve their expressive power with the accomplishment of multi-granular contexts. These then serve as the key building blocks of a new type of vision MLP, referred to as PosMLP. We evaluate the effectiveness of the proposed approach by conducting thorough experiments, demonstrating an improved or comparable performance with reduced parameter complexity. For instance, for a model trained on ImageNet1K, we achieve a performance improvement from 72.14% to 74.02% and a learnable parameter reduction from 19.4M to 18.2M. Code could be found at https://github.com/Zhicaiwww/PosMLP https://github.com/Zhicaiwww/PosMLP.
Zhicai Wang, Yanbin Hao, Xingyu Gao 0001, Hao Zhang 0047, Shuo Wang 0008, Tingting Mu, Xiangnan He 0001
ACM Multimedia5
2022 Attention in Attention: Modeling Context Correlation for Efficient Video Classification
abstract
Attention mechanisms have significantly boosted the performance of video classification neural networks thanks to the utilization of perspective contexts. However, the current research on video attention generally focuses on adopting a specific aspect of contexts (e.g., channel, spatial/temporal, or global context) to refine the features and neglects their underlying correlation when computing attentions. This leads to incomplete context utilization and hence bears the weakness of limited performance improvement. To tackle the problem, this paper proposes an efficient attention-in-attention (AIA) method for element-wise feature refinement, which investigates the feasibility of inserting the channel context into the spatio-temporal attention learning module, referred to as CinST, and also its reverse variant, referred to as STinC. Specifically, we instantiate the video feature contexts as dynamics aggregated along a specific axis with global average and max pooling operations. The workflow of an AIA module is that the first attention block uses one kind of context information to guide the gating weights calculation of the second attention that targets at the other context. Moreover, all the computational operations in attention units act on the pooled dimension, which results in quite few computational cost increase (https://github.com/haoyanbin918/Attention-in-Attention.
Yanbin Hao, Shuo Wang 0008, Pei Cao 0001, Xinjian Gao, Tong Xu 0001, Jinmeng Wu, Xiangnan He 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Spatio-Temporal Collaborative Module for Efficient Action Recognition
abstract
Efficient action recognition aims to classify a video clip into a specific action category with a low computational cost. It is challenging since the integrated spatial-temporal calculation (e.g., 3D convolution) introduces intensive operations and increases complexity. This paper explores the feasibility of the integration of channel splitting and filter decoupling for efficient architecture design and feature refinement by proposing a novel spatio-temporal collaborative (STC) module. STC splits the video feature channels into two groups and separately learns spatio-temporal representations in parallel with decoupled convolutional operators. Particularly, STC consists of two computation-efficient blocks,i.e., STand TS, where they extract either spatial (S.) or temporal (T.) features and further refine their features with either temporal (∙T) or spatial (∙S) contexts globally. The spatial/temporal context refers to information dynamics aggregated from temporal/spatial axis. To thoroughly examine our method’s performance in video action recognition tasks, we conduct extensive experiments using five video benchmark datasets requiring temporal reasoning. Experimental results show that the proposed STC networks achieve a competitive trade-off between model efficiency and effectiveness.
Yanbin Hao, Shuo Wang 0008, Yi Tan 0001, Xiangnan He 0001, Zhenguang Liu, Meng Wang 0001
IEEE Trans. Image Process.2
2022 Transductive Relation-Propagation With Decoupling Training for Few-Shot Learning
abstract
Few-shot learning, aiming to learn novel concepts from one or a few labeled examples, is an interesting and very challenging problem with many practical advantages. Existing few-shot methods usually utilize data of the same classes to train the feature embedding module and in a row, which is unable to learn adapting to new tasks. Besides, traditional few-shot models fail to take advantage of the valuable relations of the support-query pairs, leading to performance degradation. In this article, we propose a transductive relation-propagation graph neural network (GNN) with a decoupling training strategy (TRPN-D) to explicitly model and propagate such relations across support-query pairs, and empower the few-shot module the ability of transferring past knowledge to new tasks via the decoupling training. Our few-shot module, namely TRPN, treats the relation of each support-query pair as a graph node, named relational node, and resorts to the known relations between support samples, including both intraclass commonality and interclass uniqueness. Through relation propagation, the model could generate the discriminative relation embeddings for support-query pairs. To the best of our knowledge, this is the first work that decouples the training of the embedding network and the few-shot graph module with different tasks, which might offer a new way to solve the few-shot learning problem. Extensive experiments conducted on several benchmark datasets demonstrate that our method can significantly outperform a variety of state-of-the-art few-shot learning methods.
Yuqing Ma, Shihao Bai, Wei Liu 0005, Shuo Wang 0008, Yue Yu 0001, Xiao Bai 0001, Xianglong Liu 0001, Meng Wang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2021 Multi-Pretext Attention Network For Few-Shot Learning With Self-Supervision
abstract
Few-shot learning is an interesting and challenging study, which enables machines to learn from few samples like humans. Existing studies rarely exploit auxiliary information from large amount of unlabeled data. Self-supervised learning is emerged as an efficient method to utilize unlabeled data. Existing self-supervised learning methods always rely on the combination of geometric transformations for the single sample by augmentation, while seriously neglect the endogenous correlation information among different samples that is the same important for the task. In this work, we propose a Graph-driven Clustering (GC), a novel augmentation-free method for self-supervised learning, which does not rely on any auxiliary sample and utilizes the endogenous correlation information among input samples. Besides, we propose Multi-pretext Attention Network (MAN), which exploits a specific attention mechanism to combine the traditional augmentation-relied methods and our GC, adaptively learning their optimized weights to improve the performance and enabling the feature extractor to obtain more universal representations. We evaluate our MAN extensively on miniImageNet and tieredImageNet datasets and the results demonstrate that the proposed method outperforms the state-of-the-art (SOTA) relevant methods.1
Hainan Li, Renshuai Tao, Jun Li 0072, Haotong Qin, Yifu Ding 0001, Shuo Wang 0008, Xianglong Liu 0001
ICME6
2021 Few-shot learning with relation propagation and constraint
abstract
Abstract Previous deep learning methods usually required large‐scale annotated data, which is computationally exhaustive and unrealistic in certain scenarios. Therefore, few‐shot learning, where only a few annotated training images are available for training, has attracted increasing attention these days, showing huge potential in practical applications, such as portable equipment or security inspection, and so on. However, current few‐shot learning methods usually neglect the valuable semantic correlations between samples, thereby failing in extracting discriminating relations to achieve accurate predictive results. In this work, extending on a recent state‐of‐the‐art few‐shot learning method, transductive relation‐propagation network (TRPN), which considers the correlations between training samples, a constrained relation‐propagation network is proposed to further regularise the distilled correlations and thus achieve favourable few‐shot classification performance. The proposed framework contains three main components, namely preprocess module, relational propagation module, and relation constraint module. First, sample features are extracted and a relation graph node is constructed by treating the relation of each support–query pair as a graph node in the preprocess module. After that, in the relation propagation module (RPM), the valuable information of support–query pairs is modelled and propagated to directly generate the relational representations for further prediction. Then, a relation constraint module is introduced to regularise the relational representations and make it consistent with the ground‐truth relations as much as possible. With the guidance of the effective RPM and relation constraint module, the relational representations of the support–query pairs are distinguishable and thus can achieve accurate predictive results. Comprehensive experiments conducted on widely used benchmarks validate the effectiveness of our method compared to state‐of‐the‐art few‐shot classification approaches.
Huiyun Gong, Shuo Wang 0008, Yuqing Ma, Wei Liu 0005, Xianglong Liu 0001
IET Comput. Vis.2
2020 Dual Adversarial Network for Deep Active Learning
Shuo Wang 0008, Yuexiang Li, Kai Ma 0002, Ruhui Ma, Haibing Guan, Yefeng Zheng 0001
ECCV (24)1
2020 Large-Scale Few-Shot Learning via Multi-modal Knowledge Discovery
Shuo Wang 0008, Jun Yue 0004, Jianzhuang Liu, Qi Tian 0001, Meng Wang 0001
ECCV (10)1
2019 Dense Temporal Convolution Network for Sign Language Translation
abstract
The sign language translation (SLT) which aims at translating a sign language video into natural language is a weakly supervised task, given that there is no exact mapping relationship between visual actions and textual words in a sentence label. To align the sign language actions and translate them into the respective words automatically, this paper proposes a dense temporal convolution network, termed DenseTCN which captures the actions in hierarchical views. Within this network, a temporal convolution (TC) is designed to learn the short-term correlation among adjacent features and further extended to a dense hierarchical structure. In the kth TC layer, we integrate the outputs of all preceding layers together: (1) The TC in a deeper layer essentially has larger receptive fields, which captures long-term temporal context by the hierarchical content transition. (2) The integration addresses the SLT problem by different views, including embedded short-term and extended longterm sequential learning. Finally, we adopt the CTC loss and a fusion strategy to learn the featurewise classification and generate the translated sentence. The experimental results on two popular sign language benchmarks, i.e. PHOENIX and USTCConSents, demonstrate the effectiveness of our proposed method in terms of various measurements.
Dan Guo 0001, Shuo Wang 0008, Qi Tian 0001, Meng Wang 0001
IJCAI2
2019 Cross-Modality Retrieval by Joint Correlation Learning
abstract
As an indispensable process of cross-media analyzing, comprehending heterogeneous data faces challenges in the fields of visual question answering (VQA), visual captioning, and cross-modality retrieval. Bridging the semantic gap between the two modalities is still difficult. In this article, to address the problem in cross-modality retrieval, we propose a cross-modal learning model with joint correlative calculation learning. First, an auto-encoder is used to embed the visual features by minimizing the error of feature reconstruction and a multi-layer perceptron (MLP) is utilized to model the textual features embedding. Then we design a joint loss function to optimize both the intra- and the inter-correlations among the image-sentence pairs, i.e., the reconstruction loss of visual features, the relevant similarity loss of paired samples, and the triplet relation loss between positive and negative examples. In the proposed method, we optimize the joint loss based on a batch score matrix and utilize all mutual mismatched paired samples to enhance its performance. Our experiments in the retrieval tasks demonstrate the effectiveness of the proposed method. It achieves comparable performance to the state-of-the-art on three benchmarks, i.e., Flickr8k, Flickr30k, and MS-COCO.
Shuo Wang 0008, Dan Guo 0001, Xin Xu 0007, Li Zhuo 0001, Meng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.1
2018 Deep Learning based Fetal Middle Cerebral Artery Segmentation in Large-scale Ultrasound Images
Shuo Wang 0008, Yang Hua 0001, Yunyun Cao, Tao Song 0003, Zhengui Xue, Xiaoping Gong, Guanjie Wang, Ruhui Ma, Haibing Guan
BIBM1
2018 Connectionist Temporal Fusion for Sign Language Translation
abstract
Continuous sign language translation (CSLT) is a weakly supervised problem aiming at translating vision-based videos into natural languages under complicated sign linguistics, where the ordered words in a sentence label have no exact boundary of each sign action in the video. This paper proposes a hybrid deep architecture which consists of a temporal convolution module (TCOV), a bidirectional gated recurrent unit module (BGRU), and a fusion layer module (FL) to address the CSLT problem. TCOV captures short-term temporal transition on adjacent clip features (local pattern), while BGRU keeps the long-term context transition across temporal dimension (global pattern). FL concatenates the feature embedding of TCOV and BGRU to learn their complementary relationship (mutual pattern). Thus we propose a joint connectionist temporal fusion (CTF) mechanism to utilize the merit of each module. The proposed joint CTC loss optimization and deep classification score-based decoding fusion strategy are designed to boost performance. With only once training, our model under the CTC constraints achieves comparable performance to other existing methods with multiple EM iterations. Experiments are tested and verified on a benchmark, i.e. the RWTH-PHOENIX-Weather dataset, which demonstrate the effectiveness of our proposed method.
Shuo Wang 0008, Dan Guo 0001, Wengang Zhou 0001, Zhengjun Zha, Meng Wang 0001
ACM Multimedia1