Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shengwei Zhao

dblp:155/9654 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
9 papers
Language models and text generation · 22% Efficient and distributed learning · 22% Face, body and person analysis · 19%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 64% Audio and music processing · 36%

Topics — the 21 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.742020
Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency · IEEE Trans. Image Process. 2020
Efficient Low-Resolution Face Recognition via Bridge Distillation · IEEE Trans. Image Process. 2020
Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition · AAAI 2020
Machine learning › Representation and self-supervised learning
multimodal representation learning
1.422024
Multi-grained Correspondence Learning of Audio-language Models for Few-shot Audio Recognition · ACM Multimedia 2024
Multi-grained Representation Learning for Cross-modal Retrieval · SIGIR 2023
Computer vision › Face, body and person analysis
face recognition
1.342020
Efficient Low-Resolution Face Recognition via Bridge Distillation · IEEE Trans. Image Process. 2020
Low-Resolution Face Recognition in the Wild via Selective Knowledge Distillation · IEEE Trans. Image Process. 2019
Fewer-Shots and Lower-Resolutions: Towards Ultrafast Face Recognition in the Wild · ACM Multimedia 2019
Computer vision › Face, body and person analysis › face recognition › robust face recognition
low-resolution face recognition
1.232020
Efficient Low-Resolution Face Recognition via Bridge Distillation · IEEE Trans. Image Process. 2020
Low-Resolution Face Recognition in the Wild via Selective Knowledge Distillation · IEEE Trans. Image Process. 2019
Fewer-Shots and Lower-Resolutions: Towards Ultrafast Face Recognition in the Wild · ACM Multimedia 2019
Natural language and speech › Language models and text generation › retrieval-augmented generation
multimodal retrieval-augmented generation
1.012026
MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation · AAAI 2026
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation · AAAI 2026
Audio and music processing › audio analysis › audio content analysis
audio recognition
0.812024
Multi-grained Correspondence Learning of Audio-language Models for Few-shot Audio Recognition · ACM Multimedia 2024
Machine learning › Representation and self-supervised learning › representation learning
multi-level representation
0.712023
Multi-grained Representation Learning for Cross-modal Retrieval · SIGIR 2023
Multimedia analysis and retrieval › cross-modal retrieval
audio-text retrieval
0.712023
Multi-grained Representation Learning for Cross-modal Retrieval · SIGIR 2023
Multimedia analysis and retrieval
cross-modal retrieval
0.712023
Multi-grained Representation Learning for Cross-modal Retrieval · SIGIR 2023
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.612022
Uncertainty-Aware Learning against Label Noise on Imbalanced Datasets · AAAI 2022
Machine learning › Trustworthy machine learning
robustness
0.612022
Uncertainty-Aware Learning against Label Noise on Imbalanced Datasets · AAAI 2022
Machine learning › Trustworthy machine learning › uncertainty estimation
uncertainty-aware learning
0.612022
Uncertainty-Aware Learning against Label Noise on Imbalanced Datasets · AAAI 2022
Computer vision › Image recognition and object detection › visual recognition
low-resolution image recognition
0.412020
Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition · AAAI 2020
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
relational knowledge distillation
0.412020
Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition · AAAI 2020
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
spatio-temporal knowledge distillation
0.412020
Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency · IEEE Trans. Image Process. 2020
Computer vision › Video understanding and tracking
video saliency prediction
0.412020
Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency · IEEE Trans. Image Process. 2020
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
selective knowledge distillation
0.412019
Low-Resolution Face Recognition in the Wild via Selective Knowledge Distillation · IEEE Trans. Image Process. 2019
Information retrieval
multimodal retrieval
0.312026
MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation · AAAI 2026
Machine learning › Learning paradigms
class imbalance
0.212022
Uncertainty-Aware Learning against Label Noise on Imbalanced Datasets · AAAI 2022
Machine learning › Trustworthy machine learning
fairness
0.212022
Uncertainty-Aware Learning against Label Noise on Imbalanced Datasets · AAAI 2022

Methods — techniques the papers use, named apart from their topics

contrastive learning · 2.8rule-based reward · 2.0reinforcement fine-tuning · 2.0knowledge distillation · 1.6optimal transport · 1.5key-value cache · 1.5token interaction · 1.3adaptive aggregation · 1.3listwise ranking · 1.0list-wise ranking · 1.0aleatoric uncertainty · 0.6
YearPublicationVenuePosition
2026 MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation
abstract
Multi-modal Retrieval-Augmented Generation (MMRAG) enables highly credible generation by integrating external multi-modal knowledge, thus demonstrating impressive performance in complex multi-modal scenarios. However, existing MMRAG methods fail to clarify the reasoning logic behind retrieval and response generation, which limits the explainability of the results. To address this gap, we propose to introduce reinforcement learning into multi-modal retrieval-augmented generation, enhancing the reasoning capabilities of multi-modal large language models through a two-stage reinforcement fine-tuning framework to achieve explainable multi-modal retrieval-augmented generation. Specifically, in the first stage, rule-based reinforcement fine-tuning is employed to perform coarse-grained point-wise ranking of multi-modal documents, effectively filtering out those that are significantly irrelevant. In the second stage, reasoning-based reinforcement fine-tuning is utilized to jointly optimize fine-grained list-wise ranking and answer generation, guiding multi-modal large language models to output explainable reasoning logic in the MMRAG process. Our method achieves state-of-the-art results on WebQA and MultimodalQA, two benchmark datasets for multi-modal retrieval-augmented generation, and its effectiveness is validated through comprehensive ablation experiments.
Shengwei Zhao, Jingwen Yao, Sitong Wei, Linhai Xu, Yuying Liu 0007, Dong Zhang 0009, Shaoyi Du
AAAI1
2024 Multi-grained Correspondence Learning of Audio-language Models for Few-shot Audio Recognition
abstract
Large-scale pre-trained audio-language models excel in general multi-modal representation, facilitating their adaptation to downstream audio recognition tasks in a data-efficient manner. However, existing few-shot audio recognition methods based on audio-language models primarily focus on learning coarse-grained correlations, which are not sufficient to capture the intricate matching patterns between the multi-level information of audio and the diverse characteristics of category concepts. To address this gap, we propose multi-grained correspondence learning for bootstrapping audio-language models to improve audio recognition with few training samples. This approach leverages generative models to enrich multi-modal representation learning, mining the multi-level information of audio alongside the diverse characteristics of category concepts. Multi-grained matching patterns are then established through multi-grained key-value cache and multi-grained cross-modal contrast, enhancing the alignment between audio and category concepts. Additionally, we incorporate optimal transport to tackle temporal misalignment and semantic intersection issues in fine-grained correspondence learning, enabling flexible fine-grained matching. Our method achieves state-of-the-art results on multiple benchmark datasets for few-shot audio recognition, with comprehensive ablation experiments validating its effectiveness.
Shengwei Zhao, Linhai Xu, Yuying Liu 0007, Shaoyi Du
ACM Multimedia1
2023 CMFG: Cross-Model Fine-Grained Feature Interaction for Text-Video Retrieval
Shengwei Zhao, Yuying Liu 0007, Shaoyi Du, Linhai Xu
MMM (2)1
2023 Multi-grained Representation Learning for Cross-modal Retrieval
abstract
The purpose of audio-text retrieval is to learn a cross-modal similarity function between audio and text, enabling a given audio/text to find similar text/audio from a candidate set. Recent audio-text retrieval models aggregate multi-modal features into a single-grained representation. However, single-grained representation is difficult to solve the situation that an audio is described by multiple texts of different granularity levels, because the association pattern between audio and text is complex. Therefore, we propose an adaptive aggregation strategy to automatically find the optimal pool function to aggregate the features into a comprehensive representation, so as to learn valuable multi-grained representation. And multi-grained comparative learning is carried out in order to focus on the complex correlation between audio and text in different granularity. Meanwhile, text-guided token interaction is used to reduce the impact of redundant audio clips. We evaluated our proposed method on two audio-text retrieval benchmark datasets of Audiocaps and Clotho, achieving the state-of-the-art results in text-to-audio and audio-to-text retrieval. Our findings emphasize the importance of learning multi-modal multi-grained representation.
Shengwei Zhao, Linhai Xu, Yuying Liu 0007, Shaoyi Du
SIGIR1
2022 Uncertainty-Aware Learning against Label Noise on Imbalanced Datasets
abstract
Learning against label noise is a vital topic to guarantee a reliable performance for deep neural networks.Recent research usually refers to dynamic noise modeling with model output probabilities and loss values, and then separates clean and noisy samples.These methods have gained notable success. However, unlike cherry-picked data, existing approaches often cannot perform well when facing imbalanced datasets, a common scenario in the real world.We thoroughly investigate this phenomenon and point out two major issues that hinder the performance, i.e., inter-class loss distribution discrepancy and misleading predictions due to uncertainty.The first issue is that existing methods often perform class-agnostic noise modeling. However, loss distributions show a significant discrepancy among classes under class imbalance, and class-agnostic noise modeling can easily get confused with noisy samples and samples in minority classes.The second issue refers to that models may output misleading predictions due to epistemic uncertainty and aleatoric uncertainty, thus existing methods that rely solely on the output probabilities may fail to distinguish confident samples. Inspired by our observations, we propose an Uncertainty-aware Label Correction framework(ULC) to handle label noise on imbalanced datasets. First, we perform epistemic uncertainty-aware class-specific noise modeling to identify trustworthy clean samples and refine/discard highly confident true/corrupted labels.Then, we introduce aleatoric uncertainty in the subsequent learning process to prevent noise accumulation in the label noise modeling process. We conduct experiments on several synthetic and real-world datasets. The results demonstrate the effectiveness of the proposed method, especially on imbalanced datasets.
Yingsong Huang, Shengwei Zhao
AAAI3
2020 Look One and More: Distilling Hybrid Order Relational Knowledge for Cross-Resolution Image Recognition
abstract
In spite of great success in many image recognition tasks achieved by recent deep models, directly applying them to recognize low-resolution images may suffer from low accuracy due to the missing of informative details during resolution degradation. However, these images are still recognizable for subjects who are familiar with the corresponding high-resolution ones. Inspired by that, we propose a teacher-student learning approach to facilitate low-resolution image recognition via hybrid order relational knowledge distillation. The approach refers to three streams: the teacher stream is pretrained to recognize high-resolution images in high accuracy, the student stream is learned to identify low-resolution images by mimicking the teacher's behaviors, and the extra assistant stream is introduced as bridge to help knowledge transfer across the teacher to the student. To extract sufficient knowledge for reducing the loss in accuracy, the learning of student is supervised with multiple losses, which preserves the similarities in various order relational structures. In this way, the capability of recovering missing details of familiar low-resolution images can be effectively enhanced, leading to a better knowledge transfer. Extensive experiments on metric learning, low-resolution image classification and low-resolution face recognition tasks show the effectiveness of our approach, while taking reduced models.
Shiming Ge, Kangkai Zhang, Yingying Hua, Shengwei Zhao, Xin Jin 0015
AAAI5
2020 Occluded Face Recognition in the Wild by Identity-Diversity Inpainting
abstract
Face recognition has achieved advanced development by using convolutional neural network (CNN) based recognizers. Existing recognizers typically demonstrate powerful capacity in recognizing un-occluded faces, but often suffer from accuracy degradation when directly identifying occluded faces. This is mainly due to insufficient visual and identity cues caused by occlusions. On the other hand, generative adversarial network (GAN) is particularly suitable when it needs to reconstruct visually plausible occlusions by face inpainting. Motivated by these observations, this paper proposes identity-diversity inpainting to facilitate occluded face recognition. The core idea is integrating GAN with an optimized pre-trained CNN recognizer which serves as the third player to compete with the generator by distinguishing diversity within the same identity class. To this end, a collect of identity-centered features is applied in the recognizer as supervision to enable the inpainted faces clustering towards their identity centers. In this way, our approach can benefit from GAN for reconstruction and CNN for representation, and simultaneously addresses two challenging tasks, face inpainting and face recognition. Experimental results compared with 4 state-of-the-arts prove the efficacy of the proposed approach.
Shiming Ge, Chenyu Li 0001, Shengwei Zhao, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Efficient Low-Resolution Face Recognition via Bridge Distillation
abstract
Face recognition in the wild is now advancing towards light-weight models, fast inference speed and resolution-adapted capability. In this paper, we propose a bridge distillation approach to turn a complex face model pretrained on private high-resolution faces into a light-weight one for low-resolution face recognition. In our approach, such a cross-dataset resolution-adapted knowledge transfer problem is solved via two-step distillation. In the first step, we conduct cross-dataset distillation to transfer the prior knowledge from private high-resolution faces to public high-resolution faces and generate compact and discriminative features. In the second step, the resolution-adapted distillation is conducted to further transfer the prior knowledge to synthetic low-resolution faces via multi-task learning. By learning low-resolution face representations and mimicking the adapted high-resolution knowledge, a light-weight student model can be constructed with high efficiency and promising accuracy in recognizing low-resolution faces. Experimental results show that the student model performs impressively in recognizing low-resolution faces with only 0.21M parameters and 0.057MB memory. Meanwhile, its speed reaches up to 14,705, 934 and 763 faces per second on GPU, CPU and mobile phone, respectively.
Shiming Ge, Shengwei Zhao, Chenyu Li 0001, Yu Zhang 0035, Jia Li 0003
IEEE Trans. Image Process.2
2020 Spatiotemporal Knowledge Distillation for Efficient Estimation of Aerial Video Saliency
abstract
The performance of video saliency estimation techniques has achieved significant advances along with the rapid development of Convolutional Neural Networks (CNNs). However, devices like cameras and drones may have limited computational capability and storage space so that the direct deployment of complex deep saliency models becomes infeasible. To address this problem, this paper proposes a dynamic saliency estimation approach for aerial videos via spatiotemporal knowledge distillation. In this approach, five components are involved, including two teachers, two students and the desired spatiotemporal model. The knowledge of spatial and temporal saliency is first separately transferred from the two complex and redundant teachers to their simple and compact students, while the input scenes are also degraded from high-resolution to low-resolution to remove the probable data redundancy so as to greatly speed up the feature extraction process. After that, the desired spatiotemporal model is further trained by distilling and encoding the spatial and temporal saliency knowledge of two students into a unified network. In this manner, the inter-model redundancy can be removed for the effective estimation of dynamic saliency on aerial videos. Experimental results show that the proposed approach is comparable to 11 state-of-the-art models in estimating visual saliency on aerial videos, while its speed reaches up to 28,738 FPS and 1,490.5 FPS on the GPU and CPU platforms, respectively.
Jia Li 0003, Kui Fu, Shengwei Zhao, Shiming Ge
IEEE Trans. Image Process.3
2019 Fewer-Shots and Lower-Resolutions: Towards Ultrafast Face Recognition in the Wild
abstract
Is it possible to train an effective face recognition model with fewer shots that works efficiently on low-resolution faces in the wild? To answer this question, this paper proposes a few-shot knowledge distillation approach to learn an ultrafast face recognizer via two steps. In the first step, we initialize a simple yet effective face recognition model on synthetic low-resolution faces by distilling knowledge from an existing complex model. By removing the redundancies in both face images and the model structure, the initial model can provide an ultrafast speed with impressive recognition accuracy. To further adapt this model into the wild scenarios with fewer faces per person, the second step refines the model via few-shot learning by incorporating a relation module that compares low-resolution query faces with faces in the support set. In this manner, the performance of the model can be further enhanced with only fewer low-resolution faces in the wild. Experimental results show that the proposed approach performs favorably against state-of-the-arts in recognizing low-resolution faces with an extremely low memory of 30KB and runs at an ultrafast speed of 1,460 faces per second on CPU or 21,598 faces per second on GPU.
Shiming Ge, Shengwei Zhao, Xindi Gao, Jia Li 0003
ACM Multimedia2
2019 Low-Resolution Face Recognition in the Wild via Selective Knowledge Distillation
abstract
Typically, the deployment of face recognition models in the wild needs to identify low-resolution faces with extremely low computational cost. To address this problem, a feasible solution is compressing a complex face model to achieve higher speed and lower memory at the cost of minimal performance drop. Inspired by that, this paper proposes a learning approach to recognize low-resolution faces via selective knowledge distillation. In this approach, a two-stream convolutional neural network (CNN) is first initialized to recognize high-resolution faces and resolution-degraded faces with a teacher stream and a student stream, respectively. The teacher stream is represented by a complex CNN for high-accuracy recognition, and the student stream is represented by a much simpler CNN for low-complexity recognition. To avoid significant performance drop at the student stream, we then selectively distil the most informative facial features from the teacher stream by solving a sparse graph optimization problem, which are then used to regularize the finetuning process of the student stream. In this way, the student stream is actually trained by simultaneously handling two tasks with limited computational resources: approximating the most informative facial cues via feature regression, and recovering the missing facial cues via low-resolution face classification. Experimental results show that the student stream performs impressively in recognizing low-resolution faces and costs only 0.15MB memory and runs at 418 faces per second on CPU and 9; 433 faces per second on GPU.
Shiming Ge, Shengwei Zhao, Chenyu Li 0001, Jia Li 0003
IEEE Trans. Image Process.2
2017 Compressing deep neural networks for efficient visual inference
abstract
The deployments of deep neural network models on mobile or embedded devices have been challenged due to two main reasons: 1) the large model size for storage, and 2) the large memory bandwidth for inference. To address these issues, this paper develops a deep neural network compression framework to reduce the resource usage for efficient visual inference. By reviewing the trained deep model, we propose a hybrid model compression algorithm via four major modules. Approximation module reduces the number of weights in each fully connected layer with low rank approximation. Then, quantization module analyzes weight distribution in each layer and represents them with low precision fixed point, which reduces the bits for storing each weight. After that, pruning module suppresses small weights to further reduce the number of parameters. Finally, coding module joint optimizes the representation and encoding of the sparse structure of the pruned weights with relative index by Huffman coding. Beyond the compression of model size, we propose an adaptive fixed point memory allocation algorithm to reduce memory footprint in inference. The proposed framework, along with the model compression and memory allocation algorithms, can provide 20-30x compression rate with negligible accuracy loss. We conduct an evaluation on two representative models, AlexNet and VGG-16, for object recognition and face verification tasks, which demonstrate the effectiveness of our proposed compression framework.
Shiming Ge, Zhao Luo, Shengwei Zhao, Xin Jin 0015, Xiaoyu Zhang 0002
ICME3
2015 Exudates and optic disk detection in retinal images of diabetic patients
abstract
SUMMARY Diabetic retinopathy is the progressive pathological alterations in the retinal microvasculature that very often causes blindness. Because of its clinical significance, it will be helpful to have regular cost‐effective eye screening for diabetic patients by developing algorithms to perform retinal image analysis, fundus image enhancement, and monitoring. The two cost‐effective algorithms are proposed for exudates detection and optic disk extraction aimed for retinal images classification and diagnosis assistance. They represent the effort made to offer a cost‐effective algorithm for optic disk identification, which will enable easier exudates extraction, exudates detection and retinal images classification aimed to assist ophthalmologists while making diagnoses. The proposed algorithms apply mathematical modeling, which enables light intensity levels emphasis, easier optic disk and exudates detection, efficient and correct classification of retinal images. The algorithm is robust to various appearance changes of retinal fundus images and shows very promising results. Fundus images are classified into those that are healthy and those affected by diabetes, based on the detected optic disk and exudates. The obtained results indicate that the proposed algorithm successfully and correctly classifies more than 98% of the observed retinal images because of the changes in the appearance of retinal fundus images typically encountered in clinical environments. Copyright © 2014 John Wiley & Sons, Ltd.
Vesna Zeljkovic, Milena Bojic, Shengwei Zhao, Claude Tameze, Ventzeslav Valev
Concurr. Comput. Pract. Exp.3