EDBT 2026 Demo / reviewers in the wild / expert
Li Liu 0036
dblp:33/4528-36
· DBLP profile ↗
69ranked-venue papers
6as first author
62since 2021 · last 2026
0000-0002-4497-0135ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 48 · 6 first-author · 43 since 2021Artificial intelligence and machine learning · 37 · 2 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting ASR Robustness via Test-Time Reinforcement Learning with Audio-Text Semantic RewardsabstractRecently, Automatic Speech Recognition (ASR) systems (e.g., Whisper) have achieved remarkable accuracy improvements but remain highly sensitive to real-world unseen data (data with large distribution shifts), including noisy environments and diverse accents. To address this issue, test-time adaptation (TTA) has shown great potential in improving the model adaptability at inference time without ground-truth labels, and existing TTA methods often rely on pseudo-labeling or entropy minimization. However, by treating model confidence as a learning signal, these methods may reinforce high-confidence errors, leading to confirmation bias that undermines adaptation. To overcome these limitations, we present ASR-TRA, a novel Test-time Reinforcement Adaptation framework inspired by causal intervention. More precisely, our method introduces a learnable decoder prompt and utilizes temperature-controlled stochastic decoding to generate diverse transcription candidates. These are scored by a reward model that measures audio-text semantic alignment, and the resulting feedback is used to update both model and prompt parameters via reinforcement learning. Comprehensive experiments on LibriSpeech with synthetic noise and L2 Arctic accented English datasets demonstrate that our method significantly outperforms existing state-of-the-art (SOTA), including SUTA and SGEM, in both accuracy and inference speed. Ablation studies further confirm the effectiveness of combining audio and language-based rewards, highlighting our method's enhanced stability and interpretability. Overall, our approach provides a practical and robust solution for deploying ASR systems in challenging real-world conditions. Linghan Fang, Tianxin Xie, Li Liu 0036 |
AAAI | 3 |
| 2026 | TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization AbilityabstractAchieving zero-shot adversarial robustness without sacrificing generalization remains challenging for foundation models such as CLIP, especially under large adversarial perturbations. Through empirical analyses, we identify three critical yet overlooked issues: (1) Logit margins exhibit a stable offset between small and large adversarial perturbations, suggesting that explicitly adjusting margins could improve robustness against unseen large perturbations. (2) A significant negative correlation exists between logit margin and inter-class semantic similarity, indicating that semantic structures are insufficiently leveraged by existing methods. (3) Existing methods for adjusting text embeddings disrupt the intrinsic semantic consistency established by pre-trained models, undermining generalization capability. Motivated by these findings, we propose a novel Text-Image Mutual Awareness (TIMA) framework, including a Text-Aware Image (TAI) tuning module with an Adaptive Semantic-Aware Margin (ASAM) to explicitly calibrate logit margins, and an Image-Aware Text (IAT) tuning module with Semantic Consistent Minimum Hyperspherical Energy (SC-MHE) to preserve semantic consistency. Comprehensive experiments validate that TIMA significantly outperforms existing approaches by effectively addressing the identified limitations. Fengji Ma, Hei Victor Cheng, Chenxing Li, Li Liu 0036 |
AAAI | 4 |
| 2026 | Cueing Without Gapping: Cuer-Independent Cued Speech Recognition Powered by Cross-Cuer Invariant ModelingabstractAutomatic Cued Speech Recognition (ACSR) is a vital communication system designed to enhance spoken language accessibility for the hearing-impaired by combining lip movements and hand gestures to encode phonemes. Despite its effectiveness, current ACSR methods face significant challenges, including poor generalization to unseen cuers due to the limited scale of CS datasets, which restricts the ability of existing visual encoder to capture cuer-invariant CS visual features. Additionally, previous approaches relying on Connectionist Temporal Classification (CTC) decoding fail to incorporate prior linguistic sequence knowledge, further limiting their performance. To address these issues, we propose a novel Two Auxiliary Modalities guided Cross-cuer Invariant Adaptation method (TACIA), introducing pose and text modalities to help extract cuer-invariant motion and semantic features, thereby improving generalization. In addition, we introduce a Visual-guided Cued Token Prediction (VG-NTP) method, inspired by large language models. This method replaces CTC decoding by incorporating language modeling, leveraging rich linguistic knowledge, including semantics, to address the suboptimal issues present in the CTC decoding process. Extensive experiments demonstrate the superiority of our approach to the state-of-the-art (SOTA) on Chinese and British CS datasets, significantly advancing the accuracy and quality of ACSR systems. Fengji Ma, Chenxing Li, Li Liu 0036 |
AAAI | 3 |
| 2026 | UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech GenerationabstractCued Speech (CS) enhances lipreading via hand coding, offering visual phonemic cues that support precise speech perception for the hearing-impaired. The task of CS Video-to-Speech generation (CSV2S) aims to convert CS videos into intelligible speech signals. Most existing research focuses on CS Recognition (CSR), which transcribes video content into text. Consequently, a common solution for CSV2S is to integrate CSR with a text-to-speech (TTS) system. However, this pipeline relies on text as an intermediate medium, which may lead to error propagation and temporal misalignment between speech and CS video dynamics. In contrast, directly generating audio speech from CS video (direct CSV2S) often suffer from the inherent multimodal complexity and the limited availability of CS data. To address these challenges, we propose UniCUE, the first unified framework for CSV2S that directly generates speech from CS videos without relying on intermediate text. The core innovation of UniCUE lies in integrating a understanding task (CSR) that provides fine-grained CS visual-semantic cues to to guide the speech generation. Specifically, UniCUE incorporates a pose-aware visual processor, a semantic alignment pool that enables precise visual–semantic mapping, and a VisioPhonetic adapter to bridge the understanding and generation tasks within a unified architecture. To support this framework, we construct UniCUE-HI, a large-scale Mandarin CS dataset containing 11,282 videos from 14 cuers, including both hearing-impaired and normal-hearing individuals. Extensive experiments conducted on this dataset demonstrate that UniCUE achieves state-of-the-art (SOTA) performance across multiple evaluation metrics. Jinting Wang, Shan Yang 0001, Chenxing Li, Dong Yu 0001, Li Liu 0036 |
AAAI | 5 |
| 2026 | Attacks in Adversarial Machine Learning: A Systematic Survey from the Lifecycle Perspective
Baoyuan Wu, Zihao Zhu 0001, Li Liu 0036, Qingshan Liu 0001, Zhaofeng He 0001, Siwei Lyu |
Int. J. Comput. Vis. | 3 |
| 2026 | WPDA: frequency-based backdoor attack with wavelet packet decomposition
Zhengyao Song, Danni Yuan, Li Liu 0036, Shaokui Wei, Baoyuan Wu |
Neural Networks | 4 |
| 2026 | Versatile Backdoor Attack With Visible, Semantic, Sample-Specific and Compatible TriggersabstractDeep neural networks (DNNs) can be manipulated to exhibit specific behaviors when exposed to specific trigger patterns, without affecting their performance on benign samples, dubbed backdoor attack. Currently, implementing backdoor attacks in physical scenarios still faces significant challenges. Physical attacks are labor-intensive and time-consuming, and the triggers are selected in a manual and heuristic way. Moreover, expanding digital attacks to physical scenarios faces many challenges due to their sensitivity to visual distortions and the absence of counterparts in the real world. To address these challenges, we define a novel trigger called the Visible, Semantic, Sample-specific, and Compatible (VSSC) trigger, to achieve effective, stealthy and robust simultaneously, which can also be effectively deployed in the physical scenario using corresponding objects. To implement the VSSC trigger, we propose an automated pipeline comprising three modules: a trigger selection module that systematically identifies suitable triggers leveraging large language models, a trigger insertion module that employs generative models to seamlessly integrate triggers into images, and a quality assessment module that ensures the natural and successful insertion of triggers through vision-language models. Extensive experimental results and analysis validate the effectiveness, stealthiness, and robustness of the VSSC trigger. It can not only maintain robustness under visual distortions but also demonstrates strong practicality in the physical scenario. By providing the first automated pipeline, VSSC transforms physical backdoor attacks from a labor-intensive craft into a systematic and realistic threat to real-world AI systems. We hope the proposed VSSC trigger and implementation approach could inspire future studies on designing more practical triggers in backdoor attacks. Ruotong Wang 0008, Hongrui Chen, Zihao Zhu 0001, Li Liu 0036, Baoyuan Wu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Defenses in Adversarial Machine Learning: A Systematic Survey From the Lifecycle PerspectiveabstractAdversarial phenomena have been widely observed in machine learning (ML) systems, especially those using deep neural networks. These phenomena describe situations where ML systems may produce predictions that are inconsistent and incomprehensible to humans in certain specific cases. Such behavior poses a serious security threat to the practical application of ML systems. To exploit this vulnerability, several advanced attack paradigms have been developed, mainly including backdoor attacks, weight attacks, and adversarial examples. For each individual attack paradigm, various defense mechanisms have been proposed to enhance the robustness of models against the corresponding attacks. However, due to the independence and diversity of these defense paradigms, it is challenging to assess the overall robustness of an ML system against different attack paradigms. This survey aims to provide a systematic review of all existing defense paradigms from a unified lifecycle perspective. Specifically, we decompose a complete ML system into five stages: pre-training, training, post-training, deployment, and inference. We then present a clear taxonomy to categorize representative defense methods at each stage. The unified perspective and taxonomy not only help us analyze defense mechanisms but also enable us to understand the connections and differences among different defense paradigms. It inspires future research to develop more advanced and comprehensive defense strategies. Baoyuan Wu, Mingli Zhu, Meixi Zheng, Zihao Zhu 0001, Shaokui Wei, Hongrui Chen, Danni Yuan, Li Liu 0036, Qingshan Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2026 | FauForensics: Boosting Audio-Visual Deepfake Detection With Facial Action UnitsabstractThe rapid evolution of generative AI has intensified the threat of realistic audio-visual deepfakes, demanding robust and generalizable detection methods. Existing solutions primarily address unimodal (e.g., audio, visual) forgeries but struggle with multimodal manipulations due to inadequate handling of heterogeneous modality features and poor cross-dataset generalization. We propose FauForensics, a novel framework leveraging biologically invariant facial action units (FAUs), which are quantitative descriptors of facial muscle activity linked to emotion physiology. They serve as forgery-resistant representations that reduce domain dependency while capturing subtle synthetic-content disruptions. In addition, unlike prior clip-level comparisons, our method computes frame-wise audio-visual similarities via a fusion module with learnable cross-modal queries, dynamically aligning lip-audio relationships and mitigating feature heterogeneity. Experiments on four publicly available datasets show state-of-the-art performance with 5.17% average cross-dataset improvement over existing methods. Jian Wang 0129, Baoyuan Wu, Li Liu 0036, Qingshan Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Fusing Pruned and Backdoored Models: Optimal Transport-based Data-free Backdoor MitigationabstractBackdoor attacks present a serious security threat to deep neuron networks (DNNs). Although numerous effective defense techniques have been proposed in recent years, they inevitably rely on the availability of either clean or poisoned data. In contrast, data-free defense techniques have evolved slowly and still lag significantly in performance. To address this issue, different from the traditional approach of pruning followed by fine-tuning, we propose a novel data-free defense method named Optimal Transport-based Backdoor Repairing (OTBR) in this work. This method, based on our findings on neuron weight changes (NWCs) of random unlearning, uses optimal transport (OT)-based model fusion to combine the advantages of both pruned and backdoored models. Specifically, we first demonstrate our findings that the NWCs of random unlearning are positively correlated with those of poison unlearning. Based on this observation, we propose a random-unlearning NWC pruning technique to eliminate the backdoor effect and obtain a backdoor-free pruned model. Then, motivated by the OT-based model fusion, we propose the pruned-to-backdoored OT-based fusion technique, which fuses pruned and backdoored models to combine the advantages of both, resulting in a model that demonstrates high clean accuracy and a low attack success rate. To our knowledge, this is the first work to apply OT and model fusion techniques to backdoor defense. Extensive experiments show that our method successfully defends against all seven backdoor attacks across three benchmark datasets, outperforming both state-of-the-art (SOTA) data-free and data-dependent methods. Weilin Lin, Li Liu 0036, Jianze Li |
AAAI | 2 |
| 2025 | Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyabstractText-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style.Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has become a rapidly growing research area.This survey provides the first comprehensive review of controllable TTS methods, from traditional control techniques to emerging approaches using natural language prompts.We categorize model architectures, control strategies, and feature representations, while also summarizing challenges, datasets, and evaluations in controllable TTS.This survey aims to guide researchers and practitioners by offering a clear taxonomy and highlighting future directions in this fast-evolving field.One can visit https://github.com/imxtx/ awesome-controllabe-speech-synthesis for a comprehensive paper list and updates. Tianxin Xie, Yan Rong, Pengfei Zhang 0005, Wenwu Wang 0001, Li Liu 0036 |
EMNLP | 5 |
| 2025 | Teaching Others Teaches Yourself: Semi-supervised Ensembled Pseudo-labeling Method for Image ClassificationabstractSemi-supervised methods have recently received significant attention in deep learning because they are able to reduce the dependence on labeled data while ensuring good performance. The pseudo-label based method has been widely used as a classic semi-supervised method but still suffers from the confirmation bias problem, which causes significant damage to the model’s training process. To solve this problem, we propose an ensembled semi-supervised framework, which can effectively improve the quality of pseudo-labels by ensembling the prediction results of multiple models to generate pseudo-labels. Simultaneously, we innovatively design a Ensembled Divergence Promotion (EDP) method to increase the diversity among different models for better ensembling results. Extensive experiments are conducted on CIFAR-10 and CIFAR-100, showing superior performance of the proposed method to the state-of-the-art (SOTA) semi-supervised method. Meanwhile, extensive ablation experiments and visualization results are conducted to prove the effectiveness of our method. Wentao Lei, Li Liu 0036 |
ICASSP | 2 |
| 2025 | Multi-Modal Rhythmic Generative Model for Chinese Cued Speech Gestures GenerationabstractCued speech (CS) is a novel visual coding system, which combines lip reading with several specific hand codings to help hearing-impaired people to communicate effectively. This work focuses on the audio/text-driven CS gestures (i.e., continuous lip and hand gestures movements) generation. Previous work used template-based statistical methods for the French CS generation. However, these methods are fragile since they need careful hand-crafted pre-processing to fit models, resulting in poor robustness. Furthermore, the natural rhythm in generated CS gesture sequences, which is essential for a coding system of spoken languages, was overlooked in prior studies. To solve the above-mentioned problems, we innovatively propose a two-branched rhythmic CS gesture generation framework, which contains a multi-modal adversarial semantic generator (MASG) to generate accurate multi-modal CS gestures (i.e., lip, hand shape and hand position movements), and an audio-driven rhythm generator (ARG) to extract the rhythm information. Moreover, we design a new Gesture Audio Difference (GAD) metric to evaluate the rhythm coherence considering the issue of asynchrony between CS hand gestures and lip movements. Extensive experimental results are presented on two datasets of two tasks (a CS dataset named MCCS-2024 and a co-speech TED dataset) with comprehensive ablation analysis and user study, demonstrating the effectiveness of our method. The code and dataset with multi-modal annotations were made public at https://mccs-2024.github.io/. Li Liu 0036, Wentao Lei, Wenwu Wang 0001 |
ICASSP | 1 |
| 2025 | MotionComposer: Enhancing Rhythmic Music Generation with Adaptive Retrieval ReferenceabstractWith the rise of the AIGC era, rhythmic music generation has extensive applications, particularly with the surge in motion video creation. However, generating music that is rhythmically synchronized and stylistically aligned with motion video presents significant challenges. Although existing methods have made progress, they still face difficulties in producing high-quality long-term music, particularly when addressing complex rhythmic patterns and maintaining style-consistent musical chords. In this work, we present MotionComposer, a novel retrieval-augmented, easy-to-hard training approach designed to enhance rhythmic music generation. By leveraging the inherent alignment between motion rhythms and music beats, we first tackle the simpler task of beat prediction with BeatNet, which predicts music beats by analyzing motion patterns. To address the complex musical chord generation, we propose ChordNet, a retrieval-augmented network that integrates external data to enrich chord generation. Additionally, to minimize the impact of irrelevant retrievals, we design RAGate, a retrieval adaptive module that selectively filters out low-relevance retrieval references during the retrieval process. Extensive experiments across three scenarios (i.e., dance, figure skating, and floor exercise) demonstrate that our approach significantly enhances video soundtrack generation, achieving new state-of-the-art performance. Our project is available at https://beria-moon.github.io/Soundtrackyour-Motion/. Jinting Wang, Li Liu 0036 |
ICASSP | 2 |
| 2025 | Fine-portraitist: Visualizing the Speaker's Face Portrait during Speech ListeningabstractSpeech-to-portrait generation (S2P) plays a crucial role in speech-driven, human-centered creative content generation, aiming to synthesize a speaker’s face portrait with identity consistency from a given speech clip. However, existing S2P methods can typically only preserve attribute consistency, e.g., gender and age, while failing to capture the more important part-appearance consistency due to the coarse speech-face correlation. In this work, we propose Fine-portraitist, a novel retrieval-augmented, easy-to-hard generation framework designed to tackle this problem. Specifically, Fine-portraitist enhances identity consistency in S2P through two key innovations: 1) We first explore the fine-grained speech-face correlation by decomposing the face portrait into speech-related and speech-unrelated parts. Based on this, we propose a two-stage, diffusion-based pipeline to progressively achieve S2P; 2) A retrieval prior is introduced, selected from a retrieval database based on speech feature similarity, providing supplementary external information for more accurate and realistic generation results. Extensive experiments on two datasets, i.e., AVSpeech and VoxCeleb, demonstrate that Fine-portraitist significantly outperforms existing S2P methods. Jinting Wang, Li Liu 0036 |
ICASSP | 2 |
| 2025 | Inter- and Intra-Sentence Cuer-Invariant Representation Learning for Generalizable Cued Speech RecognitionabstractCued Speech (CS) is a visual coding system that combines lip movements and hand gestures to represent spoken languages for hearing-impaired people. Automatic Cued Speech Recognition (ACSR) is an emerging research topic, but the cuer (i.e., people who perform CS) generalization problem of ACSR remains unexplored. Moreover, the asynchrony between lip and hand modalities further aggravates the challenges associated with the generalization problem. Therefore, we propose a novel multi-modal, cuer-invariant representation learning framework that facilitates generalizable ACSR through contrastive learning, enabling our model to recognize unseen cuers. The proposed approach comprises three key components: 1) an inter-sentence contrastive learning module that learns cuer-invariant representations for lip and hand modalities to solve the cuer generalization problem in ACSR; and 2) a multi-level intra-sentence cross-attention module that synchronizes the lip movements and hand gestures to address the modality asynchrony issue in ACSR; 3) an easy-to-hard progressive learning strategy to stabilize the learning process and prevent performance degradation on hard examples. Extensive experiments on two available CS datasets show that our method outperforms previous works by a large margin. We also conduct ablation studies and visualization to demonstrate the effectiveness of the proposed method. Tianxin Xie, Li Liu 0036 |
ICASSP | 2 |
| 2025 | Reliable Imputed-Sample Assisted Vertical Federated LearningabstractVertical Federated Learning (VFL) is a well-known FL variant that enables multiple parties to collaboratively train a model without sharing their raw data. Existing VFL approaches focus on overlapping samples among different parties, while their performance is constrained by the limited number of these samples, leaving numerous non-overlapping samples unexplored. Some previous work has explored techniques for imputing missing values in samples, but often without adequate attention to the quality of the imputed samples. To address this issue, we propose a Reliable Imputed-Sample Assisted (RISA) VFL framework to effectively exploit non-overlapping samples by selecting reliable imputed samples for training VFL models. Specifically, after imputing non-overlapping samples, we introduce evidence theory to estimate the uncertainty of imputed samples, and only samples with low uncertainty are selected. In this way, high-quality non-overlapping samples are utilized to improve VFL model. Experiments on two widely used datasets demonstrate the significant performance gains achieved by the RISA, especially with the limited overlapping samples, e.g., a 48% accuracy gain on CIFAR-10 with only 1% overlapping samples. Yaopei Zeng, Lei Liu 0049, Shaoguo Liu, Hongjian Dou, Baoyuan Wu, Li Liu 0036 |
ICASSP | 6 |
| 2025 | MambaTrack: Exploiting Dual-Enhancement for Night UAV TrackingabstractNight unmanned aerial vehicle (UAV) tracking is impeded by the challenges of poor illumination, with previous daylight-optimized methods demonstrating suboptimal performance in low-light conditions, limiting the utility of UAV applications. To this end, we propose an efficient mamba-based tracker, leveraging dual enhancement techniques to boost night UAV tracking. The mamba-based low-light enhancer, equipped with an illumination estimator and a damage restorer, achieves global image enhancement while preserving the details and structure of low-light images. Additionally, we advance a cross-modal mamba network to achieve efficient interactive learning between vision and language modalities. Extensive experiments showcase that our method achieves advanced performance and exhibits significantly improved computation and memory efficiency. For instance, our method is 2.8× faster than CiteTracker and reduces 50.2% GPU memory. Our codes are available at https://github.com/983632847/Awesome-Multimodal-Object-Tracking. Chunhui Zhang 0001, Li Liu 0036, Xi Zhou 0001, Yanfeng Wang 0001 |
ICASSP | 2 |
| 2025 | Learning Class Unique Features in Fine-Grained Visual ClassificationabstractA major challenge in Fine-Grained Visual Classification (FGVC) is distinguishing various categories with high inter-class similarity by learning the feature that differentiates the details. Conventional cross-entropy trained Convolutional Neural Network (CNN) fails this challenge as they may suffer from producing inter-class invariant features in FGVC. In this work, we innovatively propose to regularize the training of CNN by enforcing the uniqueness of the features of each category from an information-theoretic perspective. To achieve this goal, we formulate a minimax loss based on a game-theoretic framework, where a Nash equilibrium is proved to be consistent with this regularization objective. Besides, to avoid getting a solution that produces redundant features, we present a Feature Redundancy Loss (FRL) based on the normalized inner product between each selected feature map pair to complement the proposed minimax loss. The proposed method is versatile, as it can be utilized as a regularizer for features in the mid-level or the penultimate layer, and can be combined with any architectures. Extensive experimental results on several influential benchmarks along with visualization show that our method obtains significant improvement over the baseline model without extra cost and achieves state-of-the-art results. Runkai Zheng, Li Liu 0036, Zhijia Yu, Yinqi Zhang, Hei Victor Cheng, Chris Ding |
ICASSP | 2 |
| 2025 | Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech RecognitionabstractBackdoor attacks have posed a significant threat to the security of deep neural networks (DNNs). Despite considerable strides in developing defenses against backdoor attacks in the visual domain, the specialized defenses for the audio domain remain empty. Furthermore, the defenses adapted from the visual to audio domain demonstrate limited effectiveness. To fill this gap, we propose Gradient Norm-based Fine-Tuning (GN-FT), a novel defense strategy against the attacks in the audio domain, based on the observation from the corresponding backdoored models. Specifically, we first empirically find that the backdoored neurons exhibit greater gradient values compared to other neurons, while clean neurons stay the lowest. On this basis, we fine-tune the backdoored model by incorporating the gradient norm regularization, aiming to weaken and reduce the backdoored neurons. We further approximate the loss computation for lower implementation costs. Extensive experiments on two speech recognition datasets across five models demonstrate the superior performance of our proposed method. To the best of our knowledge, this work is the first specialized and effective defense against backdoor attacks in the audio domain. Nanjun Zhou, Weilin Lin, Li Liu 0036 |
ICASSP | 3 |
| 2025 | Failure Cases Are Better Learned but Boundary Says Sorry: Facilitating Smooth Perception Change for Accuracy-Robustness Trade-Off in Adversarial TrainingabstractAdversarial Training (AT) is one of the most effective methods to train robust Deep Neural Networks (DNNs). However, AT creates an inherent trade-off between clean accuracy and adversarial robustness, which is commonly attributed to the more complicated decision boundary caused by the insufficient learning of hard adversarial samples. In this work, we reveal a counterintuitive fact for the first time: From the perspective of perception consistency, hard adversarial samples that can still attack the robust model after AT are already learned better than those successfully defended. Thus, different from previous views, we argue that it is rather the over-sufficient learning of hard adversarial samples that degrades the decision boundary and contributes to the trade-off problem. Specifically, the excessive pursuit of perception consistency would force the model to view the perturbations as noise and ignore the information within them, which should have been utilized to induce a smoother perception transition towards the decision boundary to support its establishment to an appropriate location. In response, we define a new AT objective named Robust Perception, encouraging the model perception to change smoothly with input perturbations, based on which we propose a novel Robust Perception Adversarial Training (RPAT) method, effectively mitigating the current accuracy-robustness trade-off. Experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet with ResNet-18, PreActResNet-18, and WideResNet-34-10 demonstrate the effectiveness of our method beyond four common baselines and 12 state-of-the-art (SOTA) works. The code is available at https://github.com/FlaAI/RPAT. Yanyun Wang 0003, Li Liu 0036 |
ICCV | 2 |
| 2025 | Activation Gradient based Poisoned Sample Detection Against Backdoor AttacksabstractThis work studies the task of poisoned sample detection for defending against data poisoning based backdoor attacks. Its core challenge is finding a generalizable and discriminative metric to distinguish between clean and various types of poisoned samples (e.g., various triggers, various poisoning ratios). Inspired by a common phenomenon in backdoor attacks that the backdoored model tend to map significantly different poisoned and clean samples within the target class to similar activation areas, we introduce a novel perspective of the circular distribution of the gradients w.r.t. sample activation, dubbed gradient circular distribution (GCD). And, we find two interesting observations based on GCD. One is that the GCD of samples in the target class is much more dispersed than that in the clean class. The other is that in the GCD of target class, poisoned and clean samples are clearly separated. Inspired by above two observations, we develop an innovative three-stage poisoned sample detection approach, called Activation Gradient based Poisoned sample Detection (AGPD). First, we calculate GCDs of all classes from the model trained on the untrustworthy dataset. Then, we identify the target class(es) based on the difference on GCD dispersion between target and clean classes. Last, we filter out poisoned samples within the identified target class(es) based on the clear separation between poisoned and clean samples. Extensive experiments under various settings of backdoor attacks demonstrate the superior detection performance of the proposed method to existing poisoned detection approaches according to sample activation-based metrics. Danni Yuan, Shaokui Wei, Li Liu 0036, Baoyuan Wu |
ICLR | 4 |
| 2025 | Cued-Agent: A Collaborative Multi-Agent System for Automatic Cued Speech RecognitionabstractCued Speech (CS) is a visual communication system that combines lip-reading with hand coding to facilitate communication for individuals with hearing impairments. Automatic CS Recognition (ACSR) aims to convert CS hand gestures and lip movements into text via AI-driven methods. Traditionally, the temporal asynchrony between hand and lip movements requires the design of complex modules to facilitate effective multimodal fusion. However, constrained by limited data availability, current methods demonstrate insufficient capacity for adequately training these fusion mechanisms, resulting in suboptimal performance. Recently, multi-agent systems have shown promising capabilities in handling complex tasks with limited data availability. To this end, we propose the first collaborative multi-agent system for ACSR, named Cued-Agent. It integrates four specialized sub-agents: a Multimodal Large Language Model-based Hand Recognition agent that employs keyframe screening and CS expert prompt strategies to decode hand movements, a pretrained Transformer-based Lip Recognition agent that extracts lip features from the input video, a Hand Prompt Decoding agent that dynamically integrates hand prompts with lip features during inference in a training-free manner, and a Self-Correction Phoneme-to-Word agent that enables post-processing and end-to-end conversion from phoneme sequences to natural language sentences for the first time through semantic refinement. To support this study, we expand the existing Mandarin CS dataset by collecting data from eight hearing-impaired cuers, establishing a mixed dataset of fourteen subjects. Extensive experiments demonstrate that our Cued-Agent performs superbly in both normal and hearing-impaired scenarios compared with state-of-the-art methods. The implementation is available at https://github.com/DennisHgj/Cued-Agent. Guanjie Huang, Danny H. K. Tsang, Shan Yang 0001, Guangzhi Lei, Li Liu 0036 |
ACM Multimedia | 5 |
| 2025 | AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation
Yan Rong, Jinting Wang, Guangzhi Lei, Shan Yang 0001, Li Liu 0036 |
ACM Multimedia | 5 |
| 2025 | BackdoorDM: A Comprehensive Benchmark for Backdoor Learning on Diffusion ModelabstractBackdoor learning is a critical research topic for understanding the vulnerabilities of deep neural networks. While the diffusion model (DM) has been broadly deployed in public over the past few years, the understanding of its backdoor vulnerability is still in its infancy compared to the extensive studies in discriminative models. Recently, many different backdoor attack and defense methods have been proposed for DMs, but a comprehensive benchmark for backdoor learning on DMs is still lacking. This absence makes it difficult to conduct fair comparisons and thoroughly evaluate existing approaches, thus hindering future research progress. To address this issue, we propose BackdoorDM, the first comprehensive benchmark designed for backdoor learning on DMs. It comprises nine state-of-the-art (SOTA) attack methods, four SOTA defense strategies, and three useful visualization analysis tools. We first systematically classify and formulate the existing literature in a unified framework, focusing on three different backdoor attack types and five backdoor target types, which are restricted to a single type in discriminative models. Then, we systematically summarize the evaluation metrics for each type and propose a unified backdoor evaluation method based on multimodal large language model (MLLM). Finally, we conduct a comprehensive evaluation and highlight several important conclusions. We believe that BackdoorDM will help overcome current barriers and contribute to building a trustworthy artificial intelligence generated content (AIGC) community. The codes are released in https://github.com/linweiii/BackdoorDM. Weilin Lin, Nanjun Zhou, Yanyun Wang 0003, Jianze Li, Li Liu 0036 |
NeurIPS | 6 |
| 2025 | Boosting Nighttime UAV Tracking via Self-prompting Autoregressive Learning and a New Benchmark
Chunhui Zhang 0001, Li Liu 0036, Xi Zhou 0001, Yanfeng Wang 0001 |
PRCV (16) | 2 |
| 2025 | BackdoorBench: A Comprehensive Benchmark and Analysis of Backdoor Learning
Baoyuan Wu, Hongrui Chen, Zihao Zhu 0001, Shaokui Wei, Danni Yuan, Mingli Zhu, Ruotong Wang 0008, Li Liu 0036 |
Int. J. Comput. Vis. | 9 |
| 2024 | Leveraging Noisy Labels of Nearest Neighbors for Label Correction and Sample SelectionabstractDealing with noisy labels (LNL) emerges as a critical challenge when applying deep learning (DL) in practical settings. Previous methodologies primarily concentrated on harnessing model predictions to mitigate the impact of noisy labels. Nevertheless, their efficacy is strongly contingent on the accuracy of model predictions, a factor that cannot be assured in the context of LNL. Our empirical analysis shows that in noisy datasets, the spatial information of latent feature representation combined with original noisy labels is more robust than the methods using model predictions. To mitigate the unreliability introduced by model predictions, we propose a novel Feature Representation method, which utilizes noisy labels of nearest neighbors for label Correction and sample Selection (FRCS). Extensive experiments on various benchmark datasets demonstrate the superiority of FRCS compared with SOTA methods. Our codes are available at https://github.com/tianfangjh/FRCS-Noisy-Labels. Yixiong Chen, Li Liu 0036, Xiaoguang Han 0001, Xiao-Ping Zhang 0002 |
ICASSP | 3 |
| 2024 | Bridge to Non-Barrier Communication: Gloss-Prompted Fine-Grained Cued Speech Gesture Generation with Diffusion Model
Wentao Lei, Li Liu 0036 |
IJCAI | 2 |
| 2024 | Prior-free Balanced Replay: Uncertainty-guided Reservoir Sampling for Long-Tailed Continual LearningabstractEven in the era of large models, one of the well-known issues in continual learning (CL) is catastrophic forgetting, which is significantly challenging when the continual data stream exhibits a long-tailed distribution, termed as Long-Tailed Continual Learning (LTCL). Existing LTCL solutions generally require the label distribution of the data stream to achieve re-balance training. However, obtaining such prior information is often infeasible in real scenarios since the model should learn without pre-identifying the majority and minority classes. To this end, we propose a novel Prior-free Balanced Replay (PBR) framework to learn from long-tailed data stream with less forgetting. Concretely, motivated by our experimental finding that the minority classes are more likely to be forgotten due to the higher uncertainty, we newly design an uncertainty-guided reservoir sampling strategy to prioritize rehearsing minority data without using any prior information, which is based on the mutual dependence between the model and samples. Additionally, we incorporate two prior-free components to further reduce the forgetting issue: (1) Boundary constraint is to preserve uncertain boundary supporting samples for continually re-estimating task boundaries. (2) Prototype constraint is to maintain the consistency of learned class prototypes along with training. Our approach is evaluated on three standard long-tailed benchmarks, demonstrating superior performance to existing CL methods and previous SOTA LTCL approach in both task- and class-incremental learning settings, as well as ordered- and shuffled-LTCL settings. © 2024 ACM. Lei Liu 0049, Li Liu 0036, Yawen Cui |
ACM Multimedia | 2 |
| 2024 | Unveiling and Mitigating Backdoor Vulnerabilities based on Unlearning Weight Changes and Backdoor ActivenessabstractThe security threat of backdoor attacks is a central concern for deep neural networks (DNNs). Recently, without poisoned data, unlearning models with clean data and then learning a pruning mask have contributed to backdoor defense. Additionally, vanilla fine-tuning with those clean data can help recover the lost clean accuracy. However, the behavior of clean unlearning is still under-explored, and vanilla fine-tuning unintentionally induces back the backdoor effect. In this work, we first investigate model unlearning from the perspective of weight changes and gradient norms, and find two interesting observations in the backdoored model: 1) the weight changes between poison and clean unlearning are positively correlated, making it possible for us to identify the backdoored-related neurons without using poisoned data; 2) the neurons of the backdoored model are more active (*i.e.*, larger gradient norm) than those in the clean model, suggesting the need to suppress the gradient norm during fine-tuning. Then, we propose an effective two-stage defense method. In the first stage, an efficient *Neuron Weight Change (NWC)-based Backdoor Reinitialization* is proposed based on observation 1). In the second stage, based on observation 2), we design an *Activeness-Aware Fine-Tuning* to replace the vanilla fine-tuning. Extensive experiments, involving eight backdoor attacks on three benchmark datasets, demonstrate the superior performance of our proposed method compared to recent state-of-the-art backdoor defense approaches. The code is available at https://github.com/linweiii/TSBD.git. Weilin Lin, Li Liu 0036, Shaokui Wei, Jianze Li |
NeurIPS | 2 |
| 2024 | WebUOT-1M: Advancing Deep Underwater Object Tracking with A Million-Scale BenchmarkabstractUnderwater Object Tracking (UOT) is essential for identifying and tracking submerged objects in underwater videos, but existing datasets are limited in scale, diversity of target categories and scenarios covered, impeding the development of advanced tracking algorithms. To bridge this gap, we take the first step and introduce WebUOT-1M, \ie, the largest public UOT benchmark to date, sourced from complex and realistic underwater environments. It comprises 1.1 million frames across 1,500 video clips filtered from 408 target categories, largely surpassing previous UOT datasets, \eg, UVOT400. Through meticulous manual annotation and verification, we provide high-quality bounding boxes for underwater targets. Additionally, WebUOT-1M includes language prompts for video sequences, expanding its application areas, \eg, underwater vision-language tracking. Given that most existing trackers are designed for open-air conditions and perform poorly in underwater environments due to domain gaps, we propose a novel framework that uses omni-knowledge distillation to train a student Transformer model effectively. To the best of our knowledge, this framework is the first to effectively transfer open-air domain knowledge to the UOT model through knowledge distillation, as demonstrated by results on both existing UOT datasets and the newly proposed WebUOT-1M. We have thoroughly tested WebUOT-1M with 30 deep trackers, showcasing its potential as a benchmark for future UOT research. The complete dataset, along with codes and tracking results, are publicly accessible at \href{https://github.com/983632847/Awesome-Multimodal-Object-Tracking}{\color{magenta}{here}}. Chunhui Zhang 0001, Li Liu 0036, Guanjie Huang, Xi Zhou 0001, Yanfeng Wang 0001 |
NeurIPS | 2 |
| 2024 | Less confidence, less forgetting: Learning with a humbler teacher in exemplar-free Class-Incremental learning
Zijian Gao, Kele Xu, Huiping Zhuang, Li Liu 0036, Xinjun Mao, Bo Ding 0001, Huaimin Wang 0001 |
Neural Networks | 4 |
| 2024 | Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech RecognitionabstractCued Speech (CS) is a pure visual coding method used by hearing-impaired people that combines lip reading with several specific hand shapes to make the spoken language visible. Automatic CS recognition (ACSR) seeks to transcribe visual cues of speech into text, which can help hearing-impaired people to communicate effectively. The visual information of CS contains lip reading and hand cueing, thus the fusion of them plays an important role in ACSR. However, most previous fusion methods struggle to capture the global dependency present in long sequence inputs of multi-modal CS data. As a result, these methods generally fail to learn the effective cross-modal relationships that contribute to the fusion. Recently, attentionbased transformers have been a prevalent idea for capturing the global dependency over the long sequence in multi-modal fusion, but existing multi-modal fusion transformers suffer from both poor recognition accuracy and inefficient computation for the ACSR task. To address these problems, we develop a novel computation and parameter efficient multi-modal fusion transformer by proposing a novel Token-Importance-Aware Attention mechanism (TIAA), where a token utilization rate (TUR) is formulated to select the important tokens from the multi-modal streams. More precisely, TIAA firstly models the modality-specific fine-grained temporal dependencies over all tokens of each modality, and then learns the efficient cross-modal interaction for the modality-shared coarse-grained temporal dependencies over the important tokens of different modalities. Besides, a lightweight gated hidden projection is designed to control the feature flows of TIAA. The resulting model, named Economical Cued Speech Fusion Transformer (EcoCued), achieves state-of-the-art performance on all existing CS datasets (i.e., Mandarin Chinese, French, and British CS), compared with existing transformerbased fusion methods and ACSR fusion methods. Notably, our method dramatically reduces the computational complexity from O(T2) to O(T). We will release the source code and data as open source. Lei Liu 0049, Li Liu 0036, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2024 | FHVAC: Feature-Level Hybrid Video Adaptive Configuration for Machine-Centric Live StreamingabstractWith the widespread deployment of edge computing, the focus has shifted to machine-centric live video streaming, where endpoint-collected videos are transmitted over networks to edge servers for analysis. Unlike maximizing user's Quality of Experience (QoE), machine-centric video streaming optimizes the machine's Quality of Inference (QoI) by balancing the inference accuracy, inference delay, and transmission latency with video adaptive configuration. Traditional heuristic configuration adaption methods are reliable but unable to respond to erratic network fluctuations. Reinforcement learning (RL) based algorithms exhibit superior flexibility but suffer from exploration mechanisms, resulting in long-tail effects on upload latency. In this paper, we propose FHVAC, which dynamically selects video encoding parameters for live streaming by coherently fusing rule-based and RL-based agent at the feature level. We initially develop a robust rule-based approach for ensuring the low latency in transmission, and employ imitation learning to convert it into a neural network equivalently. Subsequently, we design a novel module to combine the two approaches and assess various fusion mechanisms. Our evaluation of FHVAC across two vision tasks (pose estimation and semantic segmentation) in two scenarios (trace-driven simulation and testbed-based experiment) shows that FHVAC enhances the average QoI, and reduces 10.61%-65.27% latency tail performance compared to prior work. Yuanhong Zhang, Weizhan Zhang, Haipeng Du, Caixia Yan, Li Liu 0036 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2023 | TAOTF: A Two-Stage Approximately Orthogonal Training Framework in Deep Neural NetworksabstractThe orthogonality constraints, including the hard and soft ones, have been used to normalize the weight matrices of Deep Neural Network (DNN) models, especially the Convolutional Neural Network (CNN) and Vision Transformer (ViT), to reduce model parameter redundancy and improve training stability. However, the robustness to noisy data of these models with constraints is not always satisfactory. In this work, we propose a novel two-stage approximately orthogonal training framework (TAOTF) to find a trade-off between the orthogonal solution space and the main task solution space to solve this problem in noisy data scenarios. In the first stage, we propose a novel algorithm called polar decomposition-based orthogonal initialization (PDOI) to find a good initialization for the orthogonal optimization. In the second stage, unlike other existing methods, we apply soft orthogonal constraints for all layers of DNN model. We evaluate the proposed model-agnostic framework both on the natural image and medical image datasets, which show that our method achieves stable and superior performances to existing methods. Supplementary materials can be found in https://github.com/nonameinformation/anonymous/tree/main. Taoyong Cui, Jianze Li, Yuhan Dong, Li Liu 0036 |
ECAI | 4 |
| 2023 | Spatio-Temporal Structure Consistency for Semi-Supervised Medical Image ClassificationabstractIntelligent medical diagnosis has shown remarkable progress on the large-scale datasets with full annotations. However, very few labeled images are available due to significantly expensive annotations by experts. To efficiently leverage abundant unlabeled data, we propose a novel Spatio-Temporal Structure Consistent (STSC) learning framework to combine both spatial and temporal structure consistency. Specifically, a gram matrix is derived to capture the structural similarity among different training samples in the representation space. At the spatial level, our framework explicitly enforces the consistency of structural similarity among different samples under perturbations. At the temporal level, we desire to maintain the consistency of the structural similarity in different training iterations by digging out the stable sub-structures in a relation graph. Experiments on two medical image datasets (i.e., ISIC 2018 and ChestX-ray14) show that our method outperforms state-of-the-art Semi-Supervised Learning (SSL) methods. Furthermore, extensive qualitative analysis on the Gram matrices and heatmaps by Grad-CAM are presented to validate the effectiveness of our method. Wentao Lei, Lei Liu 0049, Li Liu 0036 |
ICASSP | 3 |
| 2023 | Cross-Modal Mutual Learning for Cued Speech RecognitionabstractAutomatic Cued Speech Recognition (ACSR) provides an intelligent human-machine interface for visual communications, where the Cued Speech (CS) system utilizes lip movements and hand gestures to code spoken language for hearing-impaired people. Previous ACSR approaches often utilize direct feature concatenation as the main fusion paradigm. However, the asynchronous modalities (i.e., lip, hand shape and hand position) in CS may cause interference for feature concatenation. To address this challenge, we propose a transformer based cross-modal mutual learning framework to prompt multi-modal interaction. Compared with the vanilla self-attention, our model forces modality-specific information of different modalities to pass through a modality-invariant codebook, concatenating linguistic representations with tokens of each modality. Then the shared linguistic knowledge is used to re-synchronize multi-modal sequences. Moreover, we establish a novel large-scale multi-speaker CS dataset for Mandarin Chinese. To our knowledge, this is the first work on ACSR for Mandarin Chinese. Extensive experiments are conducted for different languages (i.e., Chinese, French, and British English). Results demonstrate that our model exhibits superior recognition performance to the state-of-the-art by a large margin. Lei Liu 0049, Li Liu 0036 |
ICASSP | 2 |
| 2023 | Two-Stream Joint-Training for Speaker Independent Acoustic-to-Articulatory InversionabstractAcoustic-to-articulatory inversion (AAI) aims to estimate the parameters of articulators from speech audio. There are two common challenges in AAI, which are the limited data and the unsatisfactory performance in speaker independent scenario. Most current works focus on extracting features directly from speech and ignoring the importance of phoneme information which may limit the performance of AAI. To this end, we propose a novel network called SPN that uses two different streams to carry out the AAI task. Firstly, to improve the performance of speaker-independent experiment, we propose a new phoneme stream network to estimate the articulatory parameters as the phoneme features. To the best of our knowledge, this is the first work that extracts the speaker-independent features from phonemes to improve the performance of AAI. Secondly, in order to better represent the speech information, we train a speech stream network to combine the local features and the global features. Compared with state-of-the-art (SOTA), the proposed method reduces 0.18mm on RMSE and increases 6.0% on Pearson correlation coefficient in the speaker-independent experiment. The code has been released at https://github.com/liujinyu123/AAINetwork-SPN. Jianrong Wang, Xuewei Li 0001, Mei Yu 0004, Jie Gao 0008, Qiang Fang 0003, Li Liu 0036 |
ICASSP | 7 |
| 2023 | Memory-Augmented Contrastive Learning for Talking Head GenerationabstractGiven one reference facial image and a piece of speech as input, talking head generation aims to synthesize a realistic-looking talking head video. However, generating a lip-synchronized video with natural head movements is challenging. The same speech clip can generate multiple possible lip and head movements, that is, there is no one-to-one mapping relationship between them. To overcome this problem, we propose a Speech Feature Extractor (SFE) based on memory-augmented self-supervised contrastive learning, which introduces the memory module to store multiple different speech mapping results. In addition, we introduce the Mixed Density Networks (MDN) into the landmark regression task to generate multiple predicted facial landmarks. Extensive qualitative and quantitative experiments show that the quality of our facial animation is significantly superior to that of the state-of-the-art (SOTA). The code has been released at https://github.com/Yaxinzhao97/MACL.git. Jianrong Wang, Yaxin Zhao, Hongkai Fan, Li Liu 0036 |
ICASSP | 7 |
| 2023 | Global Balanced Experts for Federated Long-Tailed LearningabstractFederated learning (FL) is a prevalent distributed machine learning approach that enables collaborative training of a global model across multiple devices without sharing local data. However, the presence of long-tailed data can negatively deteriorate the model’s performance in real-world FL applications. Moreover, existing re-balance strategies are less effective for the federated long-tailed issue when directly utilizing local label distribution as the class prior at the clients’ side. To this end, we propose a novel Global Balanced Multi-Expert (GBME) framework to optimize a balanced global objective, which does not require additional information beyond the standard FL pipeline. In particular, a proxy is derived from the accumulated gradients uploaded by the clients after local training, and is shared by all clients as the class prior for re-balance training. Such a proxy can also guide the client grouping to train a multi-expert model, where the knowledge from different clients can be aggregated via the ensemble of different experts corresponding to different client groups. To further strengthen the privacy-preserving ability, we present a GBME-p algorithm with a theoretical guarantee to prevent privacy leakage from the proxy. Extensive experiments on long-tailed decentralized datasets demonstrate the effectiveness of GBME and GBME-p, both of which show superior performance to state-of-the-art methods. The code is available at here. Yaopei Zeng, Lei Liu 0049, Li Liu 0036, Li Shen 0008, Shaoguo Liu, Baoyuan Wu |
ICCV | 3 |
| 2023 | A Novel Interpretable and Generalizable Re-synchronization Model for Cued Speech based on a Multi-Cuer CorpusabstractCued Speech (CS) is a multi-modal visual coding system combining lip reading with several hand cues at the phonetic level to make the spoken language visible to the hearing impaired. Previous studies solved asynchronous problems between lip and hand movements by a cuer-dependent piecewise linear model for English and French CS. In this work, we innovatively propose three statistical measure on the lip stream to build an interpretable and generalizable model for predicting hand preceding time (HPT), which achieves cuer-independent by a proper normalization. Particularly, we build the first Mandarin CS corpus comprising annotated videos from five speakers including three normal and two hearing impaired individuals. Consequently, we show that the hand preceding phenomenon exists in Mandarin CS production with significant differences between normal and hearing impaired people. Extensive experiments demonstrate that our model outperforms the baseline and the previous state-of-the-art methods. Lufei Gao, Li Liu 0036 |
INTERSPEECH | 3 |
| 2023 | MAVD: The First Open Large-Scale Mandarin Audio-Visual Dataset with Depth Information
Jianrong Wang, Yuchen Huo, Li Liu 0036 |
INTERSPEECH | 3 |
| 2023 | Emotional Talking Head Generation based on Memory-Sharing and Attention-Augmented NetworksabstractGiven an audio clip and a reference face image, the goal of the talking head generation is to generate a high-fidelity talking head video. Although some audio-driven methods of generating talking head videos have made some achievements in the past, most of them only focused on lip and audio synchronization and lack the ability to reproduce the facial expressions of the target person. To this end, we propose a talking head generation model consisting of a Memory-Sharing Emotion Feature extractor (MSEF) and an Attention-Augmented Translator based on U-net (AATU). Firstly, MSEF can extract implicit emotional auxiliary features from audio to estimate more accurate emotional face landmarks. Secondly, AATU acts as a translator between the estimated landmarks and the photo-realistic video frames. Extensive qualitative and quantitative experiments have shown the superiority of the proposed method to the previous works. Codes will be made publicly available. Jianrong Wang, Yaxin Zhao, Li Liu 0036 |
INTERSPEECH | 3 |
| 2023 | MetaLR: Meta-tuning of Learning Rates for Transfer Learning in Medical Imaging
Yixiong Chen, Li Liu 0036, Jingxian Li, Chris Ding, Zongwei Zhou |
MICCAI (1) | 2 |
| 2023 | Cuing Without Sharing: A Federated Cued Speech Recognition Framework via Mutual Knowledge DistillationabstractCued Speech (CS) is a visual coding tool to encode spoken languages at the phonetic level, which combines lip-reading and hand gestures to effectively assist communication among people with hearing impairments. The Automatic CS Recognition (ACSR) task aims to recognize CS videos into linguistic texts, which involves both lips and hands as two distinct modalities conveying complementary information. However, the traditional centralized training approach poses potential privacy risks due to the use of facial and gesture videos in CS data. To address this issue, we propose a new Federated Cued Speech Recognition (FedCSR) framework to train an ACSR model over the decentralized CS data without sharing private information. In particular, a mutual knowledge distillation method is proposed to maintain cross-modal semantic consistency of the Non-IID CS data, which ensures learning a unified feature space for linguistic and visual information. On the server side, a globally shared linguistic model is trained to capture the long-term dependencies in the text sentences, which is aligned with the visual information from the local clients via visual-to-linguistic distillation. On the client side, the visual model of each client is trained with its own local data, assisted by linguistic-to-visual distillation treating the linguistic model as the teacher. To the best of our knowledge, this is the first approach to consider the federated ACSR task for privacy protection. Experimental results on the Chinese CS dataset with multiple cuers demonstrate that our approach outperforms both mainstream federated learning baselines and existing centralized state-of-the-art ACSR methods, achieving 9.7% performance improvement for character error rate (CER) and 15.0% for word error rate (WER). The Chinese CS dataset and our code will be open-sourced. Lei Liu 0049, Li Liu 0036 |
ACM Multimedia | 3 |
| 2023 | All in One: Exploring Unified Vision-Language Tracking with Multi-Modal AlignmentabstractCurrent mainstream vision-language (VL) tracking framework consists of three parts,i.e., a visual feature extractor, a language feature extractor, and a fusion model. To pursue better performance, a natural modus operandi for VL tracking is employing customized and heavier unimodal encoders, and multi-modal fusion models. Albeit effective, existing VL trackers separate feature extraction and feature integration, resulting in extracted features that lack semantic guidance and have limited target-aware capability in complex scenarios, e.g., similar distractors and extreme illumination. In this work, inspired by the recent success of exploring foundation models with unified architecture for both natural language and computer vision tasks, we propose an All-in-One framework, which learns joint feature extraction and interaction by adopting a unified transformer backbone. Specifically, we mix raw vision and language signals to generate language-injected vision tokens, which we then concatenate before feeding into the unified backbone architecture. This approach achieves feature integration in a unified backbone, removing the need for carefully-designed fusion modules and resulting in a more effective and efficient VL tracking framework. To further improve the learning efficiency, we introduce a multi-modal alignment module based on cross-modal and intra-modal contrastive objectives, providing more reasonable representations for the unified All-in-One transformer backbone. Extensive experiments on five benchmarks, i.e., OTB99-L, TNL2K, LaSOT, LaSOTExt and WebUAV-3M, demonstrate the superiority of the proposed tracker against existing state-of-the-art (SOTA) methods on VL tracking. Codes will be available at https://github.com/983632847/All-in-One here. Chunhui Zhang 0001, Xin Sun 0020, Yiqian Yang, Li Liu 0036, Xi Zhou 0001, Yanfeng Wang 0001 |
ACM Multimedia | 4 |
| 2023 | FedAds: A Benchmark for Privacy-Preserving CVR Estimation with Vertical Federated LearningabstractConversion rate (CVR) estimation aims to predict the probability of conversion event after a user has clicked an ad. Typically, online publisher has user browsing interests and click feedbacks, while demand-side advertising platform collects users' post-click behaviors such as dwell time and conversion decisions. To estimate CVR accurately and protect data privacy better, vertical federated learning (vFL) is a natural solution to combine two sides' advantages for training models, without exchanging raw data. Both CVR estimation and applied vFL algorithms have attracted increasing research attentions. However, standardized and systematical evaluations are missing: due to the lack of standardized datasets, existing studies adopt public datasets to simulate a vFL setting via hand-crafted feature partition, which brings challenges to fair comparison. We introduce FedAds, the first benchmark for CVR estimation with vFL, to facilitate standardized and systematical evaluations for vFL algorithms. It contains a large-scale real world dataset collected from Alibaba's advertising platform, as well as systematical evaluations for both effectiveness and privacy aspects of various vFL algorithms. Besides, we also explore to incorporate unaligned data in vFL to improve effectiveness, and develop perturbation operations to protect privacy well. We hope that future research work in vFL and CVR estimation benefits from the FedAds benchmark. Penghui Wei, Hongjian Dou, Shaoguo Liu, Rongjun Tang, Li Liu 0036, Liang Wang 0001, Bo Zheng 0007 |
SIGIR | 5 |
| 2023 | WebUAV-3M: A Benchmark for Unveiling the Power of Million-Scale Deep UAV TrackingabstractUnmanned aerial vehicle (UAV) tracking is of great significance for a wide range of applications, such as delivery and agriculture. Previous benchmarks in this area mainly focused on small-scale tracking problems while ignoring the amounts of data, types of data modalities, diversities of target categories and scenarios, and evaluation protocols involved, greatly hiding the massive power of deep UAV tracking. In this article, we propose WebUAV-3M, the largest public UAV tracking benchmark to date, to facilitate both the development and evaluation of deep UAV trackers. WebUAV-3M contains over 3.3 million frames across 4,500 videos and offers 223 highly diverse target categories. Each video is densely annotated with bounding boxes by an efficient and scalable semi-automatic target annotation (SATA) pipeline. Importantly, to take advantage of the complementary superiority of language and audio, we enrich WebUAV-3M by innovatively providing both natural language specifications and audio descriptions. We believe that such additions will greatly boost future research in terms of exploring language features and audio cues for multi-modal UAV tracking. In addition, a fine-grained UAV tracking-under-scenario constraint (UTUSC) evaluation protocol and seven challenging scenario subtest sets are constructed to enable the community to develop, adapt and evaluate various types of advanced trackers. We provide extensive evaluations and detailed analyses of 43 representative trackers and envision future research directions in the field of deep UAV tracking and beyond. The dataset, toolkits, and baseline results are available at https://github.com/983632847/WebUAV-3M. Chunhui Zhang 0001, Guanjie Huang, Li Liu 0036, Shiming Ge, Dacheng Tao |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Generating and Weighting Semantically Consistent Sample Pairs for Ultrasound Contrastive LearningabstractWell-annotated medical datasets enable deep neural networks (DNNs) to gain strong power in extracting lesion-related features. Building such large and well-designed medical datasets is costly due to the need for high-level expertise. Model pre-training based on ImageNet is a common practice to gain better generalization when the data amount is limited. However, it suffers from the domain gap between natural and medical images. In this work, we pre-train DNNs on ultrasound (US) domains instead of ImageNet to reduce the domain gap in medical US applications. To learn US image representations based on unlabeled US videos, we propose a novel meta-learning-based contrastive learning method, namely Meta Ultrasound Contrastive Learning (Meta-USCL). To tackle the key challenge of obtaining semantically consistent sample pairs for contrastive learning, we present a positive pair generation module along with an automatic sample weighting module based on meta-learning. Experimental results on multiple computer-aided diagnosis (CAD) problems, including pneumonia detection, breast cancer classification, and breast tumor segmentation, show that the proposed self-supervised method reaches state-of-the-art (SOTA). The codes are available at https://github.com/Schuture/Meta-USCL. Yixiong Chen, Chunhui Zhang 0001, Chris Ding, Li Liu 0036 |
IEEE Trans. Medical Imaging | 4 |
| 2022 | HiCo: Hierarchical Contrastive Learning for Ultrasound Video Model Pretraining
Chunhui Zhang 0001, Yixiong Chen, Li Liu 0036, Xi Zhou 0001 |
ACCV (6) | 3 |
| 2022 | Boosting Black-Box Attack with Partially Transferred Conditional Adversarial DistributionabstractThis work studies black-box adversarial attacks against deep neural networks (DNNs), where the attacker can only access the query feedback returned by the attacked DNN model, while other information such as model parameters or the training datasets are unknown. One promising approach to improve attack performance is utilizing the adversarial transferability between some white-box surrogate models and the target model (i.e., the attacked model). However, due to the possible differences on model architectures and training datasets between surrogate and target models, dubbed “surrogate biases”, the contribution of adversarial transferability to improving the attack performance may be weakened. To tackle this issue, we innovatively propose a black-box attack method by developing a novel mechanism of adversarial transferability, which is robust to the surrogate biases. The general idea is transferring partial parameters of the conditional adversarial distribution (CAD) of surrogate models, while learning the untransferred parameters based on queries to the target model, to keep the flexibility to adjust the CAD of the target model on any new benign sample. Extensive experiments on benchmark datasets and attacking against real-world API demonstrate the superior attack performance of the proposed method. The code will be available at https://github.com/Kira0096/CGATTACK. Baoyuan Wu, Yanbo Fan, Li Liu 0036, Zhifeng Li 0001, Shutao Xia |
CVPR | 4 |
| 2022 | Data-Free Backdoor Removal Based on Channel Lipschitzness
Runkai Zheng, Rongjun Tang, Jianze Li, Li Liu 0036 |
ECCV (5) | 4 |
| 2022 | Acoustic-to-Articulatory Inversion Based on Speech Decomposition and Auxiliary FeatureabstractAcoustic-to-articulatory inversion (AAI) is to obtain the movement of articulators from speech signals. Until now, achieving a speaker-independent AAI remains a challenge given the limited data. Besides, most current works only use audio speech as input, causing an inevitable performance bottleneck. To solve these problems, firstly, we pre-train a speech decomposition network to decompose audio speech into speaker embedding and content embedding as the new personalized speech features to adapt to the speaker-independent case. Secondly, to further improve the AAI, we propose a novel auxiliary feature network to estimate the lip auxiliary features from the above personalized speech features. Experimental results on three public datasets show that, compared with the state-of-the-art only using the audio speech feature, the proposed method reduces the average RMSE by 0.25 and increases the average correlation coefficient by 2.0% in the speaker-dependent case. More importantly, the average RMSE decreases by 0.29 and the average correlation coefficient increases by 5.0% in the speaker-independent case. Jianrong Wang, Longxuan Zhao, Shanyu Wang, Li Liu 0036 |
ICASSP | 6 |
| 2022 | Residual-Guided Personalized Speech Synthesis based on Face ImageabstractPrevious works derive personalized speech features by training the model on a large dataset composed of his/her audio sounds. It was reported that face information has a strong link with the speech sound. Thus in this work, we innovatively extract personalized speech features from human faces to synthesize personalized speech using neural vocoder. A Face-based Residual Personalized Speech Synthesis Model (FR-PSS) containing a speech encoder, a speech synthesizer and a face encoder is designed for PSS. In this model, by designing two speech priors, a residual-guided strategy is introduced to guide the face feature to approach the true speech feature in the training. Moreover, considering the error of feature’s absolute values and their directional bias, we formulate a novel tri-item loss function for face encoder. Experimental results show that the speech synthesized by our model is comparable to the personalized speech synthesized by training a large amount of audio data in previous works. Jianrong Wang, Xiaosheng Hu, Xuewei Li 0001, Qiang Fang 0003, Li Liu 0036 |
ICASSP | 6 |
| 2022 | MVNet: Memory Assistance and Vocal Reinforcement Network for Speech Enhancement
Jianrong Wang, Xuewei Li 0001, Mei Yu 0004, Qiang Fang 0003, Li Liu 0036 |
ICONIP (2) | 6 |
| 2022 | Pre-activation Distributions Expose Backdoor NeuronsabstractConvolutional neural networks (CNN) can be manipulated to perform specific behaviors when encountering a particular trigger pattern without affecting the performance on normal samples, which is referred to as backdoor attack. The backdoor attack is usually achieved by injecting a small proportion of poisoned samples into the training set, through which the victim trains a model embedded with the designated backdoor. In this work, we demonstrate that backdoor neurons are exposed by their pre-activation distributions, where populations from benign data and poisoned data show significantly different moments. This property is shown to be attack-invariant and allows us to efficiently locate backdoor neurons. On this basis, we make several proper assumptions on the neuron activation distributions, and propose two backdoor neuron detection strategies based on (1) the differential entropy of the neurons, and (2) the Kullback-Leibler divergence between the benign sample distribution and a poisoned statistics based hypothetical distribution. Experimental results show that our proposed defense strategies are both efficient and effective against various backdoor attacks. Runkai Zheng, Rongjun Tang, Jianze Li, Li Liu 0036 |
NeurIPS | 4 |
| 2021 | Self-Supervised Depth Estimation Via Implicit Cues from VideosabstractIn self-supervised monocular depth estimation, the depth discontinuity and motion objects' artifacts are still challenging problems. Existing self-supervised methods usually utilize two views to train the depth estimation network and use one single view to make predictions. Compared with static views, abundant dynamic properties between video frames are beneficial to refining depth estimation, especially for dynamic objects. In this work, we improve the self-supervised learning framework for depth estimation using consecutive frames from monocular and stereo videos. The main idea is to exploit an implicit depth cue extractor which leverages dynamic and static cues to generate useful depth proposals. These cues can predict distinguishable motion contours and geometric scene structures. Moreover, a new high-dimensional attention module is proposed to extract a clear global transformation, which effectively suppresses the uncertainty of local descriptors in high-dimensional space, resulting in a more reliable optimization in the learning framework. Experiments demonstrate that the proposed framework outperforms the state-of-the-art on KITTI and Make3D datasets. Jianrong Wang, Xuewei Li 0001, Li Liu 0036 |
ICASSP | 5 |
| 2021 | An Attention Self-Supervised Contrastive Learning Based Three-Stage Model for Hand Shape Feature Representation in Cued SpeechabstractCued Speech (CS) is a communication system for deaf people or hearing impaired people, in which a speaker uses it to aid a lipreader in phonetic level by clarifying potentially ambiguous mouth movements with hand shape and positions.Feature extraction of multi-modal CS is a key step in CS recognition.Recent supervised deep learning based methods suffer from noisy CS data annotations especially for hand shape modality.In this work, we first propose a self-supervised contrastive learning method to learn the feature representation of image without using labels.Secondly, a small amount of manually annotated CS data are used to fine-tune the first module.Thirdly, we present a module, which combines Bi-LSTM and self-attention networks to further learn sequential features with temporal and contextual information.Besides, to enlarge the volume and the diversity of the current limited CS datasets, we build a new British English dataset containing 5 native CS speakers.Evaluation results on both French and British English datasets show that our model achieves over 90% accuracy in hand shape recognition.Significant improvements of 8.75% (for French) and 10.09% (for British English) are achieved in CS phoneme recognition correctness compared with the state-of-the-art. Jianrong Wang, Nan Gu, Mei Yu 0004, Xuewei Li 0001, Qiang Fang 0003, Li Liu 0036 |
Interspeech | 6 |
| 2021 | Cross-Modal Knowledge Distillation Method for Automatic Cued Speech RecognitionabstractCued Speech (CS) is a visual communication system for the deaf or hearing impaired people. It combines lip movements with hand cues to obtain a complete phonetic repertoire. Current deep learning based methods on automatic CS recognition suffer from a common problem, which is the data scarcity. Until now, there are only two public single speaker datasets for French (238 sentences) and British English (97 sentences). In this work, we propose a cross-modal knowledge distillation method with teacher-student structure, which transfers audio speech information to CS to overcome the limited data problem. Firstly, we pretrain a teacher model for CS recognition with a large amount of open source audio speech data, and simultaneously pretrain the feature extractors for lips and hands using CS data. Then, we distill the knowledge from teacher model to the student model with frame-level and sequence-level distillation strategies. Importantly, for frame-level, we exploit multi-task learning to weigh losses automatically, to obtain the balance coefficient. Besides, we establish a five-speaker British English CS dataset for the first time. The proposed method is evaluated on French and British English CS datasets, showing superior CS recognition performance to the state-of-the-art (SOTA) by a large margin. Jianrong Wang, Ziyue Tang, Xuewei Li 0001, Mei Yu 0004, Qiang Fang 0003, Li Liu 0036 |
Interspeech | 6 |
| 2021 | USCL: Pretraining Deep Ultrasound Image Diagnosis Model Through Video Contrastive Representation Learning
Yixiong Chen, Chunhui Zhang 0001, Li Liu 0036, Changfeng Dong, Yongfang Luo |
MICCAI (8) | 3 |
| 2021 | Re-Synchronization Using the Hand Preceding Model for Multi-Modal Fusion in Automatic Continuous Cued Speech RecognitionabstractCued Speech (CS) is an augmented lip reading system complemented by hand coding, and it is very helpful to the deaf people. Automatic CS recognition can help communications between the deaf people and others. Due to the asynchronous nature of lips and hand movements, fusion of them in automatic CS recognition is a challenging problem. In this work, we propose a novel re-synchronization procedure for multi-modal fusion, which aligns the hand features with lips feature. It is realized by delaying hand position and hand shape with their optimal hand preceding time which is derived by investigating the temporal organizations of hand position and hand shape movements in CS. This re-synchronization procedure is incorporated into a practical continuous CS recognition system that combines convolutional neural network (CNN) with multi-stream hidden markov model (MSHMM). A significant improvement of about 4.6% has been achieved retaining 76.6% CS phoneme recognition correctness compared with the state-of-the-art architecture (72.04%), which did not take into account the asynchrony issue of multi-modal fusion in CS. To our knowledge, this is the first work to tackle the asynchronous multi-modal fusion in the automatic continuous CS recognition. Li Liu 0036, Gang Feng 0002, Denis Beautemps, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 1 |
| 2020 | Three-Dimensional Lip Motion Network for Text-Independent Speaker RecognitionabstractLip motion reflects behavior characteristics of speakers, and thus can be used as a new kind of biometrics in speaker recognition. In the literature, lots of works used two-dimensional (2D) lip images to recognize speaker in a text-dependent context. However, 2D lip easily suffers from various face orientations. To this end, in this work, we present a novel end-to-end 3D lip motion Network (3LMNet) by utilizing the sentence-level 3D lip motion (S3DLM) to recognize speakers in both the text-independent and text-dependent contexts. A new regional feedback module (RFM) is proposed to obtain attentions in different lip regions. Besides, prior knowledge of lip motion is investigated to complement RFM, where landmark-level and frame-level features are merged to form a better feature representation. Moreover, we present two methods, i.e., coordinate transformation and face posture correction to pre-process the LSD-AV dataset, which contains 68 speakers and 146 sentences per speaker. The evaluation results on this dataset demonstrate that our proposed 3LMNet is superior to the baseline models, i.e., LSTM, VGG-16 and ResNet-34, and outperforms the state-of-the-art using 2D lip image as well as the 3D face. The code of this work is released at https://github.com/wutong18/Three-Dimensional-Lip-Motion-Network-for-Text-Independent-Speaker-Recognition. Jianrong Wang, Shanyu Wang, Mei Yu 0004, Qiang Fang 0003, Ju Zhang 0001, Li Liu 0036 |
ICPR | 7 |
| 2020 | Semi-Supervised Active Learning for COVID-19 Lung Ultrasound Multi-symptom ClassificationabstractUltrasound (US) is a non-invasive yet effective medical diagnostic imaging technique for the COVID-19 global pandemic. However, due to complex feature behaviors and expensive annotations of US images, it is difficult to apply Artificial Intelligence (AI) assisting approaches for the lung's multi-symptom (multi-label) classification. To overcome these difficulties, we propose a novel semi-supervised Two-Stream Active Learning (TSAL) method to model complicated features and reduce labeling costs in an iterative manner. The core component of TSAL is the multi-label learning mechanism, in which label correlation information is used to design a multi-label margin (MLM) strategy and a confidence validation for automatically selecting informative samples and confident labels. In this framework, a multi-symptom multi-label (MSML) classification network is proposed to learn discriminative features of lung symptoms, and a human-machine interaction (HMI) is exploited to confirm the final annotations that are used to fine-tune MSML. Moreover, a novel lung US dataset named COVID19-LUSMS is built, currently containing 71 clinical patients with 6,836 images sampled from 678 videos. Experimental evaluations show that TSAL can achieve superior performance to the baseline and the state-of-the-art using only 20% data. Qualitatively, visualization of the attention map confirms a good consistency between the model prediction and the clinical knowledge. Lei Liu 0049, Wentao Lei, Li Liu 0036, Yongfang Luo |
ICTAI | 4 |
| 2019 | Automatic Detection of the Temporal Segmentation of Hand Movements in British English Cued Speech
Li Liu 0036, Jianze Li, Gang Feng 0002, Xiao-Ping Zhang 0002 |
INTERSPEECH | 1 |
| 2019 | A Light-Weight Context-Aware Self-Attention Model for Skin Lesion Segmentation
Hao Wu 0039, Jun Sun 0008, Chunjing Yu, Li Liu 0036 |
PRICAI (3) | 5 |
| 2018 | Automatic Temporal Segmentation of Hand Movements for Hand Positions Recognition in French Cued SpeechabstractIn the context of Cued Speech (CS) recognition, the recognition of lips and hand movements is a key task. As we know, a good temporal segmentation is necessary for the supervised recognition system. However, lips and hand streams cannot share the same temporal segmentation since they are not synchronized. In this work, we propose a hand preceding model to predict temporal segmentations of hand movements automatically by exploring the relationship between hand preceding time and the vowel positions in sentences. To evaluate the performance of the proposed method, we apply the hand preceding model to a multi -speakers database. Hand positions recognition is realized with the multi-Gaussian and Long-Short Term Memory (LSTM). The results show that using the predicted temporal segmentation significantly improves the recognition performance compared with that using the audio based segmentation. To the best of our knowledge, this is the first automatic method to predict the temporal segmentation for hand movements only from the audio based segmentation in CS. Li Liu 0036, Gang Feng 0002, Denis Beautemps |
ICASSP | 1 |
| 2018 | Visual Recognition of Continuous Cued Speech Using a Tandem CNN-HMM ApproachabstractInternational audience Li Liu 0036, Thomas Hueber, Gang Feng 0002, Denis Beautemps |
INTERSPEECH | 1 |
| 2017 | Automatic dynamic template tracking of inner lips based on CLNFabstractIn this paper, a novel automatic approach to extract the inner lips contour of speakers without using artifices is proposed. This method is based on a recent facial contour extraction model developed in computer vision, called Constrained Local Neural Field (CLNF), which provides 8 characteristic points (landmarks) defining the inner lips contour. However, directly applied to our visual data including Cued Speech (CS) data, CLNF failed in about 50% of cases. We propose a Modified CLNF to estimate inner lips contour based on original CLNF landmarks. A dynamic template using the first derivative of smoothed luminance variation is explored in this new model. This method gives precise estimation of aperture for inner lips. It is evaluated on 4800 images of three French speakers. The proposed method corrects 95% CLNF errors and total RMSE of one pixel (i.e. 0.05cm in average) is reached, instead of four pixels using original CLNF. Li Liu 0036, Gang Feng 0002, Denis Beautemps |
ICASSP | 1 |