VLDB 2026 Research / reviewers in the wild / expert
Kai Li 0047
dblp:181/2853-47
· DBLP profile ↗
20ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0002-4960-3019ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FGNet: Leveraging Feature-Guided Attention to Refine SAM2 for 3D EM Neuron SegmentationabstractAccurate segmentation of neural structures in Electron Microscopy (EM) images is paramount for neuroscience. However, this task is challenged by intricate morphologies, low signal-to-noise ratios, and scarce annotations, limiting the accuracy and generalization of existing methods. To address these challenges, we seek to leverage the priors learned by visual foundation models on a vast amount of natural images to better tackle this task. Specifically, we propose a novel framework that can effectively transfer knowledge from Segment Anything 2 (SAM2), which is pre-trained on natural images, to the EM domain. We first use SAM2 to extract powerful, general-purpose features. To bridge the domain gap, we introduce a Feature-Guided Attention module that leverages semantic cues from SAM2 to guide a lightweight encoder, the Fine-Grained Encoder (FGE), in focusing on these challenging regions. Finally, a dual-affinity decoder generates both coarse and refined affinity maps. Experimental results demonstrate that our method achieves performance comparable to state-of-the-art (SOTA) approaches with the SAM2 weights frozen. Upon further fine-tuning on EM data, our method significantly outperforms existing SOTA methods. This study validates that transferring representations pre-trained on natural images, when combined with targeted domain-adaptive guidance, can effectively address the specific challenges in neuron segmentation. Zhenghua Li, Hang Chen 0004, Kai Li 0047, Xiaolin Hu 0001 |
AAAI | 4 |
| 2026 | Critical Information Only: A Content Privacy-Preserving Framework for Detecting Audio DeepfakesabstractText-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals. Existing countermeasures largely focus on determining the genuineness of speech based on complete original audio recordings, which however often contain private content. This oversight may refrain deepfake detection from many applications, particularly in scenarios involving sensitive information like business secrets. In this paper, we propose SafeEar, a novel framework that aims to detect deepfake audios without relying on accessing the speech content within. Our key idea is to devise a neural audio codec into a novel decoupling model that well separates the semantic and acoustic information from audio samples, and only use the acoustic information (e.g., prosody and timbre) for deepfake detection. In this way, no semantic content will be exposed to the detector. To overcome the challenge of identifying diverse deepfake audio without semantic clues, we enhance our deepfake detector with real-world augmentation, such as codecs and reverbs. Extensive experiments conducted on five benchmark datasets demonstrate SafeEar's effectiveness in detecting various deepfake techniques with an equal error rate (EER) down to 2.41%. Simultaneously, it shields f ive-language speech content from being deciphered by both machine and human auditory analysis, demonstrated by word error rates (WERs) all above 93.74% and our user study. Furthermore, our benchmark constructed for anti-deepfake and anti-content recovery evaluation helps provide a basis for future research in the realms of audio privacy preservation and deepfake detection. Xinfeng Li, Yifan Zheng 0001, Chen Yan 0001, Kai Li 0047, Chang Zeng, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Time-Frequency-Based Attention Cache Memory Model for Real-Time Speech SeparationabstractExisting causal speech separation models often under perform compared to non-causal models due to difficulties in retaining historical information. To address this, we propose the Time-Frequency Attention Cache Memory (TFACM) model, which effectively captures spatio-temporal relationships through an attention mechanism and cache memory (CM) for historical information storage. In TFACM, an LSTM layer captures frequency-relative positions, while causal modeling is applied to the time dimension using local and global representations. The CM module stores past information, and the causal attention refinement (CAR) module further enhances time-based feature representations for finer granularity. Experimental results showed that TFACM achieved comparable performance to the SOTA TF-GridNet-Causal model, with significantly lower complexity and fewer trainable parameters. For more details, visit the project page: https://anonymous.4open.science/w/TFACM-Page/. Kai Li 0047, Runxuan Yang, Xiaolin Hu 0001 |
ASRU | 2 |
| 2025 | A Fast and Lightweight Model for Causal Audio-Visual Speech SeparationabstractAudio-visual speech separation (AVSS) aims to extract a target speech signal from a mixed signal by leveraging both auditory and visual (lip movement) cues. However, most existing AVSS methods exhibit complex architectures and rely on future context, operating offline, which renders them unsuitable for real-time applications. Inspired by the pipeline of RTFSNet, we propose a novel streaming AVSS model, named Swift-Net, which enhances the causal processing capabilities required for real-time applications. Swift-Net adopts a lightweight visual feature extraction module and an efficient fusion module for audio-visual integration. Additionally, Swift-Net employs Grouped SRUs to integrate historical information across different feature spaces, thereby improving the utilization efficiency of historical information. We further propose a causal transformation template to facilitate the conversion of non-causal AVSS models into causal counterparts. Experiments on three standard benchmark datasets (LRS2, LRS3, and VoxCeleb2) demonstrated that under causal conditions, our proposed Swift-Net exhibited outstanding performance, highlighting the potential of this method for processing speech in complex environments. Wendi Sang, Kai Li 0047, Runxuan Yang, Jianqiang Huang 0002, Xiaolin Hu 0001 |
ECAI | 2 |
| 2025 | Mastering Visual Reinforcement Learning via Positive Unlabeled Policy-Guided Contrast
Zehua Zang, Qirui Ji, Rui Wang 0079, Kai Li 0047, Fuchun Sun 0001 |
ICIC (12) | 4 |
| 2025 | SPMamba: Leveraging Long-Sequence Modeling with State Space Models for Speech SeparationabstractExisting CNN-based speech separation models face local receptive field limitations and cannot effectively capture long time dependencies. Although LSTM and Transformer-based speech separation models can avoid this problem, their high complexity causes them to face the challenge of computational resources and inference efficiency when dealing with long audio. To address this challenge, we introduce an innovative speech separation method called SPMamba. This model builds upon the robust TF-GridNet architecture, replacing its traditional BLSTM modules with bidirectional Mamba modules. These modules effectively model the spatiotemporal relationships between the time and frequency dimensions, allowing SPMamba to capture long-range dependencies with linear computational complexity. Specifically, the bidirectional processing within the Mamba modules enables the model to utilize both past and future contextual information, thereby enhancing separation performance. Extensive experiments were conducted on public datasets, including the WSJ0-2Mix and WHAM! and Libri2Mix, as well as the newly constructed Echo2Mix dataset, demonstrated that SPMamba achieved superior results to previous state-of-the-art (SOTA) models with reduced computational complexity. These findings highlight the effectiveness of SPMamba in addressing the intricate challenges of speech separation in complex environments. The source code for SPMamba is publicly accessible at https://anonymous.4open.science/r/SPMamba-ICME/. Kai Li 0047, Runxuan Yang, Xiaolin Hu 0001 |
ICME | 1 |
| 2025 | EAGLE: An Efficient Global Attention Lesion Segmentation Model for Hepatic Echinococcosis
Jiayan Chen, Kai Li 0047, Yulu Zhao, Jianqiang Huang 0002 |
PRCV (14) | 2 |
| 2025 | GDSR: Global-Detail Integration Through Dual-Branch Network With Wavelet Losses for Remote Sensing Image Super-ResolutionabstractIn recent years, deep neural networks, including Convolutional Neural Networks, Transformers, and State Space Models, have achieved significant progress in Remote Sensing Image (RSI) Super-Resolution (SR). However, existing SR methods typically overlook the complementary relationship between global and local dependencies. These methods either focus on capturing local information or prioritize global information, which results in models that are unable to effectively capture both global and local features simultaneously. Moreover, their computational cost becomes prohibitive when applied to large-scale RSIs. To address these challenges, we introduce the novel application of Receptance Weighted Key Value (RWKV) to RSI-SR, which captures long-range dependencies with linear complexity. To simultaneously model global and local features, we propose the Global-Detail dual-branch structure, GDSR, which performs SR by paralleling RWKV and convolutional operations to handle large-scale RSIs. Furthermore, we introduce the Global-Detail Reconstruction Module (GDRM) as an intermediary between the two branches to bridge their complementary roles. In addition, we propose the Dual-Group Multi-Scale Wavelet Loss, a wavelet-domain constraint mechanism via dual-group subband strategy and cross-resolution frequency alignment for enhanced reconstruction fidelity in RSI-SR. Extensive experiments under two degradation methods on several benchmarks, including AID, UCMerced, and RSSRD-QH, demonstrate that GSDR outperforms the state-of-the-art Transformer-based method HAT by an average of 0.09 dB in PSNR, while using only 63% of its parameters and 51% of its FLOPs, achieving an inference speed 3.2 times faster. Qiwei Zhu, Kai Li 0047, Guojing Zhang, Xiaoying Wang 0002, Jianqiang Huang 0002, Xilai Li |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | SafeEar: Content Privacy-Preserving Audio Deepfake Detection
Xinfeng Li, Kai Li 0047, Yifan Zheng 0001, Chen Yan 0001, Xiaoyu Ji 0001, Wenyuan Xu 0001 |
CCS | 2 |
| 2024 | RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech SeparationabstractAudio-visual speech separation methods aim to integrate different modalities to generate high-quality separated speech, thereby enhancing the performance of downstream tasks such as speech recognition. Most existing state-of-the-art (SOTA) models operate in the time domain. However, their overly simplistic approach to modeling acoustic features often necessitates larger and more computationally intensive models in order to achieve SOTA performance. In this paper, we present a novel time-frequency domain audio-visual speech separation method: Recurrent Time-Frequency Separation Network (RTFS-Net), which applies its algorithms on the complex time-frequency bins yielded by the Short-Time Fourier Transform. We model and capture the time and frequency dimensions of the audio independently using a multi-layered RNN along each dimension. Furthermore, we introduce a unique attention-based fusion technique for the efficient integration of audio and visual information, and a new mask separation approach that takes advantage of the intrinsic spectral nature of the acoustic features for a clearer separation. RTFS-Net outperforms the prior SOTA method in both inference speed and separation quality while reducing the number of parameters by 90% and MACs by 83%. This is the first time-frequency domain audio-visual speech separation method to outperform all contemporary time-domain counterparts. Samuel Pegg, Kai Li 0047, Xiaolin Hu 0001 |
ICLR | 2 |
| 2024 | IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech SeparationabstractRecent research has made significant progress in designing fusion modules for audio-visual speech separation. However, they predominantly focus on multi-modal fusion at a single temporal scale of auditory and visual features without employing selective attention mechanisms, which is in sharp contrast with the brain. To address this, We propose a novel model called intra- and inter-attention network (IIANet), which leverages the attention mechanism for efficient audio-visual feature fusion. IIANet consists of two types of attention blocks: intra-attention (IntraA) and inter-attention (InterA) blocks, where the InterA blocks are distributed at the top, middle and bottom of IIANet. Heavily inspired by the way how human brain selectively focuses on relevant content at various temporal scales, these blocks maintain the ability to learn modality-specific features and enable the extraction of different semantics from audio-visual features. Comprehensive experiments on three standard audio-visual separation benchmarks (LRS2, LRS3, and VoxCeleb2) demonstrate the effectiveness of IIANet, outperforming previous state-of-the-art methods while maintaining comparable inference time. In particular, the fast version of IIANet (IIANet-fast) has only 7% of CTCNet’s MACs and is 40% faster than CTCNet on CPUs while achieving better separation quality, showing the great potential of attention mechanism for efficient and effective multimodal fusion. Kai Li 0047, Runxuan Yang, Fuchun Sun 0001, Xiaolin Hu 0001 |
ICML | 1 |
| 2024 | CMIX: Causal Value Decomposition for Cooperative Multi-Agent Reinforcement LearningabstractValue decomposition plays a pivotal role in ensuring effective credit assignment within Multi-Agent Reinforcement Learning (MARL), particularly in cooperative multi-agent tasks where agents are limited to accessing team rewards only. However, existing methods treat the mixing network as a black box, implicitly assuming that neural networks can autonomously extract important information and achieve rational credit assignment during policy learning. This approach not only lacks interpretability but may also prove inefficient in complex scenarios. To enhance the interpretability and rationality of value decomposition, we propose an innovative approach called “Causal Value Decomposition”(CMIX). CMIX employs causal inference-based models, introducing a set of metrics beyond environmental rewards to enhance robustness and model interpretability. Specifically, CMIX establishes intricate relational structures among agents in complex environments and leverages causal relationships between agents and their surroundings to address the credit assignment challenge in MARL. By employing do-calculus, CMIX accurately measures the impact of each agent's actions on environmental states, precisely determining their contribution to the collective reward. This approach not only enhances the interpretability of existing black-box models but also improves the accuracy of credit assignment in multi-agent systems. Moreover, CMIX exhibits high scalability and complements existing value decomposition techniques. Its effectiveness and scalability have been rigorously tested across various settings, including MPE, LBF, and SMAC environments. Dunqi Yao, Chuxiong Sun, Kai Li 0047, Kaijie Zhou, Rui Wang 0079 |
SMC | 3 |
| 2024 | An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical CircuitsabstractAudio-visual approaches involving visual inputs have laid the foundation for recent progress in speech separation. However, the optimization of the concurrent usage of auditory and visual inputs is still an active research area. Inspired by the cortico-thalamo-cortical circuit, in which the sensory processing mechanisms of different modalities modulate one another via the non-lemniscal sensory thalamus, we propose a novel cortico-thalamo-cortical neural network (CTCNet) for audio-visual speech separation (AVSS). First, the CTCNet learns hierarchical auditory and visual representations in a bottom-up manner in separate auditory and visual subnetworks, mimicking the functions of the auditory and visual cortical areas. Then, inspired by the large number of connections between cortical regions and the thalamus, the model fuses the auditory and visual information in a thalamic subnetwork through top-down connections. Finally, the model transmits this fused information back to the auditory and visual subnetworks, and the above process is repeated several times. The results of experiments on three speech separation benchmark datasets show that CTCNet remarkably outperforms existing AVSS methods with considerably fewer parameters. These results suggest that mimicking the anatomical connectome of the mammalian brain has great potential for advancing the development of deep neural networks. Kai Li 0047, Fenghua Xie, Hang Chen 0004, Kexin Yuan, Xiaolin Hu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | An efficient encoder-decoder architecture with top-down attention for speech separation
Kai Li 0047, Runxuan Yang, Xiaolin Hu 0001 |
ICLR | 1 |
| 2023 | Audio-Visual Speech Separation in Noisy Environments with a Lightweight Iterative Model
Héctor Martel, Julius Richter, Kai Li 0047, Xiaolin Hu 0001, Timo Gerkmann |
INTERSPEECH | 3 |
| 2022 | On the Use of Deep Mask Estimation Module for Neural Source Separation SystemsabstractMost of the recent neural source separation systems rely on a masking-based pipeline where a set of multiplicative masks are estimated from and applied to a signal representation of the input mixture.The estimation of such masks, in almost all network architectures, is done by a single layer followed by an optional nonlinear activation function.However, recent literatures have investigated the use of a deep mask estimation module and observed performance improvement compared to a shallow mask estimation module.In this paper, we analyze the role of such deeper mask estimation module by connecting it to a recently proposed unsupervised source separation method, and empirically show that the deep mask estimation module is an efficient approximation of the so-called overseparation-grouping paradigm with the conventional shallow mask estimation layers. Kai Li 0047, Xiaolin Hu 0001 |
INTERSPEECH | 1 |
| 2022 | Inferring Mechanisms of Auditory Attentional Modulation with Deep Neural NetworksabstractHumans have an exceptional ability to extract specific audio streams of interest in a noisy environment; this is known as the cocktail party effect. It is widely accepted that this ability is related to selective attention, a mental process that enables individuals to focus on a particular object. Evidence suggests that sensory neurons can be modulated by top-down signals transmitted from the prefrontal cortex. However, exactly how the projection of attention signals to the cortex and subcortex influences the cocktail effect is unclear. We constructed computational models to study whether attentional modulation is more effective at earlier or later stages for solving the cocktail party problem along the auditory pathway. We modeled the auditory pathway using deep neural networks (DNNs), which can generate representational neural patterns that resemble the human brain. We constructed a series of DNN models in which the main structures were autoencoders. We then trained these DNNs on a speech separation task derived from the dichotic listening paradigm, a common paradigm to investigate the cocktail party effect. We next analyzed the modulation effects of attention signals during all stages. Our results showed that the attentional modulation effect is more effective at the lower stages of the DNNs. This suggests that the projection of attention signals to lower stages within the auditory pathway plays a more significant role than the higher stages in solving the cocktail party problem. This prediction could be tested using neurophysiological experiments. Ting-Yu Kuo, Yuanda Liao, Kai Li 0047, Xiaolin Hu 0001 |
Neural Comput. | 3 |
| 2021 | Speech Separation Using an Asynchronous Fully Recurrent Convolutional Neural NetworkabstractRecent advances in the design of neural network architectures, in particular those specialized in modeling sequences, have provided significant improvements in speech separation performance. In this work, we propose to use a bio-inspired architecture called Fully Recurrent Convolutional Neural Network (FRCNN) to solve the separation task. This model contains bottom-up, top-down and lateral connections to fuse information processed at various time-scales represented by stages. In contrast to the traditional approach updating stages in parallel, we propose to first update the stages one by one in the bottom-up direction, then fuse information from adjacent stages simultaneously and finally fuse information from all stages to the bottom stage together. Experiments showed that this asynchronous updating scheme achieved significantly better results with much fewer parameters than the traditional synchronous updating scheme on speech separation. In addition, the proposed model achieved competitive or better results with high efficiency as compared to other state-of-the-art approaches on two benchmark datasets. Xiaolin Hu 0001, Kai Li 0047, Jean-Marie Lemercier, Timo Gerkmann |
NeurIPS | 2 |
| 2020 | Survey of single image super-resolution reconstructionabstractImage super‐resolution reconstruction refers to a technique of recovering a high‐resolution (HR) image (or multiple images) from a low‐resolution (LR) degraded image (or multiple images). Due to the breakthrough progress in deep learning in other computer vision tasks, people try to introduce deep neural network and solve the problem of image super‐resolution reconstruction by constructing a deep‐level network for end‐to‐end training. The currently used deep learning models can divide the SISR model into four types: interpolation‐based preprocessing‐based model, original image processing based model, hierarchical feature‐based model, and high‐frequency detail‐based model, or shared the network model. The current challenges for super‐resolution reconstruction are mainly reflected in the actual application process, such as encountering an unknown scaling factor, losing paired LR–HR images, and so on. Kai Li 0047, Shenghao Yang 0003, Runting Dong, Xiaoying Wang 0002, Jianqiang Huang 0002 |
IET Image Process. | 1 |
| 2014 | Augmenting interventional ultrasound using statistical shape model for guiding percutaneous nephrolithotomy: Initial evaluation in pigs
Zhicheng Li 0001, Kai Li 0047, Hai-Lun Zhan, Ken Chen 0005, Ming-Min Chen, Yaoqin Xie, Lei Wang 0029 |
Neurocomputing | 2 |