VLDB 2026 Research / reviewers in the wild / expert
Chenxing Li
dblp:169/4579
· DBLP profile ↗
59ranked-venue papers
18as first author
43since 2021 · last 2026
0000-0002-3997-8212ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 5 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 7 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 first-author · 2 since 2021Computer networks · 5 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization AbilityabstractAchieving zero-shot adversarial robustness without sacrificing generalization remains challenging for foundation models such as CLIP, especially under large adversarial perturbations. Through empirical analyses, we identify three critical yet overlooked issues: (1) Logit margins exhibit a stable offset between small and large adversarial perturbations, suggesting that explicitly adjusting margins could improve robustness against unseen large perturbations. (2) A significant negative correlation exists between logit margin and inter-class semantic similarity, indicating that semantic structures are insufficiently leveraged by existing methods. (3) Existing methods for adjusting text embeddings disrupt the intrinsic semantic consistency established by pre-trained models, undermining generalization capability. Motivated by these findings, we propose a novel Text-Image Mutual Awareness (TIMA) framework, including a Text-Aware Image (TAI) tuning module with an Adaptive Semantic-Aware Margin (ASAM) to explicitly calibrate logit margins, and an Image-Aware Text (IAT) tuning module with Semantic Consistent Minimum Hyperspherical Energy (SC-MHE) to preserve semantic consistency. Comprehensive experiments validate that TIMA significantly outperforms existing approaches by effectively addressing the identified limitations. Fengji Ma, Hei Victor Cheng, Chenxing Li, Li Liu 0036 |
AAAI | 3 |
| 2026 | Cueing Without Gapping: Cuer-Independent Cued Speech Recognition Powered by Cross-Cuer Invariant ModelingabstractAutomatic Cued Speech Recognition (ACSR) is a vital communication system designed to enhance spoken language accessibility for the hearing-impaired by combining lip movements and hand gestures to encode phonemes. Despite its effectiveness, current ACSR methods face significant challenges, including poor generalization to unseen cuers due to the limited scale of CS datasets, which restricts the ability of existing visual encoder to capture cuer-invariant CS visual features. Additionally, previous approaches relying on Connectionist Temporal Classification (CTC) decoding fail to incorporate prior linguistic sequence knowledge, further limiting their performance. To address these issues, we propose a novel Two Auxiliary Modalities guided Cross-cuer Invariant Adaptation method (TACIA), introducing pose and text modalities to help extract cuer-invariant motion and semantic features, thereby improving generalization. In addition, we introduce a Visual-guided Cued Token Prediction (VG-NTP) method, inspired by large language models. This method replaces CTC decoding by incorporating language modeling, leveraging rich linguistic knowledge, including semantics, to address the suboptimal issues present in the CTC decoding process. Extensive experiments demonstrate the superiority of our approach to the state-of-the-art (SOTA) on Chinese and British CS datasets, significantly advancing the accuracy and quality of ACSR systems. Fengji Ma, Chenxing Li, Li Liu 0036 |
AAAI | 2 |
| 2026 | UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech GenerationabstractCued Speech (CS) enhances lipreading via hand coding, offering visual phonemic cues that support precise speech perception for the hearing-impaired. The task of CS Video-to-Speech generation (CSV2S) aims to convert CS videos into intelligible speech signals. Most existing research focuses on CS Recognition (CSR), which transcribes video content into text. Consequently, a common solution for CSV2S is to integrate CSR with a text-to-speech (TTS) system. However, this pipeline relies on text as an intermediate medium, which may lead to error propagation and temporal misalignment between speech and CS video dynamics. In contrast, directly generating audio speech from CS video (direct CSV2S) often suffer from the inherent multimodal complexity and the limited availability of CS data. To address these challenges, we propose UniCUE, the first unified framework for CSV2S that directly generates speech from CS videos without relying on intermediate text. The core innovation of UniCUE lies in integrating a understanding task (CSR) that provides fine-grained CS visual-semantic cues to to guide the speech generation. Specifically, UniCUE incorporates a pose-aware visual processor, a semantic alignment pool that enables precise visual–semantic mapping, and a VisioPhonetic adapter to bridge the understanding and generation tasks within a unified architecture. To support this framework, we construct UniCUE-HI, a large-scale Mandarin CS dataset containing 11,282 videos from 14 cuers, including both hearing-impaired and normal-hearing individuals. Extensive experiments conducted on this dataset demonstrate that UniCUE achieves state-of-the-art (SOTA) performance across multiple evaluation metrics. Jinting Wang, Shan Yang 0001, Chenxing Li, Dong Yu 0001, Li Liu 0036 |
AAAI | 3 |
| 2026 | Audio-Thinker: Guiding Large Audio Language Model When and How to Think via Reinforcement LearningabstractRecent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning utilizing rule-based rewards. However, the explicit reasoning process has not yet yielded substantial benefits for audio question answering, and effectively leveraging deep reasoning remains an open challenge, with LALMs still falling short of achieving human-level auditory-language reasoning. To address these limitations, we propose Audio-Thinker, a reinforcement learning framework designed to enhance the reasoning capabilities of LALMs through improved adaptability, consistency, and effectiveness. Our approach introduces an adaptive think accuracy reward, enabling the model to adjust its reasoning strategies based on task complexity. Furthermore, we incorporate an external reward model to evaluate the overall consistency and quality of the reasoning process, complemented by think-based rewards that assist the model in distinguishing between valid and flawed reasoning paths during training. Experimental results demonstrate that Audio-Thinker models outperform existing reasoning-oriented LALMs across various benchmark tasks, exhibiting superior reasoning and generalization capabilities. Chenxing Li, Wenfu Wang, Hao Zhang 0112, Hualei Wang, Meng Yu 0003, Dong Yu 0001 |
AAAI | 2 |
| 2026 | OpenST: Toward open-set source tracing for neural codec deepfake audio
Yuankun Xie, Ruibo Fu, Zhengqi Wen, Songjun Cao, Chenxing Li, Haonan Cheng, Long Ye |
Neurocomputing | 8 |
| 2026 | Lifelong Learning of Large Language Model Based Agents: A RoadmapabstractLifelong learning, also known as continual or incremental learning, is a crucial component for advancing Artificial General Intelligence (AGI) by enabling systems to continuously adapt in dynamic environments. While large language models (LLMs) have demonstrated impressive capabilities in natural language processing, existing LLM agents are typically designed for static systems and lack the ability to adapt over time in response to new challenges. This survey is the first to systematically summarize the potential techniques for incorporating lifelong learning into LLM-based agents. We categorize the core components of these agents into three modules: the perception module for multimodal input integration, the memory module for storing and retrieving evolving knowledge, and the action module for grounded interactions with the dynamic environment. We highlight how these pillars collectively enable continuous adaptation, mitigate catastrophic forgetting, and improve long-term performance. This survey provides a roadmap for researchers and practitioners working to develop lifelong learning capabilities in LLM agents, offering insights into emerging trends, evaluation metrics, and application scenarios. Junhao Zheng, Chengming Shi, Xidi Cai, Qiuke Li, Duzhen Zhang, Chenxing Li, Dong Yu 0001, Qianli Ma 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Enhancing Multimodal Continual Instruction Tuning with BranchLoRAabstractMultimodal Continual Instruction Tuning (MCIT) aims to finetune Multimodal Large Language Models (MLLMs) to continually align with human intent across sequential tasks. Existing approaches often rely on the Mixture-of-Experts (MoE) LoRA framework to preserve previous instruction alignments. However, these methods are prone to Catastrophic Forgetting (CF), as they aggregate all LoRA blocks via simple summation, which compromises performance over time. In this paper, we identify a critical parameter inefficiency in the MoELoRA framework within the MCIT context. Based on this insight, we propose BranchLoRA, an asymmetric framework to enhance both efficiency and performance. To mitigate CF, we introduce a flexible tuning-freezing mechanism within BranchLoRA, enabling branches to specialize in intra-task knowledge while fostering inter-task collaboration. Moreover, we incrementally incorporate task-specific routers to ensure an optimal branch distribution over time, rather than favoring the most recent task. To streamline inference, we introduce a task selector that automatically routes test inputs to the appropriate router without requiring task identity. Extensive experiments on the latest MCIT benchmark demonstrate that BranchLoRA significantly outperforms MoELoRA and maintains its superiority across various MLLM sizes. Duzhen Zhang, Yong Ren 0006, Zhongzhi Li, Yahan Yu, Jiahua Dong 0001, Chenxing Li, Zhilong Ji, Jinfeng Bai |
ACL (1) | 6 |
| 2025 | Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio GenerationabstractMainstream Text-to-Audio (TTA) models that rely on Mel-spectrograms often struggle to generate audio with rich content, leading to blurred or incoherent outputs. This stems from an inability to model intricate spectral details and textures. We investigate the role of U-Net components in generation and find that high-frequency components in skip-connections and the backbone are crucial for texture, while low-frequency backbone components are vital for the denoising process. Based on this, we propose “Mel-Refine,” a plug-and-play approach that enhances Mel-spectrogram quality by adjusting component weights during inference. Our method requires no additional training or fine-tuning and is fully compatible with any diffusion-based TTA architecture. Experiments show that Mel-Refine boosts the performance of the latest TTA model, Tango2, by $25 \%$, demonstrating its effectiveness. Hongming Guo, Ruibo Fu, Yizhong Geng, Shuchen Shi, Tao Wang 0074, Chunyu Qiang, Ya Li 0001, Zhengqi Wen, Xuefei Liu, Chenxing Li |
ASRU | 11 |
| 2025 | Beyond Discrete Environments: Benchmarking Regret-Based Automatic Curriculum Learning in MuJoCo
Chin-Jui Chang, Chenxing Li, Jan R. Seyler, Shahram Eivazi |
ICAART (2) | 2 |
| 2025 | MIRSim-RL: A Simulated Mobile Industry Robot Platform and Benchmarks for Reinforcement Learning
Qingkai Li, Zijian Ma, Chenxing Li, Yinlong Liu, Tobias Recker, Daniel Brauchle, Jan R. Seyler, Mingguo Zhao, Shahram Eivazi |
ICAART (1) | 3 |
| 2025 | DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-SpeechabstractIn recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models. Ruibo Fu, Zhengqi Wen, Tao Wang 0074, Chunyu Qiang, Jianhua Tao 0001, Chenxing Li, Shuchen Shi, Yuankun Xie, Xuefei Liu, Guanjun Li |
ICASSP | 7 |
| 2025 | STA-V2A: Video-to-Audio Generation with Semantic and Temporal AlignmentabstractVisual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader applications. In this paper, we propose Semantic and Temporal Aligned Video-to-Audio (STA-V2A), an approach that enhances audio generation from videos by extracting both local temporal and global semantic video features and combining these refined video features with text as cross-modal guidance. To address the issue of information redundancy in videos, we propose an onset prediction pretext task for local temporal feature extraction and an attentive pooling module for global semantic feature extraction. To supplement the insufficient semantic information in videos, we propose a Latent Diffusion Model with Text-to-Audio priors initialization and cross-modal guidance. We also introduce Audio-Audio Align, a new metric to assess audio-temporal alignment. Subjective and objective metrics demonstrate that our method surpasses existing Video-to-Audio models in generating audio with better quality, semantic consistency, and temporal alignment. The ablation experiment validated the effectiveness of each module. Audio samples are available at https://y-ren16.github.io/STAV2A. Yong Ren 0006, Chenxing Li, Manjie Xu, Rilin Chen, Dong Yu 0001 |
ICASSP | 2 |
| 2025 | Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0abstractSpeech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the previously proposed fusion methods require fine-tuning the pretrained models, resulting in excessively long training times and hindering model iteration when facing new speech synthesis technology. To address this issue, this paper proposes a feature fusion method based on the Mixture of Experts, which extracts and integrates features relevant to fake audio detection from layer features, guided by a gating network based on the last layer feature, while freezing the pretrained model. Experiments conducted on the ASVspoof2019 and ASVspoof2021 datasets demonstrate that the proposed method achieves competitive performance compared to those requiring fine-tuning. Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Yuankun Xie, Shuchen Shi, Chenxing Li, Xuefei Liu, Guanjun Li |
ICASSP | 11 |
| 2025 | EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Jiarui Hai, Yong Xu 0004, Hao Zhang 0112, Chenxing Li, Helin Wang, Mounya Elhilali, Dong Yu 0001 |
INTERSPEECH | 4 |
| 2025 | Video-to-Audio Generation with Fine-grained Temporal Semantics
Chenxing Li, Rilin Chen, Dong Yu 0001 |
INTERSPEECH | 3 |
| 2025 | Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model
Yong Ren 0006, Chenxing Li, Duzhen Zhang, Yujie Chen 0006, Manjie Xu, Ruibo Fu, Shan Yang 0001, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2025 | Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning
Chenxing Li, Yong Ren 0006, Yujie Chen 0006, Ruibo Fu, Shan Yang 0001, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2025 | Towards Diverse and Efficient Audio Captioning via Diffusion Models
Manjie Xu, Chenxing Li, Yong Ren 0006, Ruibo Fu, Dong Yu 0001 |
INTERSPEECH | 2 |
| 2025 | Spatial-Spectral Graph Convolutional Network for Hyperspectral Target DetectionabstractDeep learning-based hyperspectral target detection methods commonly face challenges such as insufficient target samples and inadequate use of spatial context. To address these limitations, we propose a novel hyperspectral target detection approach leveraging spatial-spectral graph convolutional networks. First, we introduce an innovative sample augmentation strategy utilizing pre-detection and target implantation, which can simulate different backgrounds around the target and expand the target sample, thereby enhancing the representation ability of the model. Next, a graph-based representation strategy is proposed to integrate spatial and spectral information. Finally, we develop four specialized graph network detectors (GCND, GATD, GCN-GAT1, and GCN-GAT2) with structural optimizations involving multi-scale feature fusion and dynamic neighborhood adjustments. Extensive experiments demonstrate the superior detection performance of our method compared to existing techniques, highlighting its practical significance. Chenxing Li, Dehui Zhu, Chen Wu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Information-theoretic complementary prompts for improved continual text classification
Duzhen Zhang, Yong Ren 0006, Chenxing Li, Dong Yu 0001, Tielin Zhang |
Neural Networks | 3 |
| 2025 | HEMVM: A Heterogeneous Blockchain Framework for Interoperable Virtual MachinesabstractThis paper introduces HEMVM, an innovative heterogeneous blockchain framework that seamlessly integrates diverse virtual machines (VMs), including the Ethereum Virtual Machine (EVM) and the Move Virtual Machine (MoveVM), into a unified system. This integration facilitates interoperability while retaining compatibility with existing Ethereum and Move toolchains by preserving high-level language constructs. HEMVM's unique cross-VM operations allow users to interact with contracts across various VMs using any wallet software, effectively resolving the fragmentation in user experience caused by differing VM designs. Our experimental results demonstrate that HEMVM is both fast and efficient, incurring minimal overhead (less than 4.4 %) for intra-VM transactions and achieving up to 9300 TPS for cross-VM transactions. Our results also show that the cross-VM operations in HEMVM are sufficiently expressive to support complex decentralized finance interactions across multiple VMs. Finally, the parallelized prototype of HEMVM shows performance improvements up to 44.8 % compared to the sequential version of HEMVM under workloads with mixed transaction types. Vladyslav Nekriach, Sidi Mohamed Beillahi, Chenxing Li, Peilun Li, Ming Wu 0007, Andreas G. Veneris, Fan Long |
Proc. ACM Program. Lang. | 3 |
| 2025 | HyperSIGMA: Hyperspectral Intelligence Comprehension Foundation ModelabstractAccurate hyperspectral image (HSI) interpretation is critical for providing valuable insights into various earth observation-related applications such as urban planning, precision agriculture, and environmental monitoring. However, existing HSI processing methods are predominantly task-specific and scene-dependent, which severely limits their ability to transfer knowledge across tasks and scenes, thereby reducing the practicality in real-world applications. To address these challenges, we present HyperSIGMA, a vision transformer-based foundation model that unifies HSI interpretation across tasks and scenes, scalable to over one billion parameters. To overcome the spectral and spatial redundancy inherent in HSIs, we introduce a novel sparse sampling attention (SSA) mechanism, which effectively promotes the learning of diverse contextual features and serves as the basic block of HyperSIGMA. HyperSIGMA integrates spatial and spectral features using a specially designed spectral enhancement module. In addition, we construct a large-scale hyperspectral dataset, HyperGlobal-450K, for pre-training, which contains about 450 K hyperspectral images, significantly surpassing existing datasets in scale. Extensive experiments on various high-level and low-level HSI tasks demonstrate HyperSIGMA's versatility and superior representational capability compared to current state-of-the-art methods. Moreover, HyperSIGMA shows significant advantages in scalability, robustness, cross-modal transferring capability, real-world applicability, and computational efficiency. Di Wang 0023, Meiqi Hu, Yuchun Miao, Jiaqi Yang 0005, Yichu Xu, Xiaolei Qin, Jiaqi Ma 0002, Chenxing Li, Chuan Fu, Hongruixuan Chen, Chengxi Han, Naoto Yokoya, Jing Zhang 0037, Minqiang Xu, Lefei Zhang, Chen Wu 0003, Bo Du 0001, Dacheng Tao, Liangpei Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2025 | Brain-Inspired Video Quality Assessment via Visual-EEG Feature AlignmentabstractVideo quality assessment (VQA) is crucial in applications such as video calls, real-time meetings, and surveillance, where video quality directly impacts user experience greatly. Traditional objective methods like SSIM and PSNR fail to capture the subjective perception of video quality, while subjective Quality of Experience (QoE) assessment metrics like Mean Opinion Score (MOS) are not scalable for large-scale automated VQA tasks. To overcome these limitations, deep learning approaches have emerged, but mostly focusing only on a single video modality, extracting low-level visual features such as color and texture. Recently, electroencephalography (EEG) has been shown to align with users' subjective experiences, offering valuable insights into neural responses to visual content. Hence, in this letter, we propose a brain-inspired deep learning framework for VQA that aligns EEG and video features. We build a video distortion dataset annotated with both MOS and EEG signals to analyze the impact of video distortions on EEG responses and subjective ratings. We then employ an adaptive EEG feature learning network to extract EEG features linked to video distortions, and propose a video quality prediction network that aligns both video and EEG features using a three-stage training strategy. Our method outperforms existing techniques, showing strong alignment with human subjective ratings. Experimental results validate the effectiveness of EEG in enhancing VQA with a more human-centric approach. Shuzhan Hu, Chenxing Li, Yiping Duan, Xiaoming Tao 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Branchformer-Based TDNN for Automatic Speaker VerificationabstractCurrent speaker verification techniques heavily rely on the utilization of neural networks to extract accurate and discriminative speaker representations. In this paper, we present Branchformer based TDNN (B-TDNN), a novel architecture for extracting speaker embeddings by capturing both global and local context within each computing unit. The proposed B-TDNN combines the branchformer and traditional TDNN architecture to effectively capture contextual information. Additionally, our research demonstrates the validity of the smaller model, emphasizing its capability to attain exceptional results even with fewer parameters. To further enhance the efficiency of the model, a Branch Auxiliary Training (BAT) method is introduced, that is, jointly training two branches while using only the more critical branch during inference. The BAT method competently decreases the parameter count of the model while ensuring that the performance remains uncompromised. Experimental results showcase B-TDNN sets a new benchmark in speaker verification performance, delivering state-of-the-art results with an impressive Equal Error Rate (EER) of 0.66% on the VoxCeleb1 trial file. Chenxing Li |
ICASSP | 2 |
| 2024 | Jointly Recognizing Speech and Singing Voices Based on Multi-Task Audio Source SeparationabstractIn short video and live broadcasts, speech, singing voice, and background music often overlap and obscure each other. This complexity creates difficulties in structuring and recognizing the audio content, which may impair subsequent ASR and music understanding applications. This paper proposes a multi-task audio source separation (MTASS) based ASR model called JRSV, which Jointly Recognizes Speech and singing Voices. Specifically, the MTASS module separates the mixed audio into distinct speech and singing voice tracks while removing background music. The CTC/attention hybrid recognition module recognizes both tracks. Online distillation is proposed to improve the robustness of recognition further. To evaluate the proposed methods, a benchmark dataset is constructed and released. Experimental results demonstrate that JRSV can significantly improve recognition accuracy on each track of the mixed audio. Ye Bai 0001, Chenxing Li, Hao Li 0078 |
ICME | 2 |
| 2024 | Prompt-guided Precise Audio Editing with Diffusion ModelsabstractAudio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as **PPAE**, which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely training-free. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks. Manjie Xu, Chenxing Li, Duzhen Zhang, Dan Su 0002, Dong Yu 0001 |
ICML | 2 |
| 2024 | Global Overcomplete Dictionary-Based Sparse and Nonnegative Collaborative Representation for Hyperspectral Target DetectionabstractThe combined sparse and collaborative representation-based algorithm is one of the most effective methods among hyperspectral target detection methods based on representation and dictionary learning. It encourages target atoms to compete with each other and background atoms to collaborate in the representation. However, this method suffers from several drawbacks. In sparse representation, an overcomplete dictionary is necessary, whereas, in collaborative representation, non-negative coefficients are required. Besides, the local dual window approach may result in impure background dictionaries obtained from the outer window. To address these issues, we propose a novel approach for hyperspectral target detection, referred to as the global overcomplete dictionary-based sparse and nonnegative collaborative representation (GODSNCR) detector. First, a hierarchical density clustering algorithm is used to complete the dictionary atom extraction to construct a joint overcomplete dictionary to satisfy the dictionary overcompleteness problem required for sparse representation. Second, a nonnegative constraint on the coefficient matrix and a “sum to one” constraint for the joint representation are incorporated to make it more consistent with the physical meaning. Finally, the limitation of the local dual window approach is overcome by substituting the local background dictionary with a global background dictionary. Through the aforementioned strategies, we can use a joint overcomplete dictionary for achieving the sparse representation of targets and utilize a global background dictionary for the collaborative representation of background, the final detection results are obtained by calculating the residuals. The experimental results clearly demonstrate that the proposed algorithm has significant improvement in detection accuracy and strong robustness compared to other typical representation-based hyperspectral target detection methods. Our model will be available at https://github.com/Chenxing-Li/GODSNCR. Chenxing Li, Dehui Zhu, Chen Wu 0003, Bo Du 0001, Liangpei Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | LVMT: An Efficient Authenticated Storage for BlockchainabstractAuthenticated storage access is the performance bottleneck of a blockchain, because each access can be amplified to potentially O (log n ) disk I/O operations in the standard Merkle Patricia Trie (MPT) storage structure. In this article, we propose a multi-Layer Versioned Multipoint Trie (LVMT), a novel high-performance blockchain storage with significantly reduced I/O amplifications. LVMT uses the authenticated multipoint evaluation tree vector commitment protocol to update commitment proofs in constant time. LVMT adopts a multi-layer design to support unlimited key–value pairs and stores version numbers instead of value hashes to avoid costly elliptic curve multiplication operations. In our experiment, LVMT outperforms the MPT in real Ethereum traces, delivering read and write operations 6× faster. It also boosts blockchain system execution throughput by up to 2.7×. Chenxing Li, Sidi Mohamed Beillahi, Guang Yang 0020, Ming Wu 0007, Wei Xu 0005, Fan Long |
ACM Trans. Storage | 1 |
| 2023 | Accelerate Training of Reinforcement Learning Agent by Utilization of Current and Previous Experience
Chenxing Li, Yinlong Liu, Zhenshan Bing, Fabian Schreier, Jan R. Seyler, Shahram Eivazi |
ICAART (3) | 1 |
| 2023 | Integration of Efficient Deep Q-Network Techniques Into QT-Opt Reinforcement Learning Structure
Shudao Wei, Chenxing Li, Jan R. Seyler, Shahram Eivazi |
ICAART (3) | 2 |
| 2023 | Image-driven Audio-visual Universal Source Separation
Chenxing Li, Ye Bai 0001, Feng Deng |
INTERSPEECH | 1 |
| 2023 | LVMT: An Efficient Authenticated Storage for Blockchain
Chenxing Li, Sidi Mohamed Beillahi, Guang Yang 0020, Ming Wu 0007, Wei Xu 0005, Fan Long |
OSDI | 1 |
| 2023 | Interference Suppression for RIS-Assisted Multicast CommunicationsabstractThis paper focuses on interference suppression in Reconfigurable Intelligent Surface (RIS)-assisted multicast communication, where a multi-antennas base station (BS) transmits identical messages to a group of single-antenna users, while a multi-antennas jammer simultaneously sends interference signals to the users. A RIS consisting of many reflectors is deployed near BS to enhance the multicast channel and suppress the interference channel by carefully designing reflection coefficients. Under such settings, a signal to interference plus noise ratio maximization problem is formulated to maximize the achievable rate by designing feasible reflection coefficients at the RIS and a beamforming vector at the BS. Then, an alternating optimization combined with semidefinite relaxation and bisection search over a series of the solutions of feasibility problems is adopted to iteratively optimize the reflection coefficients and beamforming vector separately. At last, simulation results show the RIS can significantly improve the multicast rate with interference. Chenxing Li, Linsong Du |
PIMRC | 2 |
| 2023 | Adjustable Dielectric Resonator Antenna With Parasitic Elements for 5G SAGOI-IoT ApplicationsabstractWith the development of modern communication and the Internet of Things (IoT) in the need to provide seamless interconnection between heterogeneous devices, we design an adjustable-distance resonant antenna suitable for 5G communication. The antenna meets the multiangle radiation requirements of the antenna for flexible positioning of IoT devices and conforms to the concept of green Internet basic hardware, with low energy and small size. In our design, three parasitic elements will couple with the higher order modes through the slot-hole excitation of a higher order mode dielectric resonator antenna with a dielectric constant of 10. By controlling the distance between the three parasitic elements and changing the capacitor at their terminals, radiation variation in multiple directions can be achieved. The proposed model focuses on the relationships among the three element distances and the effect of the third parasitic element on the radiation angle. Good results were obtained for gain, bandwidth, and radiation angle. Through simulation, the dielectric resonator antenna works successfully in the 15-GHz frequency band. Thus, the antenna array can be from −34° to 34° in the horizontal direction and from 0° to −36° in the vertical direction. At the same time, the gains are kept at a good level and the bandwidths are greater than 2 GHz. Compared with other dielectric resonant antennas, the parameters of this antenna are not reduced, and the flexibility of the radiation direction angles is increased. These evaluation parameters are considered ideal conditions for device-to-device communication in 5G IoT applications. Chenxing Li, Yiping Duan, Xiaoming Tao 0001 |
IEEE Internet Things J. | 1 |
| 2022 | EAD-Conformer: a Conformer-Based Encoder-Attention-Decoder-Network for Multi-Task Audio Source SeparationabstractIn this paper, we propose a Conformer-based network to improve the performance of multi-task audio source separation. This network, named EAD-Conformer, employs Conformer blocks to capture both local and global information, and an encoder-attention-decoder manner encourages the network to perform attentive modeling based on different sources. Specifically, EAD-Conformer first parses out the feature representations from the mixture by a Conformer-based encoder. Then, an attention module extracts selective information for each track and bridges encoder and decoders. Finally, three decoders respectively process attentive features and generate output masks for different sources. In addition, the proposed discriminate loss further enlarges the distance between different sources. Experiments demonstrate the effectiveness of EAD-Conformer, which achieves 13.37 dB, 11.41 dB, 10.56 dB signal-to-distortion ratio improvement on speech, music, noise track, respectively, and shows advantages over several well-known models. Chenxing Li, Feng Deng, Zhongyuan Wang 0006 |
ICASSP | 1 |
| 2022 | Conformer Space Neural Architecture Search for Multi-Task Audio Separation
Shun Lu 0001, Chenxing Li, Jianchao Tan, Feng Deng, Chengru Song |
INTERSPEECH | 4 |
| 2022 | WA-Transformer: Window Attention-based Transformer with Two-stage Strategy for Multi-task Audio Source Separation
Chenxing Li, Feng Deng, Shun Lu 0001, Jianchao Tan, Chengru Song |
INTERSPEECH | 2 |
| 2022 | Image Generation from Scene Graph with Object EdgesabstractSignificant progress has been made on methods for generating images from structured semantic descriptions, but the generated images only retain semantic information, and the appearance of objects cannot be constrained and effectively represented. Therefore, we propose a scene graph structure image generation method assisted by object edge information. Our model uses two graph convolution neural networks(GCN) to process scene graphs and obtains object features as well as relation features which aggregate related information. The object bounding boxes are predicted by a method a decoupling the size and position. Where auxiliary models are added to coordinate with segmentation mask network training. Our experiments show that the introduction of object edges provides clearer object appearance information for image generation, which can constrain object shapes and improve image quality greatly. Finally, the cascaded refinement network is used to generate images. Additionally, compared with other appearance features, such as object slices, edge information occupies a smaller quantity of data, which greatly improves the image quality with less increase in the input information. This feature also benefits semantic communication systems. A large number of experiments show that our method is significantly superior to the latest Sg2im method when evaluated on Visual Genome datasets. Chenxing Li, Yiping Duan, Qiyuan Du, Chengkang Pan, Guangyi Liu 0001, Xiaoming Tao 0001 |
VTC Fall | 1 |
| 2022 | Capacity Characterization for Reconfigurable Intelligent Surfaces Assisted Wireless Communications With InterfererabstractThe reconfigurable intelligent surface (RIS), which consists of many low-cost reflecting elements, can enhance the reception of the desired signal and suppress the interference. Considering the above advantages, this paper investigates the RIS-assisted wireless communication in the presence of an interferer, where a receiver receives the desired signal and interference from a transmitter and an interferer, respectively, and the RIS is deployed to assist this system. First, an interference-limited channel capacity maximization problem is formulated, and a numerical algorithm based on the fraction programming and optimality conditions is proposed to obtain the optimal solution. Then, this paper shows the upper boundary of the maximal capacity with interference and discusses what kind of channel conditions can achieve the upper boundary and the corresponding optimal phase shifts. Next, this paper analyzes the asymptotic behaviors of the maximal capacity in the scenarios that some or all of the number of reflecting elements, antennas at the transmitter, and interferer go to infinity. Finally, this paper generalizes the above results to the multiple interferers scenario. Linsong Du, Qingpeng Liang, Chenxing Li, Youxi Tang |
IEEE Trans. Commun. | 4 |
| 2021 | Multi-Task Audio Source SeparationabstractThe audio source separation tasks, such as speech enhancement, speech separation, and music source separation, have achieved impressive performance in recent studies. The powerful modeling capabilities of deep neural networks give us hope for more challenging tasks. This paper launches a new multi-task audio source separation (MTASS) challenge to separate the speech, music, and noise signals from the monaural mixture. First, we introduce the details of this task and generate a dataset of mixtures containing speech, music, and background noises. Then, we propose an MTASS model in the complex domain to fully utilize the differences in spectral characteristics of the three audio signals. In detail, the proposed model follows a two-stage pipeline, which separates the three types of audio signals and then performs signal compensation separately. After comparing different training targets, the complex ratio mask is selected as a more suitable target for the MTASS. The experimental results also indicate that the residual signal compensation module helps to recover the signals further. The proposed model shows significant advantages in separation performance over several well-known separation models. Chenxing Li, Feng Deng |
ASRU | 2 |
| 2021 | Effect of Non-Resolvable Multipath on Full-Duplex Self-Interference CancellationabstractIn full-duplex multipath self-interference (SI), each of the resolvable multipath SI components consists of a group of non-resolvable (NR) components with similar delays. Since the transceiver cannot distinguish the NR components in reality, the canceller can only generate the approximate constructed SI to imperfectly cancel the received SI. This leads to residual SI remaining. This paper analyzes the effect of the NR components on the SI cancellation, and proposes a practical scheme to eliminate the residual SI caused by the NR components. First, a closed-form expression of the residual SI is obtained. Next, the average power of the residual SI and the SI cancellation ratio are derived to characterize the effect of NR components on SI cancellation in the multipath SI channel. Last, a switch scheme is proposed to eliminate the residual SI. Linsong Du, Ying Liu 0013, Chenxing Li, Youxi Tang |
GLOBECOM | 3 |
| 2021 | Speaker and Direction Inferred Dual-Channel Speech SeparationabstractMost speech separation methods, trying to separate all channel sources simultaneously, are still far from having enough generalization capabilities for real scenarios where the number of input sounds is usually uncertain and even dynamic. In this work, we employ ideas from auditory attention with two ears and propose a speaker and direction inferred speech separation network (dubbed SDNet) to solve the cocktail party problem. Specifically, our SDNet first parses out the respective perceptual representations with their speaker and direction characteristics from the mixture of the scene in a sequential manner. Then, the perceptual representations are utilized to attend to each corresponding speech. Our model generates more precise perceptual representations with the help of spatial features and successfully deals with the problem of the unknown number of sources and the selection of outputs. The experiments on standard fully-overlapped speech separation benchmarks, WSJ0-2mix, WSJ0-3mix, and WSJ0-2&3mix, show the effectiveness, and our method achieves SDR improvements of 25.31 dB, 17.26 dB, and 21.56 dB under anechoic settings. Our codes will be released at https://github.com/aispeech-lab/SDNet. Chenxing Li, Jiaming Xu 0001, Nima Mesgarani, Bo Xu 0002 |
ICASSP | 1 |
| 2021 | One-Shot Voice Conversion Based on Speaker Aware ModuleabstractVoice conversion (VC) is a task to convert the voice of speech while preserving its linguistic content. Although several methods have been proposed to enable VC with non-parallel data, it is still difficult to model the voice without a great number of data or an adaptive process. In this paper, we propose a speaker-aware voice conversion (SAVC) system realizing one-shot voice conversion without an adaptation stage. The SAVC utilizes a speaker aware module (SAM) to disentangle speaker embeddings. The SAM comprises a dynamic reference encoder, a static speaker knowledge block (SKB), and a multi-head attention layer. The reference encoder is used to compress a variable-length utterance to a fixed-length vector, the SKB is made up of pre-extraction x-vectors, and the multi-head attention layer is designed to generate weighted combined speaker embeddings. Subsequently, phonetic pos- teriorgrams (PPGs) as context encoding are concatenated with speaker embeddings and sent to the decoder module for generating acoustic features. Experimental results on the Aishell-1 corpus show that the proposed method can improve speaker similarity and converted utterances' speech quality. Hao Che, Chenxing Li, Zhongyuan Wang 0006 |
ICASSP | 4 |
| 2020 | Shrec: bandwidth-efficient transaction relay in high-throughput blockchain systemsabstractThe success of Bitcoin and Ethereum has attracted many efforts to build high-throughput blockchain systems. This paper focuses on transaction dissemination --- a rather overlooked issue in these systems. We argue that efficient transaction dissemination is the key for a blockchain system to sustain at high-throughput --- usually thousands of transactions per second --- and the existing solutions fell short at doing so. Chenxing Li, Peilun Li, Ming Wu 0007, Dong Zhou 0006, Fan Long |
SoCC | 2 |
| 2020 | A Decentralized Blockchain with High Throughput and Fast Confirmation
Chenxing Li, Peilun Li, Dong Zhou 0006, Ming Wu 0007, Guang Yang 0020, Wei Xu 0005, Fan Long, Andrew Chi-Chih Yao |
USENIX ATC | 1 |
| 2020 | Secure multiparty computation for privacy-preserving drug discoveryabstractMOTIVATION: Quantitative structure-activity relationship (QSAR) and drug-target interaction (DTI) prediction are both commonly used in drug discovery. Collaboration among pharmaceutical institutions can lead to better performance in both QSAR and DTI prediction. However, the drug-related data privacy and intellectual property issues have become a noticeable hindrance for inter-institutional collaboration in drug discovery. RESULTS: We have developed two novel algorithms under secure multiparty computation (MPC), including QSARMPC and DTIMPC, which enable pharmaceutical institutions to achieve high-quality collaboration to advance drug discovery without divulging private drug-related information. QSARMPC, a neural network model under MPC, displays good scalability and performance and is feasible for privacy-preserving collaboration on large-scale QSAR prediction. DTIMPC integrates drug-related heterogeneous network data and accurately predicts novel DTIs, while keeping the drug information confidential. Under several experimental settings that reflect the situations in real drug discovery scenarios, we have demonstrated that DTIMPC possesses significant performance improvement over the baseline methods, generates novel DTI predictions with supporting evidence from the literature and shows the feasible scalability to handle growing DTI data. All these results indicate that QSARMPC and DTIMPC can provide practically useful tools for advancing privacy-preserving drug discovery. AVAILABILITY AND IMPLEMENTATION: The source codes of QSARMPC and DTIMPC are available on the GitHub: https://github.com/rongma6/QSARMPC_DTIMPC.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yi Li 0005, Chenxing Li, Fangping Wan, Hailin Hu 0002, Wei Xu 0005, Jianyang Zeng 0001 |
Bioinform. | 3 |
| 2020 | An Auxiliary Antenna Based Inter-User Interference Mitigation Approach in Full-Duplex Wireless Networks
Fei Wu 0013, Chenxing Li, Jiafan Wang 0003, Xiangyin Zhang |
Mob. Networks Appl. | 2 |
| 2019 | Digital Self-Interference Cancellation in the Presence of Phase Noise for Full-Duplex CommunicationsabstractIn wireless full-duplex communication systems, selfinterference (SI) cancellation is a key technique to ensure simultaneous transmission and reception on the same carrier frequency. During the SI cancellation, it has been shown that the oscillator phase noise dramatically affects the cancellation capability and thus needs consideration and suppression. In this paper, a novel time-domain digital SI cancellation approach is proposed to cancel the transmitter and receiver phase noises along with the incoming SI signal, without estimating either the SI channel response or the phase noise associated coefficients. For proofof-concept evaluation, simulations are performed on orthogonal frequency division multiplexing (OFDM)-modulated signals of 20 MHz bandwidth with 16 QAM constellations to verify the effectiveness of the proposed method. It is demonstrated that the proposed approach provides at least 3 dB improvement on the SI cancellation capability and at least 5 dB increase on the interference tolerance over the conventional cancellation approach that uses the frequency domain SI channel estimation. Chenxing Li, Ying Liu 0013, Youxi Tang |
PIMRC | 2 |
| 2018 | CBLDNN-Based Speaker-Independent Speech Separation Via Generative Adversarial TrainingabstractIn this paper, we propose a speaker-independent multi-speaker monaural speech separation system (CBLDNN-GAT) based on convolutional, bidirectional long short-term memory, deep feedforward neural network (CBLDNN) with generative adversarial training (GAT). Our system aims at obtaining better speech quality instead of only minimizing a mean square error (MSE). In the initial phase, we utilize log-mel filterbank and pitch features to warm up our CBLDNN in a multi-task manner. Thus, the information that contributes to separating speech and improving speech quality is integrated into the model. We execute GAT throughout the training, which makes the separated speech indistinguishable from the real one. We evaluate CBLDNN-GAT on WSJ0-2mix dataset. The experimental results show that the proposed model achieves 11.0d-B signal-to-distortion ratio (SDR) improvement, which is the new state-of-the-art result. Chenxing Li, Bo Xu 0002 |
ICASSP | 1 |
| 2018 | Compression of Acoustic Model via Knowledge Distillation and PruningabstractRecently, the performance of speech recognition system based on neural network has been greatly improved. Arguably, this huge improvement can be mainly attributed to deeper and wider layers. These systems are more difficult to be deployed on the embedded devices due to their large size and high computational complexity. To address these issues, we propose a method to compress deep feed-forward neural network (DNN) based acoustic model. In detail, a state-of-the-art acoustic model is trained as the baseline model. In this step, layer normalization is applied to accelerating the model convergence and improving the generalization performance. Knowledge distillation and pruning are then conducted to compress the model. Our final model can achieve 14.59× parameters reduction, 5× storage size reduction and comparable performance compared with the baseline model. Chenxing Li, Bo Xu 0002 |
ICPR | 1 |
| 2018 | Recurrent Neural Network Based Small-footprint Wake-up-word Speech Recognition System with a Score Calibration MethodabstractIn this paper, we propose a small-footprint wake-up-word speech recognition (WUWSR) system based on long short-term memory (LSTM) recurrent neural network, and we design a novel back-end calibration scoring method named modified zero normalization (MZN). First, LSTM is trained to predict posterior probability of context-dependent state. Next, MZN is adopted to transfer posterior probability to normalized score, which is then converted to confidence score by dynamic programming. Finally, a certain wake-up-word is recognized according to the confidence score. This WUWSR system can recognize multiple wake-up words and change wake-up words flexibly. This system can guarantee low latency by omitting decoding network. Equal error rate (EER) is adopted as the evaluation metric. Experimental results show that the proposed LSTM-based system achieves 33.33% relative improvement compared with a baseline system based on deep feed-forward neural network. Combining the front-end LSTM acoustic model with back-end MZN method, our WUWSR system can achieve 51.92% relative improvement. Chenxing Li, Bo Xu 0002 |
ICPR | 1 |
| 2018 | Single-channel Speech Dereverberation via Generative Adversarial TrainingabstractIn this paper, we propose a single-channel speech dereverberation system (DeReGAT) based on convolutional, bidirectional long short-term memory and deep feed-forward neural network (CBLDNN) with generative adversarial training (GAT).In order to obtain better speech quality instead of only minimizing a mean square error (MSE), GAT is employed to make the dereverberated speech indistinguishable form the clean samples.Besides, our system can deal with wide range reverberation and be well adapted to variant environments.The experimental results show that the proposed model outperforms weighted prediction error (WPE) and deep neural network-based systems.In addition, DeReGAT is extended to an online speech dereverberation scenario, which reports comparable performance with the offline case. Chenxing Li, Tieqiang Wang, Bo Xu 0002 |
INTERSPEECH | 1 |
| 2018 | Performance of auxiliary antenna-based self-interference cancellation in full-duplex radios
Fei Wu 0013, Shihai Shao, Chenxing Li, Youxi Tang |
Sci. China Inf. Sci. | 4 |
| 2017 | Mining Implicit Intention Using Attention-Based RNN Encoder-Decoder Model
Chenxing Li, Yajun Du, Sida Wang 0003 |
ICIC (3) | 1 |
| 2017 | Revealing Encryption for Partial Ordering
Helene Haagh, Yue Ji, Chenxing Li, Claudio Orlandi |
IMACC | 3 |
| 2017 | Ranking webpages using a path trust knowledge graph
Yajun Du, Chenxing Li, Xiaoliang Chen 0003 |
Neurocomputing | 2 |
| 2016 | A Novel Approach of Identifying User Intents in Microblog
Chenxing Li, Yajun Du, Jia Liu 0033, Sida Wang 0003 |
ICIC (3) | 1 |
| 2016 | A Novel Entity Relation Extraction Approach Based on Micro-Blog
Yajun Du, Sida Wang 0003, Chenxing Li, Jianbo Yang |
ICIC (3) | 4 |
| 2016 | BAH: A Bitmap Index Compression Algorithm for Fast Data RetrievalabstractEfficient retrieval of traffic archival data is a must-have technique to detect network attacks, such as APT(advanced persistent threat) attack. In order to take insight from Internet traffic, the bitmap index is increasingly used for efficiently querying over large datasets. However, a raw bitmap index leads to high space consumption and overhead on loading indexes. Various bitmap index compression algorithms are proposed to save storage while improving query efficiency. This paper proposes a new bitmap index compression algorithm called BAH (Byte Aligned Hybrid compression coding). An acceleration algorithm using SIMD is designed to increase the efficiency of AND operation over multiple compressed bitmaps. In all, BAH has a better compression ratio and faster intersection querying speed compared with several previous works such as WAH, PLWAH, COMPAX, Roaring etc. The theoretical analysis shows that the space required by BAH is no larger than 1.6 times the information entropy of the bitmap with density larger than 0.2%. In the experiments, BAH saves about 65% space and 60% space compared with WAH on two datasets. The experiments also demonstrate the query efficiency of BAH with the application in Internet Traffic and Web pages. Chenxing Li, Zhen Chen 0001, Wenxun Zheng, Yinjun Wu |
LCN | 1 |