EDBT 2026 Demo / reviewers in the wild / expert
Yukun Qian
dblp:321/6561
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0001-9095-4024ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | D²PPO: Diffusion Policy Policy Optimization with Dispersive LossabstractDiffusion policies excel at robotic manipulation by naturally modeling multimodal action distributions in high-dimensional spaces. Nevertheless, diffusion policies suffer from diffusion representation collapse: semantically similar observations are mapped to indistinguishable features, ultimately impairing their ability to handle subtle but critical variations required for complex robotic manipulation. To address this problem, we propose D²PPO (Diffusion Policy Policy Optimization with Dispersive Loss). D²PPO introduces dispersive loss regularization that combats representation collapse by treating all hidden representations within each batch as negative pairs. D²PPO compels the network to learn discriminative representations of similar observations, thereby enabling the policy to identify subtle yet crucial differences necessary for precise manipulation. In evaluation, we find that early-layer regularization benefits simple tasks, while late-layer regularization sharply enhances performance on complex manipulation tasks. On RoboMimic benchmarks, D²PPO achieves an average improvement of 22.7% in pre-training and 26.1% after fine-tuning, setting new SOTA results. In comparison with SOTA, the results of real-world experiments on a Franka Emika Panda robot show the excitingly high success rate of our method. The superiority of our method is especially evident in complex tasks. Guowei Zou, Weibing Li, Hejun Wu, Yukun Qian |
AAAI | 4 |
| 2026 | LCKPose: Laplacian Candidate Keypoints Modeling for 6D Object Pose Estimation
Zhiyang Mai, Yukun Qian, Hejun Wu, Liangliang Zhou |
MMM (2) | 2 |
| 2026 | OACodec: Audio attribute disentanglement via orthogonal disentanglement and mutual information minimization
Yukun Qian, Lianyu Zhou, Xuyi Zhuang, Mingjiang Wang |
Neural Networks | 1 |
| 2026 | MS-VBRVQ: Multi-scale variable bitrate speech residual vector quantization
Yukun Qian, Shiyun Xu, Xuyi Zhuang, Mingjiang Wang |
Speech Commun. | 1 |
| 2026 | CARVE: Content-Adaptive Rate-Variable Encoding for Neural Speech CodecsabstractNeural speech codecs that convert continuous wave forms into discrete tokens are an essential component for building compact speech representations in downstream applications. However, most existing codecs require both fixed and high frame rates to maintain reconstruction quality, which reduces the efficiency of downstream applications. To address this issue, we propose CARVE, a Content-Adaptive Rate-Variable Encoding strategy that performs compression at an externally specified frame rate within the latent space of a neural speech codec. CARVE comprises three modules: (i) Unsupervised Density Aware Scoring Module that combines inter-frame similarity with redundancy-aware cues to quantify the relative importance of each frame; (ii) Sparse Frame Segmentation Module that allocates finer segmentation to high-density information regions and coarser segmentation to low-density information regions based on these scores; and (iii) Context-Aware Frame Reconstruction Module that maps segment-level context back to frame-level latent representations. Integrating CARVE into the neural speech codec, extensive experiments show that it can perform compression at an externally specified frame rate while maintaining high reconstruction quality, thereby validating its effectiveness and practicality in low-frame-rate neural speech coding scenarios. Audio samples are available at: https://ethuil.github.io/CARVEdemo/ Yukun Qian, Yinghan Cao, Changjun He, Shiyun Xu, Mingjiang Wang |
IEEE Signal Process. Lett. | 2 |
| 2025 | CSMT: Combining Snoring and Metadata-based Text for Sleep Apnea Severity ClassificationabstractSleep apnea is a common sleep disorder that, if untreated, can lead to serious health issues. Snoring is a typical symptom of sleep apnea and can be utilized to develop a noncontact automatic detection method for sleep apnea severity classification (SASC). However, due to patient heterogeneity, the acoustic characteristics of snoring vary significantly among individuals. To address this issue, we introduced a text-audio multimodal model that leverages patient’s metadata to provide valuable supplementary information for SASC task. Specifically, we utilized text descriptions derived from metadata and snoring sounds to fine-tune a pretrained text-audio multimodal model. The metadata includes patient’s physical indicators such as gender, age, BMI, neck circumference, and blood pressure. We constructed a snoring dataset that included four sleep apnea severity levels. On this dataset, our method achieved a classification F-score of 74.34%. We conducted a series of ablation experiments to validate the effectiveness of improving SASC performance by leveraging both metadata-based text and snoring sounds. Additionally, we discussed the model’s performance in scenarios where parts of the metadata are unavailable, a situation that may occur in real-world applications. Heng Li 0013, Yukun Qian, Yun Lu 0004, Mingjiang Wang |
ICASSP | 2 |
| 2025 | A Novel Compressive Compound Word Encoding and Independent Word Attention for Symbolic Music GenerationabstractSymbolic music generation involves using symbolic encoding to represent music pieces as token sequences and using neural sequence models to create music by generating sequences of tokens. Symbolic encodings are primarily categorized into two types: independent word encoding and compound word encoding. Independent word encoding treats each token equally, where the model predicts one token at each timestep. Compound word encoding groups associated tokens and combines them into a super token, and places them in one timestep, where the model predicts multiple different tokens at each timestep. Previous works related to compound word encoding only use compressed representation constructed from super tokens for model learning without considering the independent tokens that consist of super tokens. To enable models to capture the dependencies between independent tokens in super tokens when using compound word encoding, we propose compressive compound word (CCP) encoding and independent word attention (IWA). The CCP reduces the number of ignore tokens in super tokens and timesteps required to encode one music piece. The IWA learns how the independent tokens in super tokens are organized. Experimental results demonstrate that IWA can improve the quality of generated music while CCP is more efficient when inference. Lianyu Zhou, Yukun Qian |
ICASSP | 3 |
| 2025 | Joint Training Framework for Accent and Speech Recognition Based on Conformer Low-Rank AdaptationabstractIn real-world scenarios, accent variations often reduce Automatic Speech Recognition (ASR) accuracy. Addressing this typically involves a multi-task ASR and Accent Recognition (ASR-AR) framework, but there is limited research on optimizing task-specific feature extraction and enhancing ASR with AR information. This study introduces the Conformer Low-rank Adaptation for Joint Accent and Speech Recognition (CLAnSR), employing LoRA to augment both ASR and AR capabilities using a shared pre-trained base encoder. This approach significantly reduces the model’s parameter and training resource demands while facilitating the extraction of task-specific features. Additionally, we have incorporated accent-aware multi-channel embedding layers, which through spatially independent embeddings, enhance the model’s capacity to accurately represent tokens across diverse dialectical contexts. Tested on the KeSpeech dataset, CLAnSR reaches state-of-the-art AR accuracy 80.41% and competitive ASR CER 8.39%, outperforming non-LLM systems and matching those with LLMs. It reduces parameters by 34.02% and enhances both ASR and AR performance, effectively handling speech dialect variations and advancing the field. Xuyi Zhuang, Yukun Qian, Shiyun Xu, Mingjiang Wang |
ICASSP | 2 |
| 2025 | SF-AN: A lightweight shuffle Fourier attention network for multi-channel speech enhancement
Shiyun Xu, Yinghan Cao, Yukun Qian, Changjun He, Mingjiang Wang |
Speech Commun. | 4 |
| 2025 | Hypformer: A Fast Hypothesis-Driven Rescoring Speech Recognition FrameworkabstractRecently, the performance of non-autoregressive ASR models has made significant progress but still lags behind hybrid CTC/attention systems. This paper introduces Hypformer, a fast hypothesis-driven rescoring speech recognition framework. Multiple hypothetical prefixes are realized by fast prefix generation algorithm. With two different rescoring methods, nar-ar rescoring and nar$^{2}$rescoring, Hypformer can flexibly switch between autoregressive and non-autoregressive decoding modes to perform rescoring of hypothesis prefixes. Experiments on the standard Mandarin datasets AISHELL-1 and AISHELL-2 demonstrate that Hypformer outperforms the state-of-the-art Hybrid CTC/Attention systems in ASR performance while achieving a speedup of over six times. Experiments on the Mandarin sub-dialect dataset KeSpeech indicate that Hypformer achieves more accurate recognition by leveraging richer contextual information. Xuyi Zhuang, Yukun Qian, Mingjiang Wang |
IEEE Signal Process. Lett. | 2 |
| 2024 | Lightweight Dynamic Sparse Transformer for Monaural Speech Enhancement
Xuyi Zhuang, Yukun Qian, Mingjiang Wang |
INTERSPEECH | 3 |
| 2023 | Half-Temporal and Half-Frequency Attention U2Net for Speech Signal ImprovementabstractDuring communication, volume changes, noise, and reverberation can disturb speech signals, significantly affecting the quality and intelligibility of speech. In the context of the ICASSP 2023 Signal Processing Grand Challenge, the first Speech Signal Improvement Grand Challenge (SIG) is organized to improve the quality of speech signals during communication. This paper proposes half-temporal and half-frequency attention U2Net for improving full-band speech signal. Channel-spectrum attention is proposed for the skip connection between the encoder and decoder. The proposed model achieves 0.353, 1.289, 0.604, 0.625, and 0.924 improvements in signal, noise, overall, reverberation, and loudness, respectively, in the SIG subjective test. The proposed model achieved fourth place in the SIG real-time track, showing excellent denoising and de-reverberation performance. Shiyun Xu, Xuyi Zhuang, Yukun Qian, Lianyu Zhou, Mingjiang Wang |
ICASSP | 4 |
| 2023 | Automatic Speech Recognition Transformer with Global Contextual Information Decoder
Yukun Qian, Xuyi Zhuang, Mingjiang Wang |
INTERSPEECH | 1 |
| 2022 | FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional NetworkabstractIn recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of full-band (48 kHz) speech signals is still very challenging. Because of the low energy of spectral information in the high-frequency part, it is more difficult to directly model and enhance the full-band spectrum using neural networks. To solve this problem, this paper proposes a two-stage real-time speech enhancement model with extraction-interpolation mechanism for a full-band signal. The 48 kHz full-band time-domain signal is divided into three sub-channels by extracting, and a two-stage processing scheme of ‘masking + compensation’ is proposed to enhance the signal in the complex domain. After the two-stage enhancement, the enhanced full-band speech signal is restored by interval interpolation. In the subjective listening and word accuracy test, our proposed model achieves superior performance and outperforms the baseline model overall by 0.59 MOS and 4.0% WAcc for the non-personalized speech denoising task. Lu Zhang 0055, Xuyi Zhuang, Yukun Qian, Heng Li 0013, Mingjiang Wang |
ICASSP | 4 |
| 2022 | Coarse-Grained Attention Fusion With Joint Training Framework for Complex Speech Enhancement and End-to-End Speech Recognition
Xuyi Zhuang, Yukun Qian, Mingjiang Wang |
INTERSPEECH | 4 |