Cheng Lu 0007

dblp:91/1482-7 · DBLP profile ↗
← Back
16ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-1877-7531ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 since 2021Theory of computation · 8 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 2
YearPublicationVenuePosition
2025 Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
abstract
The goal of the acoustic scene classification (ASC) task is to classify recordings into one of the predefined acoustic scene classes. However, in real-world scenarios, ASC systems often encounter challenges such as recording device mismatch, low-complexity constraints, and the limited availability of labeled data. To alleviate these issues, in this paper, a data-efficient and low-complexity ASC system is built with a new model architecture and better training strategies. Specifically, we firstly design a new low-complexity architecture named Rep-Mobile by integrating multi-convolution branches which can be reparameterized at inference. Compared to other models, it achieves better performance and less computational complexity. Then we apply the knowledge distillation strategy and provide a comparison of the data efficiency of the teacher model with different architectures. Finally, we propose a progressive pruning strategy, which involves pruning the model multiple times in small amounts, resulting in better performance compared to a single step pruning. Experiments are conducted on the TAU dataset. With Rep-Mobile and these training strategies, our proposed ASC system achieves the state-of-the-art (SOTA) results so far, while also winning the first place with a significant advantage over others in the DCASE2024 Challenge.
Bing Han 0008, Wen Huang 0004, Zhengyang Chen, Anbai Jiang, Pingyi Fan, Cheng Lu 0007, Zhiqiang Lv, Jia Liu 0001, Weiqiang Zhang 0001, Yanmin Qian
ICASSP6
2025 Adaptive Prototype Learning for Anomalous Sound Detection with Partially Known Attributes
abstract
Adapting pre-trained models has become the dominant approach for anomalous sound detection (ASD), where classifying the attributes of machine working status is commonly chosen as the deputy task for fine-tuning. However, attributes might be intractable to collect for some machines, causing the label to bear mixed granularity and thus deprecating the ASD performance. Therefore, we propose an adaptive proto-type learning scheme for fine-tuning pre-trained models, which adaptively scales coarse-grained labels to sub-centers so as to keep consistency with fine-grained labels. To deal with domain shift, we employ SMOTE to over-sample the prototypes of the target domain. The experiment on the DCASE 2024 ASD dataset demonstrates the efficacy of the proposed scheme, setting up a new milestone of 65.01% on both sets and outperforming the best system of the challenge. A detailed ablation study is also conducted to validate the effectiveness.
Anbai Jiang, Xinhu Zheng, Yihong Qiu, Pingyi Fan, Cheng Lu 0007, Jia Liu 0001
ICASSP7
2024 Exploring Large Scale Pre-Trained Models for Robust Machine Anomalous Sound Detection
abstract
Machine anomalous sound detection is a useful technique for various applications, but it often suffers from poor generalization due to the challenges of data collection and complex acoustic environment. To address this issue, we propose a robust machine anomalous sound detection model that leverages self-supervised pre-trained models on large-scale speech data. Specifically, we assign different weights to the features from different layers of the pre-trained model and then use the working condition as the label for self-supervised classification fine-tuning. Moreover, we introduce a data augmentation method that simulates different operating states of the machine to enrich the dataset. Furthermore, we devise a transformer pooling method that fuses the features of different segments. Experiments on the DCASE2023 dataset show that our proposed method outperforms the commonly used reconstruction-based autoencoder and classification-based convolutional network by a large margin, demonstrating the effectiveness of large-scale pre-training for enhancing the generalization and robustness of machine anomalous sound detection. In Task2 of DCASE2023, we achieve 2nd place with these methods.
Bing Han 0008, Zhiqiang Lv, Anbai Jiang, Wen Huang 0004, Zhengyang Chen, Yufeng Deng, Cheng Lu 0007, Weiqiang Zhang 0001, Pingyi Fan, Jia Liu 0001, Yanmin Qian
ICASSP8
2024 A graphic structure based branch-and-bound algorithm for complex quadratic optimization and applications to magnitude least-square problem
Cheng Lu 0007, Jitao Ma, Zhibin Deng, Wenxun Xing
J. Glob. Optim.1
2023 New semidefinite relaxations for a class of complex quadratic programming problems
Yingzhe Xu, Cheng Lu 0007, Zhibin Deng, Ya-Feng Liu
J. Glob. Optim.2
2021 A branch-and-bound algorithm for solving max-k-cut problem
Cheng Lu 0007, Zhibin Deng
J. Glob. Optim.1
2020 Set-completely-positive representations and cuts for the max-cut polytope and the unit modulus lifting
Florian Jarre, Felix Lieder, Ya-Feng Liu, Cheng Lu 0007
J. Glob. Optim.4
2019 On the Equivalence of Semidifinite Relaxations for MIMO Detection with General Constellations
abstract
The multiple-input multiple-output (MIMO) detection problem is a fundamental problem in modern digital communications. Semidefinite relaxation (SDR) based algorithms are a popular class of approaches to solving the problem because the algorithms have a polynomial-time worst-case complexity and generally can achieve a good detection error rate performance. In spite of the existence of various different SDRs for the MIMO detection problem in the literature, very little is known about the relationship between these SDRs. This paper aims to fill this theoretical gap. In particular, this paper shows that two existing SDRs for the MIMO detection problem, which take quite different forms and are proposed by using different techniques, are equivalent. As a byproduct of the equivalence result, the tightness of one of the above two SDRs under a sufficient condition can be obtained.
Ya-Feng Liu, Cheng Lu 0007
ICASSP3
2019 A sensitive-eigenvector based global algorithm for quadratically constrained quadratic programming
Cheng Lu 0007, Zhibin Deng, Xiaoling Guo
J. Glob. Optim.1
2018 Argument division based branch-and-bound algorithm for unit-modulus constrained complex quadratic programming
Cheng Lu 0007, Zhibin Deng, Weiqiang Zhang 0001, Shu-Cherng Fang
J. Glob. Optim.1
2017 An eigenvalue decomposition based branch-and-bound algorithm for nonconvex quadratic programming problems with convex quadratic constraints
Cheng Lu 0007, Zhibin Deng, Qingwei Jin
J. Glob. Optim.1
2015 High-performance Swahili keyword search with very limited language pack: The THUEE system for the OpenKWS15 evaluation
abstract
This paper presents the Swahili keyword search system developed by the THUEE team for the OpenKWS15 evaluation, which is conducted by NIST under the IARPA Babel program. There are several highlights in the development of the system, including automatic generation of the pronunciation lexicon, aggressive data augmentation, the multilingual bottleneck feature extractor trained from 6 languages, text selection from web data for language model training, semi-supervised training for acoustic models and language models, out-of-vocabulary keyword detection using morphemes and a rich diversity of the systems for combination. A wide variety of acoustic modeling techniques are explored and compared. Up to 12 different individual systems are used for combination. The system achieves the state-of-the-art performance in the required condition of the evaluation.
Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Jia Liu 0001
ASRU3
2015 Improved system fusion for keyword search
abstract
It has been demonstrated that system fusion can significantly improve the performance of keyword search. In this paper, we compare the performance of several widely-used arithmetic-based fusion methods using different normalization pipeline and try to find the best pipeline. A novel arithmetic-based fusion method is proposed in this work. The method supplies a more effective way to incorporate the number of systems which have non-zero scores for a detection. When tested on the development test dataset of the OpenKWS15 Evaluation, the proposed method achieves the highest maximum term-weighted value (MTWV) and actual term-weighted value (ATWV) among all other arithmetic-based fusion methods. Usually, discriminative fusion methods employing classifiers can outperform arithmetic-based fusion methods. A DNN-based fusion method is explored in this work. After word-burst information is added, the DNN-based fusion method outperforms all other methods. In addition, it is notable that our arithmetic-based method achieves the same MTWV as the DNN-based method.
Zhiqiang Lv, Cheng Lu 0007, Jian Kang 0006, Like Hui, Weiqiang Zhang 0001, Jia Liu 0001
ASRU3
2015 The THUEE system for the openKWS14 keyword search evaluation
abstract
The OpenKWS14 keyword search evaluation is one of the most challenging and influential evaluations in the field of speech recognition. Its goal is to build a high-performance keyword search system for a minority language with limited training data in a short period of time. We present the system of the Department of Electronic Engineering, Tsinghua University (THUEE team) for the OpenKWS14 keyword search evaluation. The highlights of the system include the use of convolutional maxout neural networks for acoustic modeling and the use of neural network language models for one-pass lattice generation. The final system is a fusion of 8 sub-systems. The system has achieved an actual term weighted value (ATWV) of 0.5107 for the full language pack (FullLP) condition in the evaluation, ranking third among the participating teams.
Zhiqiang Lv, Beili Song, Yongzhe Shi, Wei-lan Wu, Cheng Lu 0007, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP6
2015 Neuron sparseness versus connection sparseness in deep neural network for large vocabulary speech recognition
abstract
Exploiting sparseness in deep neural networks is an important method for reducing the computational cost. In this paper, we study neuron sparseness in deep neural networks for acoustic modeling. For the feed-forward stage, we only activate neurons whose input values are larger than a given threshold, and set the outputs of inactive nodes to zero. Thus, only a few nonzero outputs are fed to the next layer. Using this method, the output vector of each hidden layer becomes very sparse, so that the computational cost of the feed-forward algorithm can be reduced by adopting sparse matrix operations. The proposed method is evaluated in both small and large vocabulary speech recognition tasks, and results demonstrate that we can reduce the nonzero outputs to fewer than 20% of the total number of hidden nodes, without sacrificing speech recognition performance.
Jian Kang 0006, Cheng Lu 0007, Weiqiang Zhang 0001, Jia Liu 0001
ICASSP2
2015 Conic approximation to nonconvex quadratic programming with convex quadratic constraints
Zhibin Deng, Shu-Cherng Fang, Qingwei Jin, Cheng Lu 0007
J. Glob. Optim.4