EDBT 2026 Demo / reviewers in the wild / expert
Shun Lu 0001
dblp:273/1222-1
· DBLP profile ↗
18ranked-venue papers
7as first author
18since 2021 · last 2026
0000-0003-0865-4896ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 7 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SEA: Hierarchically searching efficient adapters for pre-trained models
Shun Lu 0001, Fangyuan Mao, Junkun Chen, Jilin Mei, Yu Hu 0001 |
Neural Networks | 1 |
| 2026 | PID: Physics-Informed Diffusion Model for Infrared Image Generation
Fangyuan Mao, Jilin Mei, Shun Lu 0001, Fuyang Liu, Fangzhou Zhao, Yu Hu 0001 |
Pattern Recognit. | 3 |
| 2024 | PHD-NAS: Preserving helpful data to promote Neural Architecture Search
Shun Lu 0001, Yu Hu 0001, Longxing Yang, Jilin Mei, Jianchao Tan, Chengru Song |
Neurocomputing | 1 |
| 2023 | PINAT: A Permutation INvariance Augmented Transformer for NAS PredictorabstractTime-consuming performance evaluation is the bottleneck of traditional Neural Architecture Search (NAS) methods. Predictor-based NAS can speed up performance evaluation by directly predicting performance, rather than training a large number of sub-models and then validating their performance. Most predictor-based NAS approaches use a proxy dataset to train model-based predictors efficiently but suffer from performance degradation and generalization problems. We attribute these problems to the poor abilities of existing predictors to character the sub-models' structure, specifically the topology information extraction and the node feature representation of the input graph data. To address these problems, we propose a Transformer-like NAS predictor PINAT, consisting of a Permutation INvariance Augmentation module serving as both token embedding layer and self-attention head, as well as a Laplacian matrix to be the positional encoding. Our design produces more representative features of the encoded architecture and outperforms state-of-the-art NAS predictors on six search spaces: NAS-Bench-101, NAS-Bench-201, DARTS, ProxylessNAS, PPI, and ModelNet. The code is available at https://github.com/ShunLu91/PINAT. Shun Lu 0001, Yu Hu 0001, Peihao Wang, Yan Han 0001, Jianchao Tan, Sen Yang 0004, Ji Liu 0002 |
AAAI | 1 |
| 2023 | PA&DA: Jointly Sampling PAth and DAta for Consistent NASabstractBased on the weight-sharing mechanism, one-shot NAS methods train a supernet and then inherit the pre-trained weights to evaluate sub-models, largely reducing the search cost. However, several works have pointed out that the shared weights suffer from different gradient descent directions during training. And we further find that large gradient variance occurs during supernet training, which degrades the supernet ranking consistency. To mitigate this issue, we propose to explicitly minimize the gradient variance of the supernet training by jointly optimizing the sampling distributions of PAth and DAta (PA&DA). We theoretically derive the relationship between the gradient variance and the sampling distributions, and reveal that the optimal sampling probability is proportional to the normalized gradient norm of path and training data. Hence, we use the normalized gradient norm as the importance indicator for path and training data, and adopt an importance sampling strategy for the supernet training. Our method only requires negligible computation cost for optimizing the sampling distributions of path and data, but achieves lower gradient variance during supernet training and better generalization performance for the supernet, resulting in a more consistent NAS. We conduct comprehensive comparisons with other improved approaches in various search spaces. Results show that our method surpasses others with more reliable ranking performance and higher accuracy of searched architectures, showing the effectiveness of our method. Code is available at https://github.com/ShunLu91/PA-DA. Shun Lu 0001, Yu Hu 0001, Longxing Yang, Jilin Mei, Jianchao Tan, Chengru Song |
CVPR | 1 |
| 2023 | MixPath: A Unified Approach for One-shot Neural Architecture SearchabstractBlending multiple convolutional kernels is proved advantageous in neural architecture design. However, current two-stage neural architecture search methods are mainly limited to single-path search spaces. How to efficiently search models of multi-path structures remains a difficult problem. In this paper, we are motivated to train a one-shot multi-path supernet to accurately evaluate the candidate architectures. Specifically, we discover that in the studied search spaces, feature vectors summed from multiple paths are nearly multiples of those from a single path. Such disparity perturbs the supernet training and its ranking ability. Therefore, we propose a novel mechanism called Shadow Batch Normalization (SBN) to regularize the disparate feature statistics. Extensive experiments prove that SBNs are capable of stabilizing the optimization and improving ranking performance. We call our unified multi-path one-shot approach as MixPath, which generates a series of models that achieve state-of-the-art results on ImageNet. Xiangxiang Chu, Shun Lu 0001, Xudong Li 0003, Bo Zhang 0046 |
ICCV | 2 |
| 2023 | Unleashing the Power of Gradient Signal-to-Noise Ratio for Zero-Shot NASabstractNeural Architecture Search (NAS) aims to automatically find optimal neural network architectures in an efficient way. Zero-Shot NAS is a promising technique that leverages proxies to predict the accuracy of candidate architectures without any training. However, we have observed that most existing proxies do not consistently perform well across different search spaces, and are less concerned with generalization. Recently, the gradient signal-to-noise ratio (GSNR) was shown to be correlated with neural network generalization performance. In this paper, we not only explicitly give the probability that larger GSNR at network initialization can ensure better generalization, but also theoretically prove that GSNR can ensure better convergence. Then we design the ξ-based gradient signal-to-noise ratio (ξ-GSNR) as a Zero-Shot NAS proxy to predict the network accuracy at initialization. Extensive experiments in different search spaces demonstrate that ξ-GSNR provides superior ranking consistency compared to previous proxies. Moreover, ξ-GSNR-based Zero-Shot NAS also achieves outstanding performance when directly searching for the optimal architecture in various search spaces and datasets. The source code is available at https://github.com/Sunzh1996/Xi-GSNR. Longxing Yang, Shun Lu 0001, Jilin Mei, Wen-Xiao Zhao, Yu Hu 0001 |
ICCV | 4 |
| 2023 | Sweet Gradient matters: Designing consistent and efficient estimator for Zero-shot Architecture Search
Longxing Yang, Yanxin Fu, Shun Lu 0001, Jilin Mei, Wen-Xiao Zhao, Yu Hu 0001 |
Neural Networks | 3 |
| 2022 | AGNAS: Attention-Guided Micro and Macro-Architecture SearchabstractMicro- and macro-architecture search have emerged as two popular NAS paradigms recently. Existing methods leverage different search strategies for searching micro- and macro- architectures. When using architecture parameters to search for micro-structure such as normal cell and reduction cell, the architecture parameters can not fully reflect the corresponding operation importance. When searching for the macro-structure chained by pre-defined blocks, many sub-networks need to be sampled for evaluation, which is very time-consuming. To address the two issues, we propose a new search paradigm, that is, leverage the attention mechanism to guide the micro- and macro-architecture search, namely AGNAS. Specifically, we introduce an attention module and plug it behind each candidate operation or each candidate block. We utilize the attention weights to represent the importance of the relevant operations for the micro search or the importance of the relevant blocks for the macro search. Experimental results show that AGNAS can achieve 2.46% test error on CIFAR-10 in the DARTS search space, and 23.4% test error when directly searching on ImageNet in the ProxylessNAS search space. AGNAS also achieves optimal performance on NAS-Bench-201, outperforming state-of-the-art approaches. The source code can be available at https://github.com/Sunzh1996/AGNAS. Yu Hu 0001, Shun Lu 0001, Longxing Yang, Jilin Mei, Yinhe Han 0001, Xiaowei Li 0001 |
ICML | 3 |
| 2022 | Searching for BurgerFormer with Micro-Meso-Macro Space DesignabstractWith the success of Transformers in the computer vision field, the automated design of vision Transformers has attracted significant attention. Recently, MetaFormer found that simple average pooling can achieve impressive performance, which naturally raises the question of how to design a search space to search diverse and high-performance Transformer-like architectures. By revisiting typical search spaces, we design micro-meso-macro space to search for Transformer-like architectures, namely BurgerFormer. Micro, meso, and macro correspond to the granularity levels of operation, block and stage, respectively. At the microscopic level, we enrich the atomic operations to include various normalizations, activation functions, and basic operations (e.g., multi-head self attention, average pooling). At the mesoscopic level, a hamburger structure is searched out as the basic BurgerFormer block. At the macroscopic level, we search for the depth, width, and expansion ratio of the network based on the multi-stage architecture. Meanwhile, we propose a hybrid sampling method for effectively training the supernet. Experimental results demonstrate that the searched BurgerFormer architectures achieve comparable even superior performance compared with current state-of-the-art Transformers on the ImageNet and COCO datasets. The codes can be available at https://github.com/xingxing-123/BurgerFormer. Longxing Yang, Yu Hu 0001, Shun Lu 0001, Jilin Mei, Yinhe Han 0001, Xiaowei Li 0001 |
ICML | 3 |
| 2022 | Conformer Space Neural Architecture Search for Multi-Task Audio Separation
Shun Lu 0001, Chenxing Li, Jianchao Tan, Feng Deng, Chengru Song |
INTERSPEECH | 1 |
| 2022 | WA-Transformer: Window Attention-based Transformer with Two-stage Strategy for Multi-task Audio Source Separation
Chenxing Li, Feng Deng, Shun Lu 0001, Jianchao Tan, Chengru Song |
INTERSPEECH | 4 |
| 2022 | STC-NAS: Fast neural architecture search with source-target consistency
Yu Hu 0001, Longxing Yang, Shun Lu 0001, Jilin Mei, Yinhe Han 0001, Xiaowei Li 0001 |
Neurocomputing | 4 |
| 2021 | DDSAS: Dynamic and Differentiable Space-Architecture SearchabstractNeural Architecture Search (NAS) has made remarkable progress in automatically designing neural networks. However, existing differentiable NAS and stochastic NAS methods are either biased towards exploitation and thus may converge to a local minimum, or biased towards exploration and thus converge slowly. In this work, we propose a Dynamic and Differentiable Space-Architecture Search (DDSAS) method to address the exploration-exploitation dilemma. DDSAS dynamically samples space, searches architectures in the sampled subspace with gradient descent, and leverages the Upper Confidence Bound (UCB) to balance exploitation and exploration. The whole search space is elastic, offering flexibility to evolve and to consider resource constraints. Experiments on image classification datasets demonstrate that with only 4GB memory and 3 hours for searching, DDSAS achieves 2.39% test error on CIFAR10, 16.26% test error on CIFAR100, and 23.9% test error when transferring to ImageNet. When directly searching on ImageNet, DDSAS achieves comparable accuracy with more than 6.5 times speedup over state-of-the-art methods. The source codes are available at https://github.com/xingxing-123/DDSAS. Longxing Yang, Yu Hu 0001, Shun Lu 0001, Jilin Mei, Yiming Zeng 0003, Zhi-Ping Shi 0002, Yinhe Han 0001, Xiaowei Li 0001 |
ACML | 3 |
| 2021 | SpeechNAS: Towards Better Trade-Off Between Latency and Accuracy for Large-Scale Speaker VerificationabstractRecently, x-vector [1] has been a successful and popular approach for speaker verification, which employs a time delay neural network (TDNN) and statistics pooling to extract speaker characterizing embedding from variable-length utterances. Improvement upon the x-vector has been an active research area, and enormous neural networks have been elaborately designed based on the x-vector, e.g., extended TDNN (E-TDNN) [2], factorized TDNN (F-TDNN) [3], and densely connected TDNN (D-TDNN) [4]. In this work, we try to identify the optimal architectures from a TDNN based search space employing neural architecture search (NAS), named SpeechNAS. Leveraging the recent advances in the speaker recognition, such as high-order statistics pooling, multi-branch mechanism, D-TDNN and angular additive margin softmax (AAM) loss with a minimum hyper-spherical energy (MHE), SpeechNAS automatically discovers five network architectures, from SpeechNAS-1 to SpeechNAS-5, of various numbers of parameters and GFLOPs on the large-scale text-independent speaker recognition dataset VoxCelebl. Our derived best neural network achieves an equal error rate (EER) of 1.02% on the standard test set of VoxCelebl, which surpasses previous TDNN based state-of-the-art approaches by a large margin. Wentao Zhu 0001, Tianlong Kong, Shun Lu 0001, Feng Deng, Sen Yang 0004, Ji Liu 0002 |
ASRU | 3 |
| 2021 | DU-DARTS: Decreasing the Uncertainty of Differentiable Architecture Search
Shun Lu 0001, Yu Hu 0001, Longxing Yang, Jilin Mei, Yiming Zeng 0003, Xiaowei Li 0001 |
BMVC | 1 |
| 2021 | DARTS-: Robustly Stepping out of Performance Collapse Without Indicators
Xiangxiang Chu, Xiaoxing Wang, Bo Zhang 0046, Shun Lu 0001, Xiaolin Wei, Junchi Yan |
ICLR | 4 |
| 2021 | TNASP: A Transformer-based NAS Predictor with a Self-evolution FrameworkabstractPredictor-based Neural Architecture Search (NAS) continues to be an important topic because it aims to mitigate the time-consuming search procedure of traditional NAS methods. A promising performance predictor determines the quality of final searched models in predictor-based NAS methods. Most existing predictor-based methodologies train model-based predictors under a proxy dataset setting, which may suffer from the accuracy decline and the generalization problem, mainly due to their poor abilities to represent spatial topology information of the graph structure data. Besides the poor encoding for spatial topology information, these works did not take advantage of the temporal information such as historical evaluations during training. Thus, we propose a Transformer-based NAS performance predictor, associated with a Laplacian matrix based positional encoding strategy, which better represents topology information and achieves better performance than previous state-of-the-art methods on NAS-Bench-101, NAS-Bench-201, and DARTS search space. Furthermore, we also propose a self-evolution framework that can fully utilize temporal information as guidance. This framework iteratively involves the evaluations of previously predicted results as constraints into current optimization iteration, thus further improving the performance of our predictor. Such framework is model-agnostic, thus can enhance performance on various backbone structures for the prediction task. Our proposed method helped us rank 2nd among all teams in CVPR 2021 NAS Competition Track 2: Performance Prediction Track. Shun Lu 0001, Jianchao Tan, Sen Yang 0004, Ji Liu 0002 |
NeurIPS | 1 |