Chengmin Lin

dblp:75/3381 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0002-2138-3739ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Trustworthy machine learning · 65% Efficient and distributed learning · 28% Reinforcement learning · 7%
Software engineering, system software, and programming languages
1 paper
Services computing and microservices · 50% Operating systems · 50%
Computer networks
1 paper
Edge and fog computing · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.712023
QANS: Toward Quantized Neural Network Adversarial Noise Suppression · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Machine learning › Efficient and distributed learning › model compression › quantization
quantized neural network
0.712023
QANS: Toward Quantized Neural Network Adversarial Noise Suppression · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Machine learning › Trustworthy machine learning › robustness › model robustness evaluation
quantized neural network robustness
0.712023
QANS: Toward Quantized Neural Network Adversarial Noise Suppression · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Edge and fog computing › service provisioning › service deployment
microservice deployment
0.612022
Microservice Deployment in Edge Computing Based on Deep Q Learning · IEEE Trans. Parallel Distributed Syst. 2022
Operating systems › resource management
load balancing
0.612022
Microservice Deployment in Edge Computing Based on Deep Q Learning · IEEE Trans. Parallel Distributed Syst. 2022
Services computing and microservices › microservice architecture
microservice deployment
0.612022
Microservice Deployment in Edge Computing Based on Deep Q Learning · IEEE Trans. Parallel Distributed Syst. 2022
Machine learning › Trustworthy machine learning › robustness
adversarial attack
0.212023
QANS: Toward Quantized Neural Network Adversarial Noise Suppression · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Machine learning › Reinforcement learning › deep reinforcement learning
deep q-learning
0.212022
Microservice Deployment in Edge Computing Based on Deep Q Learning · IEEE Trans. Parallel Distributed Syst. 2022

Methods — techniques the papers use, named apart from their topics

heuristic elastic scaling · 1.7deep q-learning · 1.7gaussian kernel regularization · 0.7activation quantization · 0.7
YearPublicationVenuePosition
2025 Multihardware Adaptive Latency Prediction for Neural Architecture Search
abstract
In hardware-aware neural architecture search (NAS), accurately assessing a model’s inference efficiency is crucial for search optimization. Traditional approaches, which measure numerous samples to train proxy models, are impractical across varied platforms due to the extensive resources needed to remeasure and rebuild models for each platform. To address this challenge, we propose a multihardware-aware NAS method that enhances the generalizability of proxy models across different platforms while reducing the required sample size. Our method introduces a multihardware adaptive latency prediction (MHLP) model that leverages one-hot encoding for hardware parameters and multihead attention mechanisms to effectively capture the intricate interplay between hardware attributes and network architecture features. Additionally, we implement a two-stage sampling mechanism based on probability density weighting to ensure the representativeness and diversity of the sample set. By adopting a dynamic sample allocation mechanism, our method can adjust the adaptive sample size according to the initial model state, providing stronger data support for devices with significant deviations. Evaluations on NAS benchmarks demonstrate the MHLP predictor’s excellent generalization accuracy using only 10 samples, guiding the NAS search process to identify optimal network architectures.
Chengmin Lin, Pengfei Yang 0001, Quan Wang 0006
IEEE Internet Things J.1
2025 PipeMCTS: Pipeline Inference Optimization for Edge Computing via Surrogate Model-Guided MCTS
abstract
For deep neural network (DNN) deployment on Internet of Things (IoT) devices, pipeline-based inference coordinates diverse computing resources in heterogeneous multiprocessor system-on-chips (HMPSoCs) to achieve efficient execution. However, determining the optimal pipeline configuration is challenging due to the exponentially expanding search space and intricate layer-wise dependencies. Existing methods formulate this as a single-shot optimization problem, which struggles to efficiently explore the search space and incurs substantial resource overhead from repeated evaluations of suboptimal configurations. This article proposes PipeMCTS, which reformulates pipeline deployment as a sequential optimization problem solved via Monte Carlo tree search (MCTS). By incrementally constructing the search tree in a layer-wise manner, PipeMCTS effectively prunes unpromising branches and accelerates the search process. PipeMCTS incorporates two key components to enhance search efficiency: 1) a temperature-controlled selection strategy that balances exploration and exploitation in the complex search space and 2) an uncertainty-aware simulation strategy that accelerates search convergence through Gaussian process (GP)-guided evaluation. Experimental results demonstrate that PipeMCTS achieves a 66.42% improvement in throughput compared to arm compute library (ARM-CL).
Jianjun Ding, Dan Xian, Yusheng Jin, Dejun Hua, Wenbin Hua, Wenkai Lv, Chengmin Lin
IEEE Internet Things J.8
2024 Graph-Reinforcement-Learning-Based Dependency-Aware Microservice Deployment in Edge Computing
abstract
Microservice architecture is a design philosophy that achieves decoupling by decomposing a monolithic application into multiple lightweight microservices. Meanwhile, edge computing can significantly reduce service latency and network congestion by extending computation and storage resources to the network edge. Therefore, in the microservice-oriented edge computing platform, a fundamental problem is how to efficiently deploy microservices with complex dependencies on the resource-constrained edge servers to satisfy the Quality of Service (QoS) constraints of users. Most of the existing studies ignore multiple call graphs with differentiated dependencies for an application, which often result in the violation of QoS. To address this issue, in this article, we first model the request response time of multiple instances and multiple call graphs scenario with service conflicts. Then, different from the existing heuristic or approximation algorithms which rely heavily on expert knowledge, we propose a graph-reinforcement-learning-based deployment (GRLD) framework. GRLD uses a graph convolutional network (GCN) to extract the graph data required for multiple call graphs with messages passing and aggregation, and the generated feature is fed into the underlying network of deep-reinforcement-learning (DRL). Experimental results show that GRLD outperforms counterparts in reducing service deployment overhead while satisfying QoS constraints of multiple call graphs.
Wenkai Lv, Pengfei Yang 0001, Tianyang Zheng, Chengmin Lin, Minwen Deng, Quan Wang 0006
IEEE Internet Things J.4
2024 Performance Prediction for Deep Learning Models With Pipeline Inference Strategy
abstract
For Heterogeneous Multi-Processor System-on-Chips (HMPSoCs), a reasonable pipeline design can significantly improve the inference performance of Deep Learning (DL) models. The pipeline design optimization can be modeled as a search problem where an accurate prediction model can efficiently speed up the search process. However, the performance prediction of DL models for the pipeline inference strategy is challenging because of the inter-layer effect, inference details, and variety of model structures. In this paper, we propose TPPNet, a transformer-based model for predicting the inference performance of various DL models with the pipeline inference strategy. TPPNet represents the DL model as an execution sequence with operators and hardware details to extract the hidden factors between layers. Moreover, we apply the Multi-task Learning (MTL) method to accurately predict throughput and latency metrics by constructing a predictive model. To the best of our knowledge, this is the first study dedicated to pipeline inference performance prediction for the DL model on HMPSoCs. We evaluate TPPNet on six well-known DL models using RK3399. The experimental outcomes affirm the high accuracy of TPPNet and its capability to significantly reduce the time overhead associated with pipeline exploration.
Pengfei Yang 0001, Linwei Hu, Wenkai Lv, Chengmin Lin, Quan Wang 0006
IEEE Internet Things J.6
2024 Fine-grained complexity-driven latency predictor in hardware-aware neural architecture search using composite loss
Chengmin Lin, Pengfei Yang 0001, Wenkai Lv, Quan Wang 0006
Inf. Sci.1
2024 Flexi-BOPI: Flexible granularity pipeline inference with Bayesian optimization for deep learning models on HMPSoC
Pengfei Yang 0001, Linwei Hu, Wenkai Lv, Chengmin Lin, Quan Wang 0006
Inf. Sci.6
2024 SLAPP: Subgraph-level attention-based performance prediction for deep learning models
Pengfei Yang 0001, Linwei Hu, Chengmin Lin, Wenkai Lv, Quan Wang 0006
Neural Networks5
2023 Efficient and accurate compound scaling for convolutional neural networks
Chengmin Lin, Pengfei Yang 0001, Quan Wang 0006, Zeyu Qiu, Wenkai Lv
Neural Networks1
2023 QANS: Toward Quantized Neural Network Adversarial Noise Suppression
abstract
Neural network quantization techniques play an important role in efficiently deploying deep learning models on the hardware with limited computing and storage resources. Numerous applications of this technology, such as autopilot, necessitate not just efficiency but also robustness. Research on the robustness of quantized networks against adversarial attacks is becoming one of the major points of interest. In this work, we rethink the impact of quantization on adversarial attacks and explore the boundary of the robustness of quantized neural networks. This study reveals that activation quantization can be used as a defense to weaken adversarial noise, but the robustness of quantized models is still limited by the amplification effect of network errors, including quantization errors and adversarial noise. To address this problem, we propose the quantization adversarial noise suppression (QANS) method that employs a Gaussian kernel regularization constraint to stabilize the model by restricting the perturbation error within two levels of tolerance. Extensive experiments are conducted with Wide ResNet and VGG-16 models on CIFAR-10 and street view house number datasets under different attack methods, including several white-box and black-box attacks. Experimental results show the proposed method achieves superior robustness to prior works.
Chengmin Lin, Pengfei Yang 0001, Tianbing He, Jinpeng Liang, Quan Wang 0006
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2022 Microservice Deployment in Edge Computing Based on Deep Q Learning
abstract
The microservice deployment strategy is promising in reducing the overall service response time in the microservice-oriented edge computing platform. However, existing works ignore the effect of different interaction frequencies among microservices and the decrease in service execution performance caused by the increased node loads. In this article, we first model the invocation relationships among microservices as an undirected and weighted interaction graph to characterize the communication overhead. Then, we propose a multi-objective microservice deployment problem (MMDP) in edge computing. MMDP aims to minimize the communication overhead while achieving load balance between edge nodes. Without the requirement for domain experts, we propose Reward Sharing Deep Q Learning (RSDQL), a learning-based algorithm, to solve MMDP and obtain the optimal deployment strategy. In addition, to improve the scalability of the services, we propose an Elastic Scaling algorithm (ES) based on heuristics to deal with the dynamic pressure of requests. Finally, we conduct a series of experiments in Kubernetes to evaluate the performance of our approach. Experimental results indicate that, compared with interaction-aware strategy and Kubernetes default strategy, RSDQL has shorter response times, more balanced resource loads, and makes services scale elastically according to the request pressure.
Wenkai Lv, Quan Wang 0006, Pengfei Yang 0001, Yunqing Ding, Bijie Yi, Chengmin Lin
IEEE Trans. Parallel Distributed Syst.7