VLDB 2026 Research / reviewers in the wild / expert
Chuxi Li
dblp:194/5966
· DBLP profile ↗
13ranked-venue papers
5as first author
11since 2021 · last 2025
0000-0001-9595-9959ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CSDSE: An efficient design space exploration framework for deep neural network accelerator based on cooperative search
Kaijie Feng, Xiaoya Fan, Jianfeng An, Haoyang Wang 0014, Chuxi Li |
Neurocomputing | 5 |
| 2024 | Resource-Efficient Heterogenous Federated Continual Learning on EdgeabstractFederated learning (FL) has been widely deployed on edge devices. In practical, the data collected by edge devices exhibits temporal variations. This leads to catastrophic forgetting issue. Continual learning methods can be used to address this problem. However, when deploying these methods in FL on edge devices, it is challenging to adapt to the limited resources and heterogeneous data of the deployed devices, which reduces the efficiency and effectiveness of federated continual learning (FCL). Therefore, this article proposes a resource-efficient heterogeneous FCL framework. This framework divides the global model into an adaptation part for new knowledge and a preservation part for old knowledge. The preservation part is used to address the catastrophic forgetting problem. Only the adaptation part is trained when learning new knowledge on a new task, reducing resource consumption. Additionally, the framework mitigates the impact of heterogeneous data through an aggregation method based on feature representation. Experimental results show that our method performs well in mitigating catastrophic forgetting in a resource-efficient manner. Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Haoyang Wang 0014, Meng Zhang 0047 |
DATE | 3 |
| 2024 | Efficient knowledge management for heterogeneous federated continual learning on resource-constrained edge devices
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Haoyang Wang 0014, Meng Zhang 0047 |
Future Gener. Comput. Syst. | 3 |
| 2024 | Equalized Aggregation for Heterogeneous Federated Mobile Edge LearningabstractFederated Learning (FL) is widely used in mobile edge applications. However, the heterogeneity issues of mobile edge devices pose significant challenges to the generalization of the global model in FL. In this paper, we propose LegoFL to simultaneously solve multiple heterogeneity issues in response to mobile edge computing characteristics. LegoFL identifies two types of heterogeneous behaviors in FL, namely heterogeneous parameter training and communication behaviors, to address multiple heterogeneity issues. These two types of heterogeneous behaviors result in feature and feature representation range mismatches between local communication parameters. To reduce these mismatches and improve the generalization of the global model, LegoFL dynamically distinguishes the parameter feature representation of different nodes using the global model's common feature as guidance. Then, under the connection states and system communication constraints, LegoFL dynamically selects contribution parameters on each device that can guarantee the generalization and performance of the global model for communication. Finally, to avoid the overfitting problem of the global model, heterogeneous local models are aggregated at the central server with matched feature representations. Extensive experiments on various datasets show that LegoFL achieves competitive performance. The accuracy and communication efficiency are improved by up to 12.86$\%$and 4.09× compared to state-of-the-art approaches. Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Jiaying Yang, Meng Zhang 0047 |
IEEE Trans. Mob. Comput. | 3 |
| 2023 | CSDSE: Apply Cooperative Search to Solve the Exploration-Exploitation Dilemma of Design Space Exploration
Kaijie Feng, Xiaoya Fan, Jianfeng An, Haoyang Wang 0014, Chuxi Li |
ICA3PP (4) | 5 |
| 2023 | Joint heterogeneity-aware personalized federated search for energy efficient battery-powered edge computing
Zhao Yang 0005, Shengbing Zhang, Chuxi Li, Jiaying Yang, Meng Zhang 0047 |
Future Gener. Comput. Syst. | 3 |
| 2023 | ACDSE: A Design Space Exploration Method for CNN Accelerator based on Adaptive Compression MechanismabstractCustomized accelerators for Convolutional Neural Network (CNN) can achieve better energy efficiency than general computing platforms. However, the design of a high-performance accelerator should take into account a variety of parameters and physical constraints. The increasing parameters and tighter constraints gradually complicate the design space, which poses new challenges to the capacity and efficiency of design space exploration methods. In this paper, we provide a novel design space exploration method named ACDSE for optimizing the design process of CNN accelerators. ACDSE implements the adaptive compression mechanism to dynamically adjust the search range and prune low-value design points according to the exploration states. As a result, it can focus on valuable subspace while also improving exploration capacity and efficiency. Additionally, we implement ACDSE to address the problem of CNN accelerator latency optimization. The experiment indicates that, compared to former DSE methods, ACDSE can reduce latency and increase efficiency by 1.39x-5.07x and 2.07x-43.87x, respectively, under the most stringent constraint conditions, demonstrating its superior adaptability to the complicated design space. Kaijie Feng, Xiaoya Fan, Jianfeng An, Chuxi Li, Kaiyue Di, Jiangfei Li |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2022 | DCNN search and accelerator co-design: Improve the adaptability between NAS frameworks and embedded platforms
Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047 |
Integr. | 1 |
| 2022 | Memory-Computing Decoupling: A DNN Multitasking Accelerator With Adaptive Data ArrangementabstractMultiple deep neural networks (DNNs) are increasingly used in real-world intelligent applications, such as intelligent robotics and autonomous vehicles to collectively complete complicated tasks running on edge devices. Because each layer of the subtasks prefers a distinct dataflow due to the heterogeneity in shape and scale of the network layers, a variable dataflow approach on the DNN accelerators is urgently required. On DNN accelerators that enable multiple dataflows, however, we detect a dimension mismatch between parallel processing under the dataflow approach and linear data memory arrangement. When multiple DNN tasks share partial features or weights, the issue is further exacerbated. During processing, this mismatch causes a sluggish data supply from both off-chip and on-chip memory. Consequently, the overall throughput, performance, and energy efficiency suffer since DNN models are sensitive to data density. In this work, we reveal the mechanism behind this data dimension mismatch and present a series of metrics that quantify the influence on system performance. On this foundation, we offer a framework that tracks the data tensor dimension conversion and employs a flexible data arrangement over multi-DNN computation to adapt to dataflow variability. An accelerator architecture named data arrangement multi-DNN accelerator (DARMA) that features a data arrangement and distribution circuit and hierarchical memory for data dimension conversion is also presented. Since the mismatch is mitigated, the suggested accelerator outperforms current accelerators in terms of bandwidth and processing unit utilization. Through tests on VR/AR, MLperf, and other multitask applications, the evaluation results show that the proposed architecture provides both energy-efficiency and throughput improvements. Chuxi Li, Xiaoya Fan, Xiaoti Wu, Zhao Yang 0005, Meng Zhang 0047, Shengbing Zhang |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2021 | Hardware-Aware NAS Framework with Layer Adaptive Scheduling on Embedded SystemabstractNeural Architecture Search (NAS) has been proven to be an effective solution for building Deep Convolutional Neural Network (DCNN) models automatically. Subsequently, several hardware-aware NAS frameworks incorporate hardware latency into the search objectives to avoid the potential risk that the searched network cannot be deployed on target platforms. However, the mismatch between NAS and hardware persists due to the absent of rethinking the applicability of the searched network layer characteristics and hardware mapping. A convolution neural network layer can be executed on various dataflows of hardware with different performance, with which the characteristics of on-chip data using varies to fit the parallel structure. This mismatch also results in significant performance degradation for some maladaptive layers obtained from NAS, which might achieved a much better latency when the adopted dataflow changes. To address the issue that the network latency is insufficient to evaluate the deployment efficiency, this paper proposes a novel hardware-aware NAS framework in consideration of the adaptability between layers and dataflow patterns. Beside, we develop an optimized layer adaptive data scheduling strategy as well as a coarse-grained reconfigurable computing architecture so as to deploy the searched networks with high power-efficiency by selecting the most appropriate dataflow pattern layer-by-layer under limited resources. Evaluation results show that the proposed NAS framework can search DCNNs with the similar accuracy to the state-of-the-art ones as well as the low inference latency, and the proposed architecture provides both power-efficiency improvement and energy consumption saving. Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047 |
ASP-DAC | 1 |
| 2021 | SPACE: Sparsity Propagation Based DCNN Training Accelerator on Edge
Chuxi Li, Shengbing Zhang |
ICA3PP (2) | 3 |
| 2020 | ENAS oriented layer adaptive data scheduling strategy for resource limited hardware
Chuxi Li, Xiaoya Fan, Yuling Geng, Danghui Wang |
Neurocomputing | 1 |
| 2016 | EKF based distributed cooperative localization for a multirobot teamabstractThis paper studies distributed cooperative localization problem for a multirobot team with one leader and two followers. Each robot in the team is equipped local sensors and can exchange data with its neighbors through wireless communication network. A distributed localization algorithm is developed by using extended Kalman filter (EKF) scheme. In every sampling period, each member in the team estimates its local state based on its local measurements and neighbor's state estimation information sent from its neighbors at current sampling time or last sampling time. A simulation result shows that the algorithm is feasible. Chuxi Li, Jieying Lu, Weizhou Su |
ICARCV | 1 |