Chaonong Xu

dblp:60/10412 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0003-1897-4797ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 10 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Collaborative Inference Acceleration with Non-Penetrative Tensor Partitioning
abstract
The inference of large-sized images on Internet of Things (IoT) devices is commonly hindered by limited resources, while there are often stringent latency requirements for Deep Neural Network (DNN) inference. Currently, this problem is generally addressed by collaborative inference, where the large-sized image is partitioned into multiple tiles, and each tile is assigned to an IoT device for processing. However, since significant latency will be incurred due to the communication overhead caused by tile sharing, the existing collaborative inference strategy is inefficient for convolutional computation, which is indispensable for any DNN. To reduce it, we propose Non-Penetrative Tensor Partitioning (NPTP), a fine-grained tensor partitioning method that reduces the communication latency by minimizing the communication load of tiles shared, thereby reducing inference latency. We evaluate NPTP with four widely-adopted DNN models. Experimental results demonstrate that NPTP achieves a 1.44-1.68× inference speedup relative to CoEdge, a state-of-the-art (SOTA) collaborative inference algorithm.
Zhibang Liu, Chaonong Xu, Zhenjie Lv, Zhizhuo Liu, Suyu Zhao
ICASSP2
2025 FGSMS: Fine-Grained SM Scheduling for Efficient Deep Learning Computing
Nanjian Zhou, Zhizhuo Liu, Chaonong Xu
ICIC (21)4
2024 Efficient Neural Network Fine-Tuning via Layer Contribution Analysis
Zhizhuo Liu, Nanjian Zhou, Zhibang Liu, Chaonong Xu
ICIC (4)5
2024 Low-Latency Deep Learning Inference Schedule on Multi-Core MCU
abstract
Emerging Artificial Internet-of-Things (AIoT) services based on Microcontroller Units (MCU) heavily harness Deep Learning (DL) to improve user experiences. Such DL-assisted services depend on fast Neural Network (NN) execution for high responsiveness, demanding tiny IoT devices to minimize the NN execution latency by efficiently utilizing their underlying hardware resources. However, existing inference frameworks cannot achieve satisfied real-time performance for Multi-core MCU (MMCU), now a mainstream platform for AIoT. We mention that two improvements can be made to speed up inference: 1) select appropriate data layout for each operator in NNs, and 2) exploit the capability of MMCU by properly partitioning each operator into multiple cores.In this paper, we propose a novel idea of joint data layout and Intra-Operator Parallelism (IOP) for low-latency DL inference on MMCU. We formulate the problem based on a self-built latency predictor, which predicts the execution latency of each operator within NNs on a given MMCU. An algorithm with time complexity being ${\mathcal{O}}\left({|E|M{N^3}}\right)$ is proposed to find an optimal scheduling plan, where N and M refer to the maximum number of possible layouts and IOP strategies for an operator. Our experimental evaluation demonstrates that our scheduling plan can achieve a speedup of 1.52×−3.37× for CMSIS-NN, a state-of-the-art edge inference software stack. Besides, compared with the state-of-the-art IOP execution system, our scheduling plan achieves a speedup of approximately 1.67×.
Chaonong Xu, Chao Li 0028, Weiming Kong
IJCNN1
2024 Complexity and algorithm of setting optimal location for data sink in real-time NOMA-based IIoTs
Chaonong Xu, Chao Li 0028
Comput. Commun.1
2023 Enhance Broadcasting Throughput by Associating Network Coding with UAVs Relays Deployment in Emergency Communications
Chaonong Xu
CollaborateCom (3)1
2023 Minimizing Peak Memory Footprint of Inference on IoTs Devices by Efficient Recomputation
Xiaofeng Sun 0001, Chaonong Xu
ICIC (5)2
2023 Hybrid Parallel Inference for Large Model on Heterogeneous Clusters for High Throughput
abstract
In high-throughput intelligent computing scenarios, multi-device parallelism strategies based on data parallelism or pipeline parallelism have been extensively utilized to accelerate large deep neural network model inference. Data parallelism offers nearly linear improvement in inference speed, but it is limited by the memory capacity of a single device which constrains the model size. On the other hand, pipeline parallelism can support larger models, but the total communication of the activations among devices is high, which limits the improvement of the inference speed. To address the demand for efficient model inference in high-throughput heterogeneous scenarios, we proposes a hybrid parallelism strategy that combines data parallelism and pipeline parallelism. The strategy involves grouping heterogeneous device clusters and then employing inter-group data parallelism along with intra-group pipeline parallelism. Moreover, we propose an algorithm to find an optimal hybrid parallel inference strategy with maximum throughput. The control variables of the strategy includes the number of groups, group-device assignments and model partition ratios. Our experimental evaluation demonstrates that compared to PipeEdge, a pipeline parallel inference framework for heterogeneous cluster, our strategy can achieve 1.7× −3.4× acceleration in an 8-device heterogeneous cluster without loss of accuracy.
Chaonong Xu, Weiming Kong, Chao Li 0028, Luqi Gong
ICPADS1
2023 A novel second-order learning algorithm based attention-LSTM model for dynamic chemical process modeling
Baochang Xu, Likun Yuan, Chaonong Xu
Appl. Intell.4
2023 IMF2O2: A Fully Connected Sensor Deployment Algorithm for Underwater Sensor Networks
abstract
To address the problems of node deployment schemes in existing underwater sensor networks that lack consideration of network connectivity and high deployment costs, this article constructs an optimization model that maximizes network coverage and minimizes deployment costs while ensuring full connectivity. For the NP-hard property of this optimization model, an improved moth flame optimization node deployment algorithm based on fuzzy operators (IMF 2 O 2 ) is proposed. First, comprehensively considering the two performance metrics of network coverage and network connectivity, a multi-objective selection mechanism based on fuzzy operators is proposed to improve network coverage while ensuring full connectivity. Second, a fixed number of nodes are used to monitor the target event points, transforming the node deployment of sensors into an optimal problem and proposing an improved moth flame optimization algorithm to solve this problem. Finally, the two metrics of coverage and deployment cost are measured and the fuzzy operator is used to select the optimal number of nodes to be deployed. Numerical results showed that the proposed algorithm improved network coverage rate by 10%, 22%, and 25%, and improved network connectivity rate by 12%, 20%, and 8% as compared to PSSD, RAWS, and VODA, respectively, while ensuring full connectivity.
Na Xia, Bin Chen 0006, Huazheng Du, Chaonong Xu, Rong Zheng 0001
ACM Trans. Sens. Networks5
2022 Complexity of minimum uplink power scheduling with delay bound for Backbone-assisted PD-NOMA wireless networks
Chaonong Xu, Yutong Zhu, Chao Li 0028
Comput. Networks2
2021 Optimal data sink location for real-time NOMA-based Industrial IoTs
abstract
Real-time performance is one of the most vital metrics for applications in Industrial Internet of Things (IIoTs), and the relative geographic relationship between data sink and wireless sensors has great influence on the real-time performance. Since the locations of wireless sensors are in generally fixed in IIoTs, setting reasonable location for data sink is an efficient way for improving the real-time performance. In this paper, we investigate Non-Orthogonal Multiple Access (NOMA) based IIoTs, and consider how to minimize average access delay by setting suitable location for data sink. We formulate the problem and present an algorithm by mapping the problem into the classic minimum chain covering problem, and make the problem algorithm-tractable. Simulation results reveal that due to the full exploitation of NOMA parallelism, average access delay decreases more than 60% for some typical settings, and it can even reach 70% for the linear network topology.
Chaonong Xu, Chao Li 0028
IPCCC2
2020 Reliable uplink transmissions for NOMA-based Industrial Wireless Networks with guaranteed real-time performance
Chaonong Xu, Jianxiong Wu, Chao Li 0028
Comput. Commun.1
2019 Low-complexity uplink scheduling algorithms with power control in successive interference cancellation based wireless mud-logging systems
Chaonong Xu, Haichuan Ding, Yongjun Xu 0001
Wirel. Networks1
2018 Complexity of minimum uplink scheduling in backbone-assisted successive interference cancellation-based wireless networks
Chaonong Xu, Kaichi Ma, Yongjun Xu 0001
Comput. Networks1
2017 Optimal Power Scheduling for SIC-Based Uplink Wireless Networks with Guaranteed Real-Time Performance
Chaonong Xu, Kaichi Ma, Yongjun Xu 0001
WASA1
2017 Localizability Judgment in UWSNs Based on Skeleton and Rigidity Theory
abstract
Underwater sensor networks (UWSNs) have been investigated in a variety of applications such as sea resources reconnaissance, pollution monitoring and tactical monitoring. In 3D underwater environments, it is a key topic to judge the localizability of sensor nodes given known locations of a small set of anchor nodes. In this paper, a novel localizability judgment method for UWSNs is proposed based on rigidity theory. A UWSN is modelled as an undirected graph based on acoustic connectivity. The graph is then reduced to a subgraph with global rigidity, called skeleton, from which the set of localizable sensors can determined. Furthermore, the Analytic Hierarchy Process (AHP) is used to evaluate the localization confidence of localizable sensors. Extensive simulations demonstrate that the proposed localizability judgment method can achieve low false negative rate and high efficiency networks of different sensor numbers and sensor densities. It is also shown to perform well in dynamic networks with relatively low waterflow speed.
Na Xia, Yuanxiao Ou, Shiliang Wang, Rong Zheng 0001, Huazheng Du, Chaonong Xu
IEEE Trans. Mob. Comput.6
2013 Odometer in the pocket
abstract
Some previous work has shown the feasibility of Pedestrian Dead Reckoning (PDR) using a mobile phone, but the estimation of length walked is still a big challenge. In this paper, we propose a formula for estimating walking velocity, which can be integrated to get distance.
Lin Wu 0006, Yongjun Xu 0001, Zhulin An, Chaonong Xu, Fei Wang 0014
MobiSys4