Yanming Chen 0002

dblp:68/6802-2 · DBLP profile ↗
← Back
20ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0002-2747-6637ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Corrigendum: LyDRL: Lyapunov-guided Deep Reinforcement Learning for Stable Task Offloading in Connected Autonomous Vehicles
abstract
This is a corrigendum for the article “LyDRL: Lyapunov-guided Deep Reinforcement Learning for Stable Task Offloading in Connected Autonomous Vehicles” published in ACM Trans. Autonom. Adapt. Syst. 20, 3, Article 24 (September 2025), 29 pages.
Yanming Chen 0002, Yiwen Zhang 0001, Weiwei Fang, Naixue Xiong
ACM Trans. Auton. Adapt. Syst.1
2025 Accelerating Diffusion Models via Parallel Denoising
Yanming Chen 0002, Zixin Ma, Chuanguang Yang, Zhulin An, Yiwen Zhang 0001
ACM Multimedia1
2025 LyDRL: Lyapunov-guided Deep Reinforcement Learning for Stable Task Offloading in Connected Autonomous Vehicles
abstract
Task offloading is recognized as a promising approach to enhance the computational performance of Connected Autonomous Vehicles (CAVs). Some applications of CAVs, such as metaverse applications, require substantial resources, posing significant challenges to CAVs with limited computing and storage capacities. CAVs can offload resource-intensive applications to the Vehicular Edge Computing (VEC) server, which has strong computing capabilities. To fully utilize the resources in the CAV system, partial offloading is employed. However, the local computing resources are limited for continuously generated partial offloading tasks. This results in many partially locally executed tasks experiencing long processing times or being discarded, which is detrimental to delay-sensitive tasks on CAVs. This article proposes Lyapunov function-guided reinforcement learning for the CAVs task offloading computational framework, LyDRL. Specifically, LyDRL first uses the Lyapunov function to transform the long-term objective optimization problem into subproblems determined at each time slot. In each time slot, deep reinforcement learning is used to obtain the optimal task offloading decision while satisfying the constraints. Simulation results show that compared with the existing algorithms, the proposed strategy can ensure the stability of the CAVs system and achieve the lowest system overhead.
Yanming Chen 0002, Yiwen Zhang 0001, Weiwei Fang, Naixue Xiong
ACM Trans. Auton. Adapt. Syst.1
2024 Lyapunov-guided Deep Reinforcement Learning for Vehicle task Stable offloading
abstract
Cellular vehicular-to-everything (C-V2X), a critical Internet of Vehicles (IOV) technology is promised to be enhanced and strengthened to improve road traffic safety and achieve intelligent transportation in the 5G era. However, computation-intensive and latency-sensitive computation tasks of autonomous driving have created a great challenge for computation and storage-limited vehicles. Vehicular edge computing (VEC) is envisioned as a promising approach to processing the explosive computation tasks of vehicular users (VU). In the VEC system, each VU allocates to process partial tasks through offloading and the remaining tasks through local execution. In practical scenarios, the number of vehicles and the arrival of vehicle tasks are random, leading to a highly complex environment for VEC systems. To solve this problem, we propose a novel framework, named LYDDPG, that combines the advantages of Lyapunov optimization and deep reinforcement learning (DRL) to ensure the stability of the system during task offloading.
Yanming Chen 0002, Yiwen Zhang 0001
CSCWD2
2024 End-edge collaborative DNN inference acceleration via E-CARGO and RBC
abstract
Nowadays, a wide range of intelligent applications rely on deep neural networks (DNNs), ranging from face recognition to autonomous driving. Inference on pre-trained DNN is accurate and efficient, but resource-intensive, especially for end devices such as smartphone and wearable devices. To address the associated resource constraints, DNN inference tasks are often offloaded to the edge or cloud, achieved by partitioning the DNN and offloading part of the computation to another device. However, most of the existing solutions usually adopt more complex methods such as reinforcement learning and heuristic algorithms. In contrast, this paper introduces a simple approach by modeling the end-edge collaborative DNN inference system via the Environments - Classes, Agents, Roles, Groups, Objects (E-CARGO) model, and Role-Based Collaboration (RBC) methodology. A DNN partition point selection algorithm is proposed and the DNN task assignment problem in the system is formulated as a Group Multi-Role Assignment(GMRA) problem to be solved. Extensive simulation experiments demonstrate that the proposed solution can effectively reduce the global delay of DNN inference.
Wenying Peng, Yanming Chen 0002, Haibin Zhu 0001, Yiwen Zhang 0001
CSCWD2
2024 MicroENet: An Efficient Network for MCUs with Low Model Parameters and Peak Memory
abstract
Machine learning (ML) is increasingly vital for IoT applications and Industry 4.0. Compared to uploading data to the cloud for ML inference. Utilizing ML for local low-power IoT device analysis can conserve energy and ensure data privacy by avoiding cloud uploads. However, the inference of convolutional neural networks (CNNs) usually requires large intermediate activation maps and involves substantial parameters. Most IoT devices have only <320KB SRAM and <2MB Flash. To address these resource constraints, this paper proposes a novel model MicroENet with only hundreds of KB peak memory and parameters. Firstly, the memory bottleneck lies in the first few blocks of the CNNs so that we reduce the output channels of the first layer of the model and employ efficient downsampling in the second layer to decrease the image size quickly, bypassing the large activation layer. Then, the enhanced attention depthwise blocks (MCU-Blocks) are proposed, which have high parametric efficiency. Based on these blocks, we develop a tiny architecture with an ImageNet accuracy of 63.9%. Impressively, this result is achieved using only 245KB peak memory and 0.96 million parameters. In the visual wake word experiment, the size of our model is further reduced and achieves 89.64% accuracy with 28KB peak memory. Finally, MicroENet is deployed on the STM32F746 microcontroller for the image classification task. The experimental results show that our model outperforms others with similar peak memory and parameters.
Yanming Chen 0002, Yiwen Zhang 0001
CSCWD2
2024 An Intelligent Co-Scheduling Framework for Efficient Super-Resolution on Edge Platforms With Heterogeneous Processors
abstract
Deep neural networks (DNNs) have shown remarkable performance in the super-resolution (SR) task, which can upscale low-resolution images to satisfy application demands on image quality. However, the high computational intensity of DNN models poses a challenge to executing SR tasks on resource-constrained edge platforms. To leverage heterogeneous computational resources (e.g., CPU, GPU, and NPU) to speed up image reconstruction through concurrent inference, we propose a novel framework, called ESHP, for Efficient Super-resolution on edge platforms with Heterogeneous Processors. Our proposed ESHP framework boasts several advantageous characteristics: 1) it substantially speeds up SR processing over the existing approaches by leveraging all available heterogeneous hardware; 2) it uses deep reinforcement learning (DRL) to enable adaptive and optimal scheduling based on runtime states; 3) it strikes a balance between SR performance and computational cost during inference; and 4) it does not modify the original architecture of given SR model. We have conducted extensive experiments on typical edge platforms with popular SR models and resolution datasets of different scales, which verify the effectiveness and the versatility of our ESHP against other commonly-used baselines.
Weiwei Fang, Liang Qian, Yanming Chen 0002, Naixue Xiong
IEEE Internet Things J.4
2024 Multi-scale conditional reconstruction generative adversarial network
Yanming Chen 0002, Zhulin An, Fuzhen Zhuang
Image Vis. Comput.1
2024 EdgeCI: Distributed Workload Assignment and Model Partitioning for CNN Inference on Edge Clusters
abstract
Deep learning technology has grown significantly in new application scenarios such as smart cities and driverless vehicles, but its deployment needs to consume a lot of resources. It is usually difficult to execute inference task solely on resource-constrained Intelligent Internet-of-Things (IoT) devices to meet strictly service delay requirements. CNN-based inference task is usually offloaded to the edge server or cloud. However, it may lead to unstable performance and privacy leaks. To address the above challenges, this article aims to design a low latency distributed inference framework, EdgeCI, which assigns inference tasks to locally idle, connected, and resource-constrained IoT device cluster networks. EdgeCI exploits two key optimization knobs, including: (1) Auction-based Workload Assignment Scheme (AWAS), which achieves the workload balance by assigning each workload partition to the more matching IoT device; (2) Fused-Layer parallelization strategy based on non-recursive Dynamic Programming (DPFL), which is aimed at further minimizing the inference time. We have implemented EdgeCI based on PyTorch and evaluated its performance with VGG-16 and ResNet-34 image recognition models. The experimental results prove that our proposed AWAS and DPFL outperform the typical state-of-the-art solutions. When they are well combined, EdgeCI can improve inference speed by 34.72% to 43.52%. EdgeCI outperforms the state-of-the art approaches on our edge cluster.
Yanming Chen 0002, Weiwei Fang, Naixue Xiong
ACM Trans. Internet Techn.1
2023 Medical Image Super-Resolution via Diagnosis-Guided Attention
abstract
Medical image super-resolution (SR) is an important medical image processing task and is often helpful for downstream medical analysis tasks. Most of the conventional SR methods tried to generate visually more convincing images whereas ignoring the following downstream tasks. In this paper, we take the Alzheimer’s disease diagnosis as the downstream task and propose a novel diagnosis-guided medical image SR network, which can make the SR and diagnosis be boosted by each other. The method contains two sub-networks, i.e., the SR network and the diagnosis network. To achieve better diagnosis performance, in the SR network, we apply the deformable convolution to capture the regions of interest (ROIs) with different and irregular sizes and shapes, which are important for diagnosis. Moreover, to integrate the two tasks, i.e., SR and diagnosis, more profoundly, we design a novel diagnosis-guided attention module, which makes the key regions for diagnosis can be reconstructed more clearly by the SR network. The extensive experiments on medical image data sets show that the proposed method often outperforms other state-of-the-art SR methods, which demonstrates its effectiveness. The codes of this paper are released in https://github.com/WJingwei/SRDA.
Peng Zhou 0006, Xianjun Han, Yanming Chen 0002
ICME4
2023 Intensifying The Consistency of Pseudo Label Refinement for Unsupervised Domain Adaptation Person Re-Identification
abstract
Clustering-based unsupervised domain adaptation (UDA) for person re-identification aims to learn in the unlabeled target domain. However, the noise problem of clustering-based generated pseudo labels remains under-explored, and these wrong labels can mislead the feature learning process. In this paper, we propose a consistent and intensive pseudo label refinement method in which pseudo labels generated in two different feature spaces, local and global, refine each other to improve the pseudo label quality of the final clusters. Then we utilize a quantitative criterion to measure label inaccuracy and fine-tune the target domain to reduce the noise by an inaccuracy-guided pseudo label optimization scheme. On a strong benchmark, we demonstrate the superiority of the method with extensive experiments. Specifically, our method outperforms the baseline by 7.2% mAP on the Duke2Market task and 0.5% mAP on the Market2MSMT task over the state-of-the-art.
Linfan Zha, Yanming Chen 0002, Peng Zhou 0006, Yiwen Zhang 0001
ICME2
2023 Using Less but Important Information for Feature Distillation
Yanming Chen 0002, Choonghyun Lee, Yong Gong
ICONIP (1)2
2022 FPAR: Filter Pruning Via Attention and Rank Enhancement
abstract
In recent years, deep convolutional neural networks (CNNs) have become larger than ever, thus their deployment on edge devices becomes difficult. There are numerous popular methods to accelerate the networks; however, many of these methods consider only the importance of a single filter to the network and neglect the coorelation between filters. To solve this problem, we propose a novel filter pruning method, called Filter Pruning via Attention and Rank Enhancement (FPAR), based on the attention mechanism and rank of feature maps. Moreover, the inspiration for it comes from a discovery: For a network with attention modules, irrespective of the batch of input images, the mean of channel-wise weights of the attention module is almost constant. Thus, we can use a few batches of input data to obtain this indicator to guide pruning. With extensive experiments on various datasets, demonstrate that our method outperforms the most advanced methods with similar accuracy. For example, using VGG-16, we removed 62.8% of floating-point operations (FLOPs) even with a 0.24% of the accuracy increase compared with the unpruned network.
Yanming Chen 0002, Mingrui Shuai, Shubin Lou, Zhulin An, Yiwen Zhang 0001
ICME1
2022 FRATCF: Feature-Residue Real-Time UAV Tracking Based on Automatic Spatio-Temporal Regularization Correlation Filter
abstract
Traditional discriminative correlation filter (DCF) has received widespread popularity due to its high computational efficiency. However, most of the existing DCF-based trackers improve the learning of the target object by introducing some simple regularization methods in the detection stage, which may easily lose the tracking target in scenes with background clutter, fast-moving cameras and similar targets. We propose a feature residual filter with automatic spatio-temporal regularization, namely FRATCF, which can be strengthened the filter learning by introducing the feature residual between two adjacent frames in the training phase. Extensive experiments are conducted on two challenging unmanned aerial vehicle (UAV) benchmarks, i.e., UAV123@10fps and DTB70. Results prove that our tracker runs at ∼43 FPS on an extremely cheap configuration, which is about twice the speed of AutoTrack. The performance is also better than other state-of-the-art (SOTA) trackers.
Yanming Chen 0002, Yueqing Jing, Peng Zhou 0006, Yiwen Zhang 0001
ICME2
2022 FPC: Filter pruning via the contribution of output feature map for deep convolutional neural networks acceleration
Yanming Chen 0002, Yiwen Zhang 0001, Qiang He 0001
Knowl. Based Syst.1
2021 CCPrune: Collaborative channel pruning for learning compact convolutional networks
Yanming Chen 0002, Yiwen Zhang 0001, Weisong Shi
Neurocomputing1
2020 Distributed Color-Based Particle Filter for Target Tracking in Camera Network
Yueqing Jing, Yanming Chen 0002
CollaborateCom (2)2
2020 A deep neural network compression algorithm based on knowledge transfer for edge devices
Yanming Chen 0002, Chao Li 0028, Luqi Gong, Yiwen Zhang 0001, Weisong Shi
Comput. Commun.1
2017 An image-based near-duplicate video retrieval and localization using improved Edit distance
Qingjie Zhao, Peng Lv 0001, Yanming Chen 0002
Multim. Tools Appl.5
2017 Multiple cues-based active contours for target contour tracking under sophisticated background
Peng Lv 0001, Qingjie Zhao, Yanming Chen 0002, Liujun Zhao
Vis. Comput.3