Zhenming Chen

dblp:34/3134 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Theory of computation · 3 · 2 first-authorSystems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 TTEdit: Cross-Modal Fusion with Diffusion Models for Detail-Aware Fashion Editing
Junliang Tan, Zhenming Chen
MMM (2)3
2026 Coconut: Multilevel Collaborative Deployment for Real-Time Deep Learning Tasks in Heterogeneous Edge GPU Cluster
Changyao Lin, Zhenming Chen, Jie Liu 0001
IEEE Internet Things J.2
2026 Cola: Cross-Processor Operator Parallelism for Asynchronous Deep Learning Inference
Changyao Lin, Zhenming Chen, Jie Liu 0001
IEEE Trans. Mob. Comput.2
2025 TMFit: Enhancing Fashion Image Editing Precision via Text-Driven Mask Generation
Zhenming Chen
PRCV (6)2
2025 AnyView VTON: Consistent 3D Virtual Try-On with View-Conditioned Diffusion Model
Maochun Zhang, Zhenming Chen
PRCV (10)3
2025 E3: Early Exiting with Explainable AI for Real-Time and Accurate DNN Inference in Edge-Cloud Systems
abstract
Edge intelligence applications frequently generate deep learning inference tasks with varying Service Level Objectives (SLO, such as accuracy and real-time requirements). For such tasks, recent progressive inference modes support early exit from different branches to satisfy inference requirements. However, existing edge-cloud progressive neural architectures cannot simultaneously achieve high accuracy and real-time performance for different data features. Therefore, we utilize explainable AI technique to construct and train a novel progressive neural architecture E3. E3 can progressively extract the most important features for inference, ensuring higher accuracy at early-exit points. While the less important features in the later stage are highly compressible, thereby reducing edge-cloud transmission overhead. Furthermore, E3 cooperates with online execution control to launch tasks and decide the exit point for each task, ensuring resource utilization and real-time performance, and adapting to bandwidths and deadlines. Experimental results on various edge-cloud platforms, datasets, and reference models demonstrate that E3 is more lightweight, efficient, energy-saving, and incurs almost no additional runtime overhead compared to traditional architectures. Under stringent deadlines, the average accuracy of tasks increases by > 50%, and the deadline satisfaction rate approaches 100%.
Changyao Lin, Zhenming Chen, Jie Liu 0001
SenSys2
2025 TOP: Task-Based Operator Parallelism for Asynchronous Deep Learning Inference on GPU
abstract
Current deep learning compilers have made significant strides in optimizing computation graphs for single- and multi-model scenarios. However, they lack specific optimizations for asynchronous multi-task inference systems. In such systems, tasks arrive dynamically, leading to diverse inference progress for each model. This renders traditional optimization strategies based solely on the original computation graph suboptimal or even invalid. Furthermore, existing operator scheduling methods do not account for parallel task pipelines involving the same model. Task pipelines present additional opportunities for optimization. Therefore, we propose Task-based Operator Parallelism (TOP). TOP incorporates an understanding of the impact of task arrival patterns on the inference progress of each model. It leverages the multi-agent reinforcement learning algorithm MADDPG to cooperatively optimize the task launcher and model scheduler, generating an optimal pair of dequeue frequency and computation graph. The objective of TOP is to enhance resource utilization, increase throughput, and allocate resources judiciously to prevent task backlog. To expedite the optimization process in TOP, we introduce a novel stage partition method using the GNN-based Policy Gradient (GPG) algorithm. Through extensive experiments on various devices, we demonstrate the efficacy of TOP. It outperforms the state-of-the-art in operator scheduling for both single- and multi-model task processing scenarios. Benefiting from TOP, we can significantly enhance the throughput of a single model by increasing its concurrency or batch size, thereby achieving self-acceleration.
Changyao Lin, Zhenming Chen, Jie Liu 0001
IEEE Trans. Parallel Distributed Syst.2
2024 Poster Abstract: Xpi: Real-Time Progressive Inference Serving with Explainable AI in Edge-Cloud Systems
abstract
The constrained computing and memory resources at the edge pose challenges for satisfying different service-level objectives (SLOs) of deep learning inference requests. In this paper, we propose a novel edge-cloud progressive inference framework Xpi, which integrates explainable AI technique to facilitate early-exit, and learning-based online execution control to satisfy different SLOs and optimize edge resource overheads. We implement Xpi on an edge-cloud platform, and conduct partial experiments on two datasets. Xpi outperforms several advanced edge-cloud progressive inference frameworks in terms of accuracy and deadline satisfaction rate.
Changyao Lin, Zhenming Chen, Jie Liu 0001
IPSN2
2007 A Constant Approximation Algorithm for Interference Aware Broadcast in Wireless Networks
abstract
Broadcast protocols play a vital role in multihop wireless networks. Due to the broadcast nature of radio signals, a node's interference range can be larger than its transmission range, i.e., it can interfere with other node's reception even if the latter is not within its transmission range. To design an efficient broadcast protocol, both the collision and the interference among multiple transmissions must be addressed. However, most of the previous works on wireless broadcast protocols either treated interference in the same way as collision or did not consider interference at all. In this paper, we study a more general model in which interference is distinguished from collision, and propose a simple and yet efficient interference and collision free broadcast protocol. Our objective is to minimize the makespan, i.e., the earliest time such that every node receives the message. By exploiting the geometry property of the nodes that interfere with each other, we show that our algorithm is a constant approximation algorithm, it guarantees to deliver the message to all nodes within a small constant factor of the optimal makespan. We apply our algorithm under both the unit disk graph model and the more realistic radio irregularity model. The experimental results show that our algorithm consistently outperforms the previous algorithms.
Zhenming Chen, Chunming Qiao, Jinhui Xu 0001, Taekkyeun Lee
INFOCOM1
2006 Robustness of k-gon Voronoi diagram construction
Zhenming Chen, Evanthia Papadopoulou, Jinhui Xu 0001
Inf. Process. Lett.1
2005 Efficient geometric techniques for reconstructing 3D vessel trees from biplane image
abstract
No abstract available.
Lopamudra Mukherjee, Jinhui Xu 0001, Kenneth R. Hoffmann, Zhenming Chen
SCG6
2004 An Efficient Algorithm for Determining 3-D Bi-plane Imaging Geometry
Jinhui Xu 0001, Zhenming Chen, Kenneth R. Hoffmann
ICCSA (3)3
2004 Efficient Job Scheduling Algorithms with Multi-type Contentions
Zhenming Chen, Jinhui Xu 0001
ISAAC1