EDBT 2026 Demo / reviewers in the wild / expert
Yujie Zhang 0007
dblp:33/1685-7
· DBLP profile ↗
5ranked-venue papers
3as first author
4since 2021 · last 2025
0009-0003-5209-6980ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DAOP: Data-Aware Offloading and Predictive Pre-Calculation for Efficient MoE InferenceabstractMixture-of-Experts (MoE) models, though highly effective for various machine learning tasks, face significant deployment challenges on memory-constrained devices. While GPUs offer fast inference, their limited memory compared to CPUs means not all experts can be stored on the GPU simultaneously, necessitating frequent, costly data transfers from CPU memory, often negating GPU speed advantages. To address this, we present DAOP, an on-device MoE inference engine to optimize parallel GPU-CPU execution. DAOP dynamically allocates experts between CPU and GPU based on per-sequence activation patterns, and selectively pre-calculates predicted experts on CPUs to minimize transfer latency. This approach enables efficient resource utilization across various expert cache ratios while maintaining model accuracy through a novel graceful degradation mechanism. Comprehensive evaluations across various datasets show that DAOP outperforms traditional expert caching and prefetching methods by up to 8.20x and offloading techniques by 1.35x while maintaining accuracy. Yujie Zhang 0007, Shivam Aggarwal, Tulika Mitra |
DATE | 1 |
| 2025 | Para-Pipe: Exploiting Hierarchical Operator Parallelism of ML Computational Graphs on SoCsabstractAs edge-based deep learning applications become more complex, optimizing performance on heterogeneous System-on-Chips (SoCs) presents unique challenges. Traditional pipelining techniques distributing the computation across different on-chip processing units, while effective for throughput, do not address the latency demands posed by modern neural networks with complex interdependencies and extensive operator parallelism. There is a potential in leveraging operator parallelism to enable concurrent execution across multiple processing units, thereby reducing inference latency. However, prioritizing pipelining or parallel execution often necessitates a compromise, where optimizing one performance metric adversely impacts the other. This paper introduces Para-Pipe, a hierarchical mapping framework that integrates intra-and inter-stage operator parallelism within a pipelined architecture. Para-Pipe navigates the trade-off between throughput and latency by selectively fine-tuning parallelism levels within and across pipeline stages. This strategy can significantly reduce inter-processor communication overhead, significantly improving energy efficiency. Our evaluation demonstrates that Para-Pipe generates multiple Pareto-optimal configurations, achieving a balance between throughput and latency on an Amlogic SoC equipped with ARM big.LITTLE CPUs and GPU, as well as the Black Sesame Technology SoC featuring a deep learning accelerator and two DSPs. More importantly, throughput-optimized configurations under Para-Pipe on Amlogic SoC show an average energy efficiency improvement of 11.0% over purely pipelined strategies and 23.3% relative to non-pipelined parallel execution. Yujie Zhang 0007, Huiying Lan, Ehsan Aghapour, Peng Zan, Weidong Shao, Anuj Pathania, Tulika Mitra |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | Power-Performance Characterization of TinyML SystemsabstractTinyML systems are enabling machine learning (ML) inference at the edge. However, there exists little quantitative analysis of such systems. This paper presents a systematic performance and power characterization of diverse TinyML applications on micro-controllers (MCUs), spanning neural network models, software libraries, operating systems, and hardware architectures. We focus on the impact of the multiple layers of abstractions that provide higher programmability at the expense of performance and energy efficiency. We propose a model to estimate the costs of different abstraction layers and make recommendations for minimizing those costs. Our findings can help designers with Neural Architecture Search (NAS) and CNN inference optimization on edge devices. Yujie Zhang 0007, Dhananjaya Wijerathne, Zhaoying Li 0004, Tulika Mitra |
ICCD | 1 |
| 2022 | C2S: Class-aware client selection for effective aggregation in federated learningabstractFederated learning is proposed to train distributed data in a safe manner by avoiding to send data to server. The server maintains a global model and sends it to clients in each communication round, and then aggregates the updated local models to derive a new global model. Traditionally, the clients are randomly selected in each round and aggregation is based on weighted averaging. Researches show that the performance on IID data is satisfactory while significant accuracy drop can be observed for Non-IID data. In this paper, we explore the reasons and propose a novel aggregation approach for Non-IID data in federated learning. Specifically, we propose to group the clients according to classes of data they have, and select one set in each communication round. Local models from the same set are averaged as usual and the updated global model is sent to next group of clients for further training. In this way, the parameters are only averaged on similar clients and passed among different groups. Evaluation shows that the proposed scheme has advantages in terms of model accuracy and convergence speed with highly unbalanced data distribution and complex models. Mei Cao, Yujie Zhang 0007, Zezhong Ma, Mengying Zhao |
High Confid. Comput. | 2 |
| 2020 | Q-learning Based Backup for Energy Harvesting Powered Embedded SystemsabstractNon-volatile processors (NVPs) are used in energy harvesting powered embedded systems to preserve data across interruptions. In NVP systems, volatile data are backed up to non-volatile memory upon power failures and resumed after power comes back. Traditionally, backup is triggered immediately when energy warning occurs. However, it is also possible to more aggressively utilize the residual energy for program execution to improve forward progress. In this work, we propose a Q-learning based backup strategy to achieve maximal forward progress in energy harvesting powered intermittent embedded systems. The experimental results show an average of 307.4% and 43.4% improved forward progress compared with traditional instant backup and the most related work, respectively. Yujie Zhang 0007, Weining Song, Mengying Zhao, Zhaoyan Shen, Zhiping Jia |
DATE | 2 |