EDBT 2026 Demo / reviewers in the wild / expert
Yudong Mu
dblp:414/7051
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0005-9629-1184ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A2RT: Efficient Ray Tracing Accelerator with Approximate-Accurate Computing and QuantizationabstractRay tracing (RT) has revolutionized photorealistic rendering by simulating light transport, but existing methods face a trade-off between computational efficiency and rendering accuracy. To address this, we present A2RT, a software-hardware co-designed RT accelerator employing the end to end optimization of "quantization → computation". On the software side, we introduce a customized data flow mechanism with type-specific quantization for bounding boxes, ray origins, and directions, and we organize BVH nodes into Group- and Sub-Nodes. At the hardware level, a heterogeneous RT engine allocates resources based on node criticality: accurate computing units handle Group-Nodes, while approximate units process Sub-Nodes. A custom INT-FLOAT approximate multiplier further accelerates the approximate units. Experimental results show that A2RT achieves 45.51% energy consumption and 2.29× speedup over RT Core, and consumes 81.79% of energy while delivering 1.57× performance improvement compared to state-of-the-art accelerators. Zhihua Fan, Yudong Mu, Zhen Wang 0045, Xiaochun Ye, Xuejun An |
DATE | 4 |
| 2025 | NFMap: Node Fusion Optimization for Efficient CGRA Mapping with Reinforcement Learning
Yudong Mu, Zhihua Fan, Xuejun An, Xiaochun Ye |
APPT | 1 |
| 2025 | FDHA: Fusion-Driven Heterogeneous Accelerator for Efficient Diffusion Model Inference
Yudong Mu, Zhihua Fan, Xiaoxia Yao, Honglie Wang, Xuejun An, Xiaochun Ye |
Euro-Par (2) | 1 |
| 2025 | GenCNN: A Partition-Aware Multi-Objective Mapping Framework for CNN Accelerators Based on Genetic AlgorithmabstractConvolutional Neural Networks (CNNs) require partitioning to efficiently run on CNN accelerators, which offer multiple parallel processing dimensions, such as Processing Element (PE) array topologies and Single Instruction Multiple Data (SIMD) execution. The choice of parallelization strategy directly impacts accelerator performance. However, the vast search space for CNN partitioning and parallelization makes manual optimization costly and complex, especially when addressing both aspects simultaneously. This highlights the need for an automated framework to efficiently map CNNs onto accelerators. Our key insight is that existing approaches suffer from inadequate accelerator performance modeling and a lack of multi-objective optimization strategies that jointly consider task partitioning and convolution parallelization. To address this, we propose GenCNN, a multi-objective genetic algorithm-based mapping framework for CNN accelerators. GenCNN first constructs a fine-grained performance model that captures both off-chip data access and on-chip data processing. It then applies the Non-dominated Sorting Genetic Algorithm II improved by Multi-Objective Bayesian Optimization to derive a Pareto-optimal partitioning and parallelization strategy that balances off-chip latency and PE utilization. Finally, GenCNN optimizes scheduling and routing to minimize data transfers. Experimental results show that GenCNN achieves up to 17.66× speedup in compilation and 6.47× in execution compared with state-of-the-art mapping frameworks. Yudong Mu, Zhihua Fan, Xuejun An, Dongrui Fan, Xiaochun Ye |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | A RISC-V Extended Infrastructure for CNNs Through Pipelined Computing and Data Dependence OptimizationabstractWith the rapid development of artificial intelligence (AI), convolutional neural networks (CNNs) have been widely applied in fields like computer vision and recommendation systems. This growth has intensified the demand for hardware acceleration of CNNs. Existing accelerators are either designed as co-processors or improve performance through extended instructions. While these methods can significantly improve performance, they often result in limited programming and execution flexibility. In this paper, we design custom RISC-V instructions specifically for CNNs to maximize data reuse and exploit parallelism. Then, to efficiently execute CNNs instructions, we extend a Pipelined Vector Computing Unit (PPVCU). Finally, we incorporate Pattern Detection Logic (PDL) to identify common data dependence patterns in CNNs, enabling the Data Dependence Computing Unit (DDCU) to process instructions within each pattern in parallel. Experimental results show that our approach achieves, on average, 9.54× performance improvement and 6.7× energy efficiency improvement compared to our baseline, 8.34× performance improvement and 3.1× energy efficiency improvement compared to state-of-the-art designs. Teng Luo, Tengfei Xia, Zhihua Fan, Yudong Mu, Xuejun An, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |