Yifu Zhu

dblp:362/4018 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2025
0009-0003-7938-814XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 State-driven fairness control for efficient I/O queue scheduling in NVMe virtualization
Yifu Zhu, Xin Kuang, Yanjie Tan, Huailiang Tan, Keqin Li 0001
Future Gener. Comput. Syst.2
2025 Dynamic DPU Offloading and Computational Resource Management in Heterogeneous Systems
abstract
DPU offloading has emerged as a promising way to enhance data processing efficiency and free up host CPU resources. However, unsuitable offloading may overwhelm the hardware and hurt overall system performance. It is still unclear how to make full use of the shared hardware resources and select optimal execution units for each tenant application. In this paper, we propose DORM, a dynamic DPU offloading and resource management architecture for multi-tenant cloud environments with CPU-DPU heterogeneous platforms. The primary goal of DORM is to minimize host resource consumption and maximize request processing efficiency. By establishing a joint optimization model for offloading decision and resource allocation, we abstract the problem into a mixed integer programming mathematical model. To simplify the complexity of model-solving, we decompose the model into two subproblems: a 0-1 integer programming model for offloading decision-making and a convex optimization problem for fine-grained resource allocation. Besides, DORM presents an orchestrator agent to detect load changes and dynamically adjust the scheduling strategy. Experimental results demonstrate that DORM significantly improves resource efficiency, reducing host CPU core usage by up to 83.3%, increasing per-core throughput by up to 4.61x, and lowering the latency by up to 58.5% compared to baseline systems.
Yanjie Tan, Yifu Zhu, Huailiang Tan, Keqin Li 0001
IEEE Trans. Computers3
2024 A HW/SW Co-Design of Video Dehazing Accelerator Using Decoupled Local Atmospheric Light Prior
abstract
In this paper, we introduce DLAPID, a novel decoupled parallel hardware-software co-design architecture for real-time video dehazing. From a software point of view, DLAPID isolates the atmospheric light operation from the initial transmission estimation to take full advantage of the hardware accelerators' parallelization features. For the hardware implementation, we deploy DLAPID both on FPGA and GPU platforms and validate its effectiveness. Using both real-world driving scenario testing sets and ground-truth datasets, we quantitatively and qualitatively assess the proposed method against several SOTA (state-of-the-art) video dehazing models. The outcomes of our experiments demonstrate that our approach achieves better dehazing performance with lower power consumption and has real-time processing capabilities, thereby preventing potential accidents of autonomous vehicles.
Yanjie Tan, Yifu Zhu, Feiteng Nie, Huailiang Tan
DAC2
2024 MTDA: Efficient and Fair DPU Offloading Method for Multiple Tenants
abstract
In modern cloud computing environment, the offloading potential of DPU must be fully exploited for multiple tenants. Existing DPU offloading techniques lack the capability to perform the fair allocation of a DPU domain's internal resources among tenants with various performance requirements. In this article, we propose a virtual multi-channel DPU offloading architecture for multiple tenants (MTDA) and implement it on a BlueField-2 DPU platform to achieve stability and fairness in resource allocation for generic datacenter tasks. MTDA provides an independent virtual channel for each tenant before their requests are submitted to avoid competition among tenants. Considering the diverse requirements of tenants, MTDA constructs a credit-based resource allocation model and a traffic-aware scheduling algorithm to fully utilize the rich computing resources of DPU and improve the fairness of DPU resource allocation. Experimental results show that MTDA increases the throughput by up to 101.2%, 143.2%, 36.1%, and 41.7%, lowers the latency by up to 50.3%, 58.9%, 26.6%, and 29.4%, improves the fairness by up to 98.8%, 99.0%, 98.3%, and 98.4%, and provides more stable performance for multi-tenants, compared with DPDK, iPipe, FairNIC, and LogNIC.
Yanjie Tan, Yifu Zhu, Huailiang Tan, Keqin Li 0001
IEEE Trans. Serv. Comput.3
2023 Brief Industry Paper: Real-Time Image Dehazing for Automated Vehicles
abstract
Autonomous vehicles at L2 and above are increasingly relying on stereo vision systems, where haze removal is critical to detect obstacles hidden in fog. Existing image dehazing techniques have low processing speed and high resource consumption, restricting their application scope in practice. In this work, we propose a hardware-software co-design solution for haze removal. It fully decouples calculation of the two main parameters, i.e., atmospheric light and transmittance, in the dehazing process. By eliminating the data dependency, parallelism in hardware acceleration is enhanced. Furthermore, in replacement of the conventional global homogeneous atmospheric light computation, we report a chunk-based heterogeneous method to reduce cache overhead. Our approach is implemented on FPGA, compared against five state-of-the-art (SOTA) works for image haze removal. Evaluation using test sets of real-world foggy driving scenarios shows that our object detection accuracy is over 88%, 9.5%-47.4% better than the SOTA works with neural networks (NN) on GPU, and 25.9%-52.2 % better than the SOTA works on FPGA. The processing speed varies with image resolution and our improvement is generally even more at higher resolution, which is 29.7% faster than the fastest SOTA. We have the lowest overall resource consumption, where the bottleneck BRAM usage is reduced by over 70%. The FPGA solution has circuit-level timing determinism at nanosecond, hence suitable for hard real-time applications.
Yanjie Tan, Yifu Zhu, Huailiang Tan, Wanli Chang 0001
RTSS2
2023 MAPD: An FPGA-Based Real-Time Video Haze Removal Accelerator Using Mixed Atmosphere Prior
abstract
Real-time video dehazing plays a key role in helping autonomous driving detect pedestrians or obstacles in severe foggy weather to prevent potential hazards. Existing video dehazing methods achieve good restoration performance but still suffer from oversaturation and low dehazing speed, especially for high-definition (HD, high-resolution) videos. In this article, we propose a mixed atmosphere prior information video dehazing accelerator (MAPD) and implement it on field programmable gate array (FPGA) to achieve real-time haze removal for HD video. MAPD provides a mixed atmospheric light model by applying heterogeneous atmospheric light in the foreground area to balance brightness deviation, and maintaining the global atmospheric light in the background region. Considering the parallel characteristics of FPGA, MAPD leverages the redundant information between adjacent frames to accelerate the dehazing process and designs an indirect transmission estimation to decrease resource consumption. For comparison, we also implement six dehazing solutions (DCP, color ellipsoid prior (CEP), RDCP, FFVD, MHVD, and REFD) on FPGA, and deploy a graphics processing unit (GPU)-based method$(D^{4})$on the platform with Nvidia 3080 GPU. Experiments using two widely used benchmarks show that MAPD increases the performance by up to 36.5%, 53.5%, 36.3%, 33.3%, 11.9%, and 23.3%, decreases resource consumption by up to 79.7%, 75.0%, 74.8%, 25.6%, 22.6%, and 73.9% and enhances FPS for HD videos by up to 241.6%, 145.9%, 151.7%, 68.6%, 50.6%, and 62.4%, compared with DCP, CEP, RDCP, FFVD, MHVD, and REFD. Compared to$D^{4}$, MAPD also promotes the dehazing performance by up to 21.8%, and increases FPS by up to 487.0%.
Yanjie Tan, Yifu Zhu, Huailiang Tan, Keqin Li 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2