Haikai Zhao

dblp:389/4359 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2026
0009-0001-3830-4275ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › federated learning
client selection
0.812024
FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management · SenSys 2024
Machine learning › Efficient and distributed learning
federated learning
0.812024
FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management · SenSys 2024
Machine learning › Efficient and distributed learning › federated learning › resource-efficient federated learning
memory-efficient federated learning
0.812024
FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management · SenSys 2024
Machine learning › Efficient and distributed learning › model compression › neural network compression
activation compression
0.212024
FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management · SenSys 2024
Machine learning › Efficient and distributed learning
model compression
0.212024
FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management · SenSys 2024

Methods — techniques the papers use, named apart from their topics

tensor compression · 0.8recomputation · 0.8
YearPublicationVenuePosition
2026 Understanding the Performance of Native Execution in Big Data Engines: The Good, the Bad, and How to Fix It
Haikai Zhao, Zhenman Fang
EDBT1
2026 FORCv2: A High-Throughput Streaming FPGA Accelerator for Optimized Row Columnar File Format Processing in Big Data Engines
abstract
To enhance the storage efficiency of large datasets, Big Data analytics commonly rely on columnar file formats, such as Apache ORC (Optimized Row Columnar), to encode and compress data. These formats significantly reduce storage requirements and improve query performance by supporting efficient data organization and compression techniques. However, the use of such file formats introduces a new challenge: while high-bandwidth SSDs mitigate the I/O bottleneck, the computational burden of decompressing and decoding these formats shifts the bottleneck to the CPU. This computational overhead becomes a critical challenge in achieving high-throughput data processing. This work introduces FORCv2, a high-performance streaming-based FPGA accelerator overlay designed to overcome these limitations. First, FORCv2 integrates multi-PE (Processing Engine) Zlib decompression, which leverages Xilinx’s high-performance single-PE implementation and enhances performance with multi-way parallelism. Second, it features fully pipelined data decoding, optimized to achieve throughput close to the memory bandwidth of a single High-Bandwidth Memory (HBM) channel, ensuring high-efficiency processing of ORC files. Third, FORCv2 incorporates flexible and efficient filtering capabilities, supporting both pre-stored indices and user-defined filter conditions for versatile data processing. Finally, it provides seamless dataflow integration with the Apache ORC library, enabling efficient end-to-end processing of ORC file formats. The proposed design leverages the inherent parallelism and HBM capabilities of AMD/Xilinx Alveo U280 FPGAs to deliver exceptional performance. By employing a fully pipelined architecture and resource-efficient design, FORCv2 achieves up to 2.8 GB/s of overall throughput (input throughput). Considering all operators combined, on average, it provides end-to-end speedups of approximately 52x for the synthetic dataset and 39x for the TPC-H dataset, versus the ORC C++ library running on a dual-socket system with 48 CPU threads. Experimental evaluations demonstrate the effectiveness of FORCv2 across diverse workloads, showcasing its potential to handle real-world Big Data processing scenarios efficiently. The FORCv2 implementation will be open sourced at https://github.com/SFU-HiAccel/FORC .
Abdul Wadood, Alec Lu, Haikai Zhao, Zhenman Fang
ACM Trans. Reconfigurable Technol. Syst.3
2024 FedHybrid: Breaking the Memory Wall of Federated Learning via Hybrid Tensor Management
abstract
Federated Learning (FL) emerges as a new learning paradigm that enables multiple devices to collaboratively train a shared model while preserving data privacy. However, one fundamental and prevailing challenge that hinders the deployment of FL on mobile devices is the memory limitation. This paper proposes FedHybrid, a novel framework that effectively reduces the memory footprint during the training process while guaranteeing the model accuracy and the overall training progress. Specifically, FedHybrid first selects the participating devices for each training round by jointly evaluating their memory budget, computing capability, and data diversity. After that, it judiciously analyzes the computational graph and generates an execution plan for each selected client in order to meet the corresponding memory budget while minimizing the training delay through employing a hybrid of recomputation and compression techniques according to the characteristic of each tensor. During the local training process, FedHybrid carries out the execution plan with a well-designed activation compression technique to effectively achieve memory reduction with minimum accuracy loss. We conduct extensive experiments to evaluate FedHybrid on both simulation and off-the-shelf mobile devices. The experiment results demonstrate that FedHybrid achieves up to a 39.1% increase in model accuracy and a 15.5X reduction in wall clock time under various memory budgets compared with the baselines.
Kahou Tam, Chunlin Tian, Li Li 0064, Haikai Zhao, Cheng-Zhong Xu 0001
SenSys4