Weilin Zhu

dblp:249/7966 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 LPAQMP: Multilayer Parallel Design for LPAQ Compression
abstract
The context mixing compression algorithm (CM) offers an extremely high compression ratio, effectively reducing data storage costs. However, its throughput on CPU platforms is extremely low, making it difficult to fully exploit the algorithm's inherent parallel potential: neither sustained byte-level SIMD execution is attainable, nor does the throughput scale close to linear in the number of threads on multicore CPUs. To address these issues, we propose a collaborative byte-level and task-level parallel design scheme, LPAQMP. At the byte level, LPAQMP completely eliminates residual dependencies between bits by restructuring the data path; at the task level, LPAQMP introduces a global rollback hash table (GRHT), significantly improving cache efficiency and achieving approximately linear throughput scaling as threads increase. Experiments show that while maintaining a high compression ratio, LPAQMP achieves a throughput of$152.67 \text{MB} / \mathrm{s}$, a$12.3 \times$improvement over LPAQ-CPU and$1.2 \times$higher than SOTA.
Panyue Wei, Wei Tong 0001, Weilin Zhu, Yifei Qu
DCC3
2024 LpaqHP: A High-Performance FPGA Accelerator for LPAQ Compression
abstract
LPAQ is a powerful context-based lossless compression algorithm ranking top on compression ratio on many benchmarks. However, its application is limited due to its high computational complexity, extremely slow compression speed, and large memory usage. In this paper, we introduce LpaqHP, an FPGA-based design, to accelerate LPAQ. We speed up LPAQ by eliminating the bit-level dependency within a byte in the three main components of LPAQ. In other words, LpaqHP can compress all eight bits within a byte in parallel. In the meanwhile, we managed to maintain the compression ratio in LpaqHP by implementing two dedicated schemes to compensate for the compression ratio. Experiments show that LpaqHP achieves a throughput of 67.96MB/s on Xilinx Virtex UltraScale plus VCU118 card, 234 × faster than executing on Intel Xeon E5-2650 at 2.2GHz and 5.97 × faster than the state-of-the-art work pLPAQ.
Weilin Zhu, Wei Tong 0001, Hujun Ge, Zuoxian Zhang, Mengran Zhang, Wen Zhou 0030
ICPP1
2023 Turn Waste Into Wealth: Alleviating Read/Write Interference in ZNS SSDs
abstract
The emerging NVMe Zoned Namespace (ZNS) solid-state drive (SSD) is built on high-density NAND flash memories. As write latency is much longer than read latency in NAND flash memories, the read performance of ZNS SSD is subject to the chip-blocking writes. However, current works on alleviating read/write interference are usually designed for traditional SSDs, which cannot be applied to the ZNS SSD due to its unique constraints. In this work, we find that many zones marked as full actually have some empty and wasted spaces, to which we refer as idle space. We back up popular read data with idle space. To minimize the overhead, we leverage device-side information to ensure that only one backup is needed. Secondly, we propose an I/O scheduling method through request splitting to ensure that the sole backup is not blocked by any writes and can always serve blocked reads. Experiments show that when compared to the current ZNS SSDs, our work significantly improves average read response time and read tail latency of 99thand 99.9thpercentile by up to 44.28%, 50.09%, and 46.68%. Moreover, our work prevails over the read-prioritizing scheme on read performance and write tail latency.
Weilin Zhu, Wei Tong 0001
ICCD1
2023 A Low-Latency and High-Endurance MLC STT-MRAM-Based Cache System
abstract
Spin-transfer torque magnetic random access memory (STT-MRAM) is a promising cache memory candidate due to its high density, low leakage power, and nonvolatility. Multilevel cell (MLC) STT-MRAM can further increase density by storing 2 bits in one cell’s hard and soft domain, respectively. However, MLC STT-MRAM suffers two-step write, leading to high write energy, long latency, and severe lifetime degradation. Current encoding techniques propose to encode the new data to reduce the two-step data writes. However, they have two weaknesses: 1) high area overhead, e.g., recent work TSE (Hsieh et al., 2020) needs extra 37.5% MLCs and 2) prolong the write latency due to an extra read. Therefore, we propose enhanced one-step write (EOSwrite) to write data in one step. EOSwrite includes line bypassing and four intraline encoding techniques. Line bypassing schemes can bypass the writes to zero or clean lines, leading to low write/read latency. As for the intraline techniques, we propose four write modes. They utilize the data patterns and the clean data in cache lines to write data in one step, therefore reducing the data write latency. The key idea of one-step write is to write as much data as possible in the soft domain of MLC STT-MRAM. EOSwrite can greatly relieve the weaknesses of the current encoding schemes. Evaluation results show that EOSwrite can improve the lifetime of MLC STT-MRAM by 56.96%, reduce dynamic energy by 33.95%, reduce access latency by 36.95%, and improve system performance of MLC STT-MRAM by 4.30%, respectively. While the area overhead of EOSwrite is only 7.27%.
Wei Zhao 0034, Jie Xu 0013, Xueliang Wei, Bing Wu 0001, Chengning Wang, Weilin Zhu, Wei Tong 0001, Dan Feng 0001, Jingning Liu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 HDQGF: Heterogeneous Data Quality Guarantee Framework Based on Deep Learning
abstract
Although user-generated data on the Internet contains rich information, many approaches cannot effectively guarantee data quality for data analysis from raw data. In recent years, many researches on data quality guarantee using machine learning have shown that enhancing data quality is conducive to improve the accuracy of analysis results. But the existing approaches only consider the single dimension and neglect the fusion of heterogeneous data, such as images or social graphs. To consider this element and address the above issue, we leverage the deep learning technique to guarantee the data quality by using the heterogeneous data. Our framework is named HDQGF, which is an end-to-end approach using a combination of multiple networks to fuse heterogeneous information. In order to verify the effectiveness of the model, we designed related experiments on three real datasets. According to the experimental results, our model HDQGF can enhance the performance by improving the data quality.
Zongze Jin, Weilin Zhu, Lei Chi, Weiping Wang 0005
CSCWD3
2020 DROAllocator: A Dynamic Resource-Aware Operator Allocation Framework in Distributed Streaming Processing
Zongze Jin, Weimin Mu, Weilin Zhu
NPC4
2019 BGElasor: Elastic-Scaling Framework for Distributed Streaming Processing with Deep Neural Network
Weimin Mu, Zongze Jin, Weilin Zhu, Weiping Wang 0005
NPC4