VLDB 2026 Research / reviewers in the wild / expert
Yanfeng Ding
dblp:402/7265
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Storage systems · 53% GPUs and heterogeneous computing · 32% High-performance computing · 16% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › data compression
genomic data compression |
0.9 | 1 | 2025 | PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database · KDD (2) 2025 |
Bioinformatics and computational biology › genomics
genomic data management |
0.3 | 1 | 2025 | PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database · KDD (2) 2025 |
GPUs and heterogeneous computing › GPU computing › GPU algorithms
GPU compression |
0.3 | 1 | 2025 | PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database · KDD (2) 2025 |
GPUs and heterogeneous computing
multi-GPU computing |
0.3 | 1 | 2025 | MSDZip: Universal Lossless Compression for Multi-source Data via Stepwise-parallel and Learning-based Prediction · WWW 2025 |
High-performance computing
parallel compression |
0.3 | 1 | 2025 | MSDZip: Universal Lossless Compression for Multi-source Data via Stepwise-parallel and Learning-based Prediction · WWW 2025 |
Methods — techniques the papers use, named apart from their topics
step-wise model passing · 1.7multi-knowledge learning · 1.7(s,k)-mer encoding · 1.7stepwise-parallel compression · 0.9neural network prediction · 0.9local-global-deep mixing · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive Lossless Compression for Genomics Data by Multiple (s, k)-mer Encoding and XLSTMabstractLearning-based lossless compressors have been validated to have competitive advantages in genomics data (GD) compression. However, learning-based GD-dedicated compressors typically need to be pre-trained on multi-source data and then are directly used to compress another target data, we denote them as static compressors, and they often face two challenges: limited compression ratios and bad-performed generalization due to data distribution variations. To solve these problems, we propose AGDLC, a novel Adaptive Genomics Data Lossless Compressor. It includes two critical designs: 1) We design a multiple (s, k)-mer mixer for extracting GD redundancy from multiple dimensions to improve compression ratios. 2) We introduce a recently popular XLSTM model as the backbone, which adaptively compresses GD while updating parameters, without pre-training, improving compression ratios and compression generalization at the same time. We compare AGDLC with 13 baselines on 7 real-world datasets, and the experimental results demonstrate that it achieves the best compression ratio with an average improvement of 2.162%-69.436%. The codes can be found at https://github.com/dingyanfeng/AGDLC. Hui Sun 0002, Yanfeng Ding, Liping Yi, Huidong Ma, Haonan Xie, Gang Wang 0001, Xiaoguang Liu 0001 |
ICASSP | 2 |
| 2025 | PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics DatabaseabstractLearning-based lossless compressors play a crucial role in large-scale genomic database backup, storage, transmission, and management. However, their 1) inadequate compression ratio, 2) low compression & decompression throughput, and 3) poor compression robustness limit their widespread adoption and application in both industry and academia. To solve those challenges, we propose a novel Parallel Multi-Knowledge Learning-based Compressor (PMKLC) with four crucial designs: 1) We propose an automated multi-knowledge learning-based compression framework as compressors' backbone to enhance compression ratio and robustness; 2) we design a GPU-accelerated (s,k)-mer encoder to optimize compression throughput and computing resource usage; 3) we introduce data block partitioning and Step-wise Model Passing (SMP) mechanisms for parallel acceleration; 4) We design two compression modes PMKLC-S and PMKLC-M to meet the complex application scenarios, where the former runs on a resource-constrained single GPU and the latter is multi-GPU accelerated. We benchmark PMKLC-S/M and 14 baselines (7 traditional and 7 leaning-based) on 15 real-world datasets with different species and data sizes. Compared to baselines on the testing datasets, PMKLC-S/M achieve the average compression ratio improvement up to 73.609% and 73.480%, the average throughput improvement up to 3.036X and 10.710X, respectively. Besides, PMKLC-S/M also achieve the best robustness and competitive memory cost, indicating its greater stability against datasets with different probability distribution perturbations, and its strong ability to run on memory-constrained devices. Overall, PMKLC is a balanced compression solution that optimizes compression ratio, throughput, robustness, and resource consumption. PMKLC and linkages of datasets are available at https://github.com/dingyanfeng/PMKLC. Hui Sun 0002, Yanfeng Ding, Liping Yi, Huidong Ma, Gang Wang 0001, Xiaoguang Liu 0001, Wentong Cai 0001 |
KDD (2) | 2 |
| 2025 | MSDZip: Universal Lossless Compression for Multi-source Data via Stepwise-parallel and Learning-based PredictionabstractWith the rapid development of the Internet, the huge amount of Multi-Source Data (MSD) brings challenges in data sharing and storing. Lossless data compression is the major way to solve those problems. Nowadays, neural-network technologies bring significant advantage in data modeling, making learning-based lossless compressors (LLCs) for multi-source data have emerged continuously. Compared with traditional compressors, the LLCs are more useful to catch complex redundancy patterns in MSD, and thus have great potential in enhancing compression ratio. However, existing LLCs still suffer from unsatisfactory compression ratios and lower throughput. To solve those problems, we propose a novel universal MSD lossless compressor called MSDZip via Stepwise-parallel and learning-based prediction technologies, it introduces two major designs: 1) We propose a Local-Global-Deep Mixing block in the learning-based prediction module to establish dependencies for MSD symbols, where designed Deep Mixing block solves the problem of unstable weights in the perceptual layers caused by cold-start problem to enhance the compression ratio significantly. 2) We design a Stepwise-parallel multi-GPU-accelerated compression strategy to address the compression speed and graphics memory constraints of single GPU in the face of large-scale data. The Stepwise-parallel module passes the source MSD to learning-based prediction model through the data chunking strategy, where the model of the previous chunk is used to guide the compression of the next chunk in parallel. We compare MSDZip with 5 classical learning-based and 6 traditional compressors on 12 well-studied real-world datasets. The experimental results demonstrate that MSDZip optimizes 3.418%-69.874% in terms of compression ratio and 31.171%-495.649% in terms of throughput compared to advanced LLCs. The source code of MSDZip and the linkages of the experimental datasets are available at https://github.com/mhuidong/MSDZip. Huidong Ma, Hui Sun 0002, Liping Yi, Yanfeng Ding, Xiaoguang Liu 0001, Gang Wang 0001 |
WWW | 4 |