Huidong Ma

dblp:16/1179 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
13since 2021 · last 2026
0009-0008-0882-8212ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Efficient Learned Data Compression via Dual-Stream Feature Decoupling
abstract
Huidong Ma, Xinyan Shi, Sun Hui, Xiaofei Yue, Xiaoguang Liu, Gang Wang, Wentong Cai. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Huidong Ma, Xinyan Shi, Hui Sun 0002, Xiaofei Yue, Xiaoguang Liu 0001, Gang Wang 0001, Wentong Cai 0001
ACL (1)1
2025 Genomics Data Lossless Compression with (S, K)-Mer Encoding and Deep Neural Networks
abstract
Learning-based compression shows competitive compression ratios for genomics data. It often includes three types of compressors: static, adaptive and semi-adaptive. However, these existing compressors suffer from inferior compression ratios or throughput, and adaptive compressors also faces model cold-start problems. To address these issues, we propose DeepGeCo, a novel genomics data lossless adaptive compression framework with (s,k)-mer encoding and deep neural networks, involving three compression modes (MINI for static, PLUS for adaptive, ULTRA for semi-adaptive) for flexible requirements of compression ratios or throughput. In DeepGeCo, (1) we develop BiGRU and Transformer as the backbone to build Warm-Start and Supporter models in terms of cold-start problems. (2) We introduce (s,k)-mer encoding to pre-process genomics data before feeding it into the DNN model for improve model throughput, and we propose a new metric - Ranking of Throughput and Compression Ratio (RTCR) for effective encoding parameters selection. (3) We design a threshold controller and a probabilistic mixer within the backbone to balance compression ratios and model throughput. Experiments on 10 real-world datasets show that DeepGeCo's three compression modes improve up to a 22.949X average throughput and up to a 31.095% average compression ratio improvement while occupying low CPU or GPU memory.
Hui Sun 0002, Liping Yi, Huidong Ma, Yongxia Sun, Yingfeng Zheng, Wenwen Cui, Meng Yan 0008, Gang Wang 0001, Xiaoguang Liu 0001
AAAI3
2025 Multi-source Data Lossless Compression via Parallel Expansion Mapping and xLSTM
abstract
Explosive growth of multi-source data (MSD) poses challenges in data transmitting and storing. Neural Network (NN)-based lossless compressors are an important type of compression approaches to alleviate these problems. However, existing NN-based lossless compressors suffer from poor compression ratio and high time cost at the same time. To address these issues, we propose a novel MSD Lossless Compressor (MSDLC) with two compression stages: 1) We propose a Parallel Expansion Mapper (PEM) to map redundant pieces in MSD into unused alphabet values, which not only compresses MSD but also saves time for the next stage’s NN-based lossless compression. 2) With the mapped MSD as input, we design a NN-based lossless compressor to further improve compression ratio, where we introduce the state-of-the-art xLSTM model and design a Deep Spatial Gating Module (DSGM) as the backbone of NN. We compare MSDLC with 11 baselines on 6 real-world datasets and the results validate that MSDLC obtains the best average compression ratio and time cost. Compared with baselines, compression ratios are improved by 1.103%~113.897%, and the time costs are improved by 41.367%~73.891%. The codes can be available at https://github.com/mhuidong/MSDLC.
Huidong Ma, Hui Sun 0002, Liping Yi, Xiaoguang Liu 0001, Gang Wang 0001
ICASSP1
2025 Adaptive Lossless Compression for Genomics Data by Multiple (s, k)-mer Encoding and XLSTM
abstract
Learning-based lossless compressors have been validated to have competitive advantages in genomics data (GD) compression. However, learning-based GD-dedicated compressors typically need to be pre-trained on multi-source data and then are directly used to compress another target data, we denote them as static compressors, and they often face two challenges: limited compression ratios and bad-performed generalization due to data distribution variations. To solve these problems, we propose AGDLC, a novel Adaptive Genomics Data Lossless Compressor. It includes two critical designs: 1) We design a multiple (s, k)-mer mixer for extracting GD redundancy from multiple dimensions to improve compression ratios. 2) We introduce a recently popular XLSTM model as the backbone, which adaptively compresses GD while updating parameters, without pre-training, improving compression ratios and compression generalization at the same time. We compare AGDLC with 13 baselines on 7 real-world datasets, and the experimental results demonstrate that it achieves the best compression ratio with an average improvement of 2.162%-69.436%. The codes can be found at https://github.com/dingyanfeng/AGDLC.
Hui Sun 0002, Yanfeng Ding, Liping Yi, Huidong Ma, Haonan Xie, Gang Wang 0001, Xiaoguang Liu 0001
ICASSP4
2025 PMKLC: Parallel Multi-Knowledge Learning-based Lossless Compression for Large-Scale Genomics Database
abstract
Learning-based lossless compressors play a crucial role in large-scale genomic database backup, storage, transmission, and management. However, their 1) inadequate compression ratio, 2) low compression & decompression throughput, and 3) poor compression robustness limit their widespread adoption and application in both industry and academia. To solve those challenges, we propose a novel Parallel Multi-Knowledge Learning-based Compressor (PMKLC) with four crucial designs: 1) We propose an automated multi-knowledge learning-based compression framework as compressors' backbone to enhance compression ratio and robustness; 2) we design a GPU-accelerated (s,k)-mer encoder to optimize compression throughput and computing resource usage; 3) we introduce data block partitioning and Step-wise Model Passing (SMP) mechanisms for parallel acceleration; 4) We design two compression modes PMKLC-S and PMKLC-M to meet the complex application scenarios, where the former runs on a resource-constrained single GPU and the latter is multi-GPU accelerated. We benchmark PMKLC-S/M and 14 baselines (7 traditional and 7 leaning-based) on 15 real-world datasets with different species and data sizes. Compared to baselines on the testing datasets, PMKLC-S/M achieve the average compression ratio improvement up to 73.609% and 73.480%, the average throughput improvement up to 3.036X and 10.710X, respectively. Besides, PMKLC-S/M also achieve the best robustness and competitive memory cost, indicating its greater stability against datasets with different probability distribution perturbations, and its strong ability to run on memory-constrained devices. Overall, PMKLC is a balanced compression solution that optimizes compression ratio, throughput, robustness, and resource consumption. PMKLC and linkages of datasets are available at https://github.com/dingyanfeng/PMKLC.
Hui Sun 0002, Yanfeng Ding, Liping Yi, Huidong Ma, Gang Wang 0001, Xiaoguang Liu 0001, Wentong Cai 0001
KDD (2)4
2025 MSDZip: Universal Lossless Compression for Multi-source Data via Stepwise-parallel and Learning-based Prediction
abstract
With the rapid development of the Internet, the huge amount of Multi-Source Data (MSD) brings challenges in data sharing and storing. Lossless data compression is the major way to solve those problems. Nowadays, neural-network technologies bring significant advantage in data modeling, making learning-based lossless compressors (LLCs) for multi-source data have emerged continuously. Compared with traditional compressors, the LLCs are more useful to catch complex redundancy patterns in MSD, and thus have great potential in enhancing compression ratio. However, existing LLCs still suffer from unsatisfactory compression ratios and lower throughput. To solve those problems, we propose a novel universal MSD lossless compressor called MSDZip via Stepwise-parallel and learning-based prediction technologies, it introduces two major designs: 1) We propose a Local-Global-Deep Mixing block in the learning-based prediction module to establish dependencies for MSD symbols, where designed Deep Mixing block solves the problem of unstable weights in the perceptual layers caused by cold-start problem to enhance the compression ratio significantly. 2) We design a Stepwise-parallel multi-GPU-accelerated compression strategy to address the compression speed and graphics memory constraints of single GPU in the face of large-scale data. The Stepwise-parallel module passes the source MSD to learning-based prediction model through the data chunking strategy, where the model of the previous chunk is used to guide the compression of the next chunk in parallel. We compare MSDZip with 5 classical learning-based and 6 traditional compressors on 12 well-studied real-world datasets. The experimental results demonstrate that MSDZip optimizes 3.418%-69.874% in terms of compression ratio and 31.171%-495.649% in terms of throughput compared to advanced LLCs. The source code of MSDZip and the linkages of the experimental datasets are available at https://github.com/mhuidong/MSDZip.
Huidong Ma, Hui Sun 0002, Liping Yi, Yanfeng Ding, Xiaoguang Liu 0001, Gang Wang 0001
WWW1
2025 A survey and benchmark evaluation for neural-network-based lossless universal compressors toward multi-source data
abstract
Abstract As various types of data grow explosively, large-scale data storage, backup, and transmission become challenging, which motivates many researchers to propose efficient universal compression algorithms for multi-source data. In recent years, due to the emergence of hardware acceleration devices such as GPUs, TPUs, DPUs, and FPGAs, the performance bottleneck of neural networks (NN) has been overcome, making NN-based compression algorithms increasingly practical and popular. However, the research survey for the NN-based universal lossless compressors has not been conducted yet, and there is also a lack of unified evaluation metrics. To address the above problems, in this paper, we present a holistic survey as well as benchmark evaluations. Specifically, i) we thoroughly investigate NN-based lossless universal compression algorithms toward multi-source data and classify them into 3 types: static pre-training, adaptive, and semi-adaptive. ii) We unify 19 evaluation metrics to comprehensively assess the compression effect, resource consumption, and model performance of compressors. iii) We conduct experiments more than 4600 CPU/GPU hours to evaluate 17 state-of-the-art compressors on 28 real-world datasets across data types of text, images, videos, audio, etc. iv) We also summarize the strengths and drawbacks of NN-based lossless data compressors and discuss promising research directions. We summarize the results as the NN-based Lossless Compressors Benchmark (NNLCB, See fahaihi.github.io/NNLCB website), which will be updated and maintained continuously in the future.
Hui Sun 0002, Huidong Ma, Haonan Xie, Yongxia Sun, Liping Yi, Meng Yan 0008, Xiaoguang Liu 0001, Gang Wang 0001
Frontiers Comput. Sci.2
2024 LRCB: A Comprehensive Benchmark Evaluation of Reference-free Lossless Compression Tools for Genomics Sequencing Long Reads Data
abstract
The advancement of long reads sequencing technologies has led to a significant increase in biological sequencing big data. Although several reference-free compressors are available for saving long reads data storage space, choosing the suitable one is challenging due to the shortage of thorough and systematic evaluations of their lossless compression effectiveness, both dedicated and general-purpose. In this study, we performed benchmark examinations on 30 compressors, including 11 specialized for long reads and 19 general-purpose ones, using 31 real-world datasets with differing sequencing platforms, species, and lengths. Each lossless compressor was evaluated on 13 performance measures, including compression strength, compression robustness, as well as time and peak memory required for compression and decompression. Additionally, for future long reads data compressors, we outlined investigation directions with consideration for privacy-sensitive sequences data security, hardware parallel acceleration, parameter tuning framework, and system hardware-algorithm integration design. We summarized the results as the Long Reads Compression Benchmark, available at https://github.com/fahaihi/LRCB.
Hui Sun 0002, Huidong Ma, Yingfeng Zheng, Haonan Xie, Meng Yan 0008, Xiaoguang Liu 0001, Gang Wang 0001
DCC2
2024 PQSDC: a parallel lossless compressor for quality scores data via sequences partition and run-length prediction mapping
abstract
MOTIVATION: The quality scores data (QSD) account for 70% in compressed FastQ files obtained from the short and long reads sequencing technologies. Designing effective compressors for QSD that counterbalance compression ratio, time cost, and memory consumption is essential in scenarios such as large-scale genomics data sharing and long-term data backup. This study presents a novel parallel lossless QSD-dedicated compression algorithm named PQSDC, which fulfills the above requirements well. PQSDC is based on two core components: a parallel sequences-partition model designed to reduce peak memory consumption and time cost during compression and decompression processes, as well as a parallel four-level run-length prediction mapping model to enhance compression ratio. Besides, the PQSDC algorithm is also designed to be highly concurrent using multicore CPU clusters. RESULTS: We evaluate PQSDC and four state-of-the-art compression algorithms on 27 real-world datasets, including 61.857 billion QSD characters and 632.908 million QSD sequences. (1) For short reads, compared to baselines, the maximum improvement of PQSDC reaches 7.06% in average compression ratio, and 8.01% in weighted average compression ratio. During compression and decompression, the maximum total time savings of PQSDC are 79.96% and 84.56%, respectively; the maximum average memory savings are 68.34% and 77.63%, respectively. (2) For long reads, the maximum improvement of PQSDC reaches 12.51% and 13.42% in average and weighted average compression ratio, respectively. The maximum total time savings during compression and decompression are 53.51% and 72.53%, respectively; the maximum average memory savings are 19.44% and 17.42%, respectively. (3) Furthermore, PQSDC ranks second in compression robustness among the tested algorithms, indicating that it is less affected by the probability distribution of the QSD collections. Overall, our work provides a promising solution for QSD parallel compression, which balances storage cost, time consumption, and memory occupation primely. AVAILABILITY AND IMPLEMENTATION: The proposed PQSDC compressor can be downloaded from https://github.com/fahaihi/PQSDC.
Hui Sun 0002, Yingfeng Zheng, Haonan Xie, Huidong Ma, Meng Yan 0008, Xiaoguang Liu 0001, Gang Wang 0001
Bioinform.4
2023 SR2C: A Structurally Redundant Short Reads Collapser for Optimizing DNA Data Compression
abstract
The current redundant sequence deduplication algorithms cannot remove structural repetitive DNA short reads such as mirror, reverse, paired, and complementary palindromes in high-throughput genomics sequencing data. Moreover, these methods also cannot construct indexes to recover the original sequences, thus failing to meet the requirements of lossless compression for downstream applications. To address these problems, we propose a data structure called Cycle-Hash-Linkage (CHL) and present a CPU parallelism optimization algorithm named SR2C (Structurally Redundant Short Reads Collapser) based on CHL to improve the compression ratio of DNA sequencing data. Experimental results on actual data from the NCBI database demonstrate that SR2C achieves an average residual sequence percentage improvement of 2.556% compared to the state-of-the-art redundant sequence deduplication algorithm, Minirmd. Furthermore, SR2C cascaded optimization improves the average compression ratios of compression algorithms Pigz, PBzip2, XZ, and 7Z by 92.345%, 78.999%, 10.132%, and 7.434%, respectively. By leveraging multi-core CPU parallel computation, SR2C effectively reduces time consumption, which achieves 2-5X deduplication and recovers acceleration.The same name Linux toolkit is freely available at https://github.com/fahaihi/SR2C.
Hui Sun 0002, Huidong Ma, Yingfeng Zheng, Haonan Xie, Xiaofei Wang 0001, Xiaoguang Liu 0001, Gang Wang 0001
ICPADS2
2023 ricME: Long-Read Based Mobile Element Variant Detection Using Sequence Realignment and Identity Calculation
Huidong Ma, Hui Sun 0002, Haixiang Lin
ISBRA1
2023 cnnLSV: detecting structural variants by encoding long-read alignment information and convolutional neural network
abstract
BACKGROUND: Genomic structural variant detection is a significant and challenging issue in genome analysis. The existing long-read based structural variant detection methods still have space for improvement in detecting multi-type structural variants. RESULTS: In this paper, we propose a method called cnnLSV to obtain detection results with higher quality by eliminating false positives in the detection results merged from the callsets of existing methods. We design an encoding strategy for four types of structural variants to represent long-read alignment information around structural variants into images, input the images into a constructed convolutional neural network to train a filter model, and load the trained model to remove the false positives to improve the detection performance. We also eliminate mislabeled training samples in the training model phase by using principal component analysis algorithm and unsupervised clustering algorithm k-means. Experimental results on both simulated and real datasets show that our proposed method outperforms existing methods overall in detecting insertions, deletions, inversions, and duplications. The program of cnnLSV is available at https://github.com/mhuidong/cnnLSV . CONCLUSIONS: The proposed cnnLSV can detect structural variants by using long-read alignment information and convolutional neural network to achieve overall higher performance, and effectively eliminate incorrectly labeled samples by using the principal component analysis and k-means algorithms in training model stage.
Huidong Ma, Haofa He, Feng Yang 0014
BMC Bioinform.1
2023 PMFFRC: a large-scale genomic short reads compression optimizer via memory modeling and redundant clustering
abstract
BACKGROUND: Genomic sequencing reads compressors are essential for balancing high-throughput sequencing short reads generation speed, large-scale genomic data sharing, and infrastructure storage expenditure. However, most existing short reads compressors rarely utilize big-memory systems and duplicative information between diverse sequencing files to achieve a higher compression ratio for conserving reads data storage space. RESULTS: We employ compression ratio as the optimization objective and propose a large-scale genomic sequencing short reads data compression optimizer, named PMFFRC, through novelty memory modeling and redundant reads clustering technologies. By cascading PMFFRC, in 982 GB fastq format sequencing data, with 274 GB and 3.3 billion short reads, the state-of-the-art and reference-free compressors HARC, SPRING, Mstcom, and FastqCLS achieve 77.89%, 77.56%, 73.51%, and 29.36% average maximum compression ratio gains, respectively. PMFFRC saves 39.41%, 41.62%, 40.99%, and 20.19% of storage space sizes compared with the four unoptimized compressors. CONCLUSIONS: PMFFRC rational usage big-memory of compression server, effectively saving the sequencing reads data storage space sizes, which relieves the basic storage facilities costs and community sharing transmitting overhead. Our work furnishes a novel solution for improving sequencing reads compression and saving storage space. The proposed PMFFRC algorithm is packaged in a same-name Linux toolkit, available un-limited at https://github.com/fahaihi/PMFFRC .
Hui Sun 0002, Yingfeng Zheng, Haonan Xie, Huidong Ma, Xiaoguang Liu 0001, Gang Wang 0001
BMC Bioinform.4