EDBT 2026 Demo / reviewers in the wild / expert
Yili Ma
dblp:280/9126
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0000-9405-4584ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs
Jinwu Yang, Jiaan Wu, Xinyang Ma, Hairui Zhao 0002, Yida Gu, Yuanhong Huang, Wenjing Huang 0002, Yili Ma, Zhongzhe Hu, Shaoteng Liu, Jiaxun Lu, Guangming Tan, Dingwen Tao |
ISCA | 12 |
| 2026 | HFP-SAM: Hierarchical Frequency Prompted SAM for Efficient Marine Animal SegmentationabstractMarine Animal Segmentation (MAS) aims at identifying and segmenting marine animals from complex marine environments. Most of previous deep learning-based MAS methods struggle with the long-distance modeling issue. Recently, Segment Anything Model (SAM) has gained popularity in general image segmentation. However, it lacks of perceiving fine-grained details and frequency information. To this end, we propose a novel learning framework, named Hierarchical Frequency Prompted SAM (HFP-SAM) for high-performance MAS. First, we design a Frequency Guided Adapter (FGA) to efficiently inject marine scene information into the frozen SAM backbone through frequency domain prior masks. Additionally, we introduce a Frequency-aware Point Selection (FPS) to generate highlighted regions through frequency analysis. These regions are combined with the coarse predictions of SAM to generate point prompts and integrate into SAM's decoder for fine predictions. Finally, to obtain comprehensive segmentation masks, we introduce a Full-View Mamba (FVM) to efficiently extract spatial and channel contextual information with linear computational complexity. Extensive experiments on four public datasets demonstrate the superior performance of our approach. We will make our code publicly available upon the acceptance. Tianyu Yan, Yang Liu 0066, Tongdan Tang, Yili Ma, Long Lv, Feng Tian 0001, Weibing Sun, Huchuan Lu |
IEEE Trans. Image Process. | 6 |
| 2026 | Computational Burst Buffers: Accelerating HPC I/O via In-Storage Compression OffloadingabstractBurst buffers (BBs) act as an intermediate storage layer between compute nodes and parallel file systems (PFS), effectively alleviating the I/O performance gap in high-performance computing (HPC). As scientific simulations and AI workloads generate larger checkpoints and analysis outputs, BB capacity shortages and PFS bandwidth bottlenecks are emerging, and CPU-based compression is not an effective solution due to its high overhead. We introduceComputational Burst Buffers(CBBs), a storage paradigm that embeds hardware compression engines such as application-specific integrated circuit (ASIC) inside computational storage drives (CSDs) at the BB tier. CBB transparently offloads both lossless and error-bounded lossy compression from CPUs to CSDs, thereby (i) expanding effective SSD-backed BB capacity, (ii) reducing BB–PFS traffic, and (iii) eliminating contention and energy overheads of CPU-based compression. Unlike prior CSD-based compression designs targeting databases or flash caching, CBB co-designs the burst-buffer layer and CSD hardware for HPC and quantitatively evaluates compression offload in BB–PFS hierarchies. We prototype CBB using a PCIe 5.0 CSD with an ASIC Zstd-like compressor and an FPGA prototype of an SZ entropy encoder, and evaluate CBB on a 16-node cluster. Experiments with four representative HPC applications and a large-scale workflow simulator show up to 61% lower application runtime, 8–12× higher cache hit ratios, and substantially reduced compute-node CPU utilization compared to software compression and conventional BBs. These results demonstrate that compression-aware BBs with CSDs provide a practical, scalable path to next-generation HPC storage. Xiang Chen 0028, Bing Lu 0001, Haoquan Long, Huizhang Luo, Yili Ma, Guangming Tan, Dingwen Tao, Fei Wu 0005, Tao Lu 0014 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2026 | TSUE+: An Efficient Update Framework With Swift Recycling Mechanism for Erasure-Coded Cluster File SystemsabstractCompared to replication-based storage systems, erasure-coded storage incurs significantly higher overhead during data updates. To address this issue, various parity logging methods have been proposed. Nevertheless, due to the long update path and substantial amount of random I/O involved in erasure code update processes, the resulting long latency and low through put often fail to meet the requirements of high performance applications. To address this challenge, we propose TSUE+, an efficient update framework with a swift recycling mechanism. TSUE+ divides the update process into two distinct stages: in the synchronous stage, data updates are stored in the format of replica data logs, eliminating random I/O by trading space for time; in the asynchronous stage, the recorded update logs are recycled and merged into original data and parity blocks, thereby reclaiming the storage overhead incurred in the synchronization phase. By converting random I/O operations into sequential ones based on data logs, TSUE+ effectively reduces update latency; furthermore, it significantly minimizes recycling overhead using a three-layer log structure and by leveraging the spatio-temporal locality of access patterns. We evaluated TSUE+ and other state of-the-art (SOTA) update mechanisms under diverse encoding schemes, using heterogeneous storage devices—including HDDs, SATA SSDs, NVMe SSDs, and PMEM—and multiple real-world and synthetic workloads: the MSR Cambridge trace, the Alibaba Cloud trace, the Tencent Cloud trace, and multiple synthetic worst-case workloads. Among all the platforms, TSUE+ has achieved significant performance improvements compared to other update methods, it also indicates that TSUE+ can be applied to various storage devices. Additionally, we provided percentile-based tail latency tests and update tests under the worst-case environment, which further demonstrated the ro bustness of TSUE+. Moreover, by enabling prompt log recycle and avoiding unnecessary overwrites and improving update granularity through locality-aware recycling, TSUE+ not only improves update performance but also mitigates write wear on SSD devices, thereby extending their operational lifespan. Yida Gu, Wenjing Huang 0002, Yili Ma, Dong Dai 0001, Guangming Tan, Dingwen Tao |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | SaneKV: A Swift-Adaptive and NUMA-Enhanced Persistent Key-Value StoreabstractKey-value (KV) storage is widely used in domains such as big data analytics, AI training, databases, and distributed file systems. Traditional KV systems built on DRAM or disk-based architectures struggle to meet the dual requirements of high throughput and strong data persistence demanded by modern applications. Non-Volatile Memory (NVM), with its combination of high bandwidth and persistence, offers a promising foundation for building high-performance, persistent, and largescale KV stores. Consequently, NVM-based KV storage has gained substantial research and industrial interest in recent years. However, directly porting DRAM-or disk-oriented KV designs to NVM devices often yields suboptimal performance. In Non-Uniform Memory Access (NUMA) architectures, frequent NVM persistence operations and high cross-NUMA access latency significantly limit I/O efficiency. To address these challenges, we propose SaneKV, an NVM-optimized KV store. SaneKV introduces a metadata asynchronous persistence mechanism that reduces I/O latency by aggregating metadata write-back operations, and an adaptive NVM data allocation policy to improve throughput. Experimental results show that for small-sized KV workloads, SaneKV achieves up to 75 % higher write throughput and 40 % higher read throughput than state-of-the-art NVM-based KV stores. For large-sized KV workloads, SaneKV achieves comparable peak throughput while requiring up to 40 % fewer threads to reach saturation and delivers over$2 \times$higher overall performance than prior NVM-based designs, demonstrating superior scalability and resource utilization. Shengquan Yin, Yili Ma, Dingwen Tao, Guangming Tan |
ICPADS | 2 |
| 2024 | GroupMorph: Medical Image Registration via Grouping Network With Contextual FusionabstractPyramid-based deformation decomposition is a promising registration framework, which gradually decomposes the deformation field into multi-resolution subfields for precise registration. However, most pyramid-based methods directly produce one subfield per resolution level, which does not fully depict the spatial deformation. In this paper, we propose a novel registration model, called GroupMorph. Different from typical pyramid-based methods, we adopt the grouping-combination strategy to predict deformation field at each resolution. Specifically, we perform group-wise correlation calculation to measure the similarities of grouped features. After that, n groups of deformation subfields with different receptive fields are predicted in parallel. By composing these subfields, a deformation field with multi-receptive field ranges is formed, which can effectively identify both large and small deformations. Meanwhile, a contextual fusion module is designed to fuse the contextual features and provide the inter-group information for the field estimator of the next level. By leveraging the inter-group correspondence, the synergy among deformation subfields is enhanced. Extensive experiments on four public datasets demonstrate the effectiveness of GroupMorph. Code is available at https://github.com/TVayne/GroupMorph. Zuopeng Tan, Lihe Zhang, Yanan Lv, Yili Ma, Huchuan Lu |
IEEE Trans. Medical Imaging | 4 |