Mengyang Ma

dblp:311/0831 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 MTA: A Merge-then-Adapt Framework for Personalized Large Language Models
abstract
Xiaopeng Li, Yuanjin Zheng, Wanyu Wang, Wenlin Zhang, Pengyue Jia, Yingyi Zhang, Haiying He, Mengyang Ma, Yiqi Wang, Maolin Wang, Xuetao Wei, Xiangyu Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaopeng Li 0014, Yuanjin Zheng, Wenlin Zhang 0001, Pengyue Jia, Yingyi Zhang 0001, Haiying He, Mengyang Ma, Yiqi Wang 0001, Maolin Wang 0001, Xuetao Wei, Xiangyu Zhao 0001
ACL (1)8
2026 GAAF: Fast and Scalable Graph-based Vector Similarity Search with Any-Match Label Filtering
abstract
In many practical scenarios, vector retrieval is frequently coupled with keyword constraints, particularly under Any-Match semantics. Filtered Approximate Nearest Neighbor Search (Filtered ANNS) has emerged as a widely adopted solution. Within this domain, state-of-the-art methods often utilize graph-based indices that enforce constraints via runtime filtering on a monolithic graph. However, real-world label skew degrades this monolithic design: frequent labels waste computation on largely valid neighborhoods, while rare labels suffer from graph sparsity in locating limited candidates. To address this, we propose GAAF, a frequency-aware Graph Ensemble framework that decouples the handling of high- and low-frequency labels. GAAF partitions the dataset into specialized graphs: utilizing dedicated indexes for high-frequency labels to eliminate redundant comparisons, while consolidating the rest of the labels into shared graphs to restore connectivity. Leveraging the fine-grained control afforded by this ensemble, we introduce NUMA-aware data placement to minimize remote access, and Adaptive Inter-graph Pruning to bypass redundant traversals. Experiments on diverse datasets demonstrate that GAAF significantly outperforms state-of-the-art baselines. © 2026 Copyright held by the owner/author(s).
Mengyang Ma, Xizhe Yin, Junqiao Qiu
ICS1
2026 BlossomRec: Block-level Fused Sparse Attention Mechanism for Sequential Recommendations
abstract
Transformer structures have been widely used in sequential recommender systems (SRS). However, as user interaction histories increase, computational time and memory requirements also grow. This is mainly caused by the standard attention mechanism. Although there exist many methods employing efficient attention and SSM-based models, these approaches struggle to effectively model long sequences and may exhibit unstable performance on short sequences. To address these challenges, we design a sparse attention mechanism, BlossomRec, which models both long-term and short-term user interests through attention computation to achieve stable performance across sequences of varying lengths. Specifically, we categorize user interests in recommendation systems into long-term and short-term interests, and compute them using two distinct sparse attention patterns, with the results combined through a learnable gated output. Theoretically, it significantly reduces the number of interactions participating in attention computation. Extensive experiments on four public datasets demonstrate that BlossomRec, when integrated with state-of-the-art Transformer-based models, achieves comparable or even superior performance while significantly reducing memory usage, providing strong evidence of BlossomRec's efficiency and effectiveness. The code is available at https://github.com/Applied-Machine-Learning-Lab/WWW2026_BlossomRec.
Mengyang Ma, Xiaopeng Li 0014, Zhaocheng Du, Jingtong Gao, Pengyue Jia, Yuyang Ye 0002, Yiqi Wang 0001, Yunpeng Weng, Weihong Luo, Xiao Han 0004, Xiangyu Zhao 0001
WWW1
2025 ConZone: A Zoned Flash Storage Emulator for Consumer Devices
abstract
Considering the potential benefits to lifespan and performance, zoned flash storage is expected to be incorporated into the next generation of consumer devices. However, due to the limited volatile cache and heterogeneous flash cells of consumer-grade flash storage, adopting a zone abstraction requires additional internal hardware design to maximize its benefits. To understand and efficiently improve the hardware design on consumer-grade zoned flash storage, we present ConZone—the first emulator tailored to the characteristics of consumer-grade zoned flash storage. Users can explore the internal architecture and management strategies of consumer-grade zoned flash storage and integrate the optimization with software. We validate the accuracy of ConZone by realizing a hardware architecture for consumer-grade zoned flash storage and comparing it with the state-of-the-art. We also make a case study for read performance research with ConZone to explore the design of mapping mechanisms and cache management strategies.
Dingcui Yu, Yumiao Zhao, Wentong Li 0002, Ziang Huang, Zonghuan Yan, Mengyang Ma, Liang Shi 0001
DATE7
2024 CacheTrimmer: Adaptive Cache File Trimming for Optimized Performance and Lifetime on Mobile Devices
abstract
Mobile devices always cache numerous files during application runtime, which can be trimmed to improve the user experience. However, existing cache file trimming methods are unaware of the cleaning cost within the file system and storage devices, which degrades the system performance and storage lifetime, resulting in low benefits of trimming cache files. Motivated by this, an adaptive cache file trimming (CacheTrimmer) scheme is proposed to trim cache files for performance and lifetime improvement. The basic idea is to determine the trimming timing based on the cleaning cost of the file system and storage device, maximizing the benefit of trimming cache files. Specifically, CacheTrimmer includes two components: First, a cleaning cost-aware trimming method is proposed to trim cache files by recording the index information of cache files in a list and determining the timing and size of file trimming. Second, to avoid trimming-induced intra-segment fragmentation and improve trimming efficiency, a log-structured cache scheme is further proposed to maintain the cache files in separate segments. We prototype CacheTrimmer with a real mobile platform. Experimental results under real workloads show that CacheTrimmer achieves encourage performance and lifetime improvement compared to the state-of-the-art.
Yunpeng Song, Wentong Li 0002, Yiyang Huang 0001, Dingcui Yu, Mengyang Ma, Liang Shi 0001
ICCD6
2024 Improving F2FS fsync() Latency Through Parallelizing Dnode and Data Page Writeback
abstract
F2FS improves performance and longevity through out-place updates, and is now wildly used in the real world. However, additional overhead is introduced to support such design, which leads to a significant performance bottleneck when performing fsync(). Through a series of experimental observations, the paper reveals the impact of dnode page writeback on throughput and fsync() latency. The serial flushing of data pages and dnode pages in the current fsync() design limits the potential for parallel write back and fails to fully utilize the parallelism of flash devices. We then deeply look into the current fsync() design and find the main difficulty of paralleling fsync() is the dependency between dnode pages and data pages. Based on these findings, we propose a new dual-thread design that significantly reduces the total latency of fsync() and improves throughput by pre-allocating data pages. Then, we give two optimizations to reduce overhead. A red-black tree is introduced to cache old block addresses for better node management performance. A linked list is introduced to avoid contention of the page cache for better page performance. We implemented our method in Linux Kernel, and the experimental results show that our dual-thread design can decrease fsync() latency and increase write throughput in different situations.
Mengyang Ma, Yumiao Zhao, Yunpeng Song, Shouzhen Gu
NAS1
2021 BuYang-HuanWu-Tang Alleviates Rheumatoid Arthritis' Hypoxia via BNIP3 and PI3K/ATK
abstract
In the development of rheumatoid arthritis (RA), hypoxia occurs in the process of synovial proliferation, neovascularization, and leukocyte extravasation. BuYang-HuanWu-Tang (BYHWT), a Chinese medicine formula, can alleviate hypoxia in RA patients with the syndrome of qi stagnation and blood stasis. However, the BYHWT mechanism network against hypoxia is not clear. In this study, bioinformatics analysis was deployed via bio-active compounds, target protein, protein interactions, and RA microarray data for mechanism exploration and validation. As a result, the BYHWT may alleviate hypoxia via down-regulating BNIP3, MAPK8, and the PI3K/AKT pathway. Literature review and microarray data can validate the results. This study may provide insights of BYHWT against hypoxia in RA for alternative therapy.
Junping Zhan, Mengyang Ma, Huilian Wang, Xiyun Miao, Qingliang Meng, Guang Zheng
BIBM2