EDBT 2026 Demo / reviewers in the wild / expert
Songzhu Mei
dblp:11/10487
· DBLP profile ↗
22ranked-venue papers
1as first author
14since 2021 · last 2026
0000-0002-4926-5953ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 6 since 2021Security and privacy · 2Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HierMine: Accelerating Graph Pattern Mining via Hierarchical SamplingabstractCurrent Approximate Graph Pattern Mining (AGPM) systems rely on uniform and static sampler allocation strategies, which result in significant computational redundancy because most regions of the sampling space contribute little to the final estimate. However, existing systems overlook this critical characteristic and fail to account for the hierarchical memory architecture of modern servers, leading to slow convergence and substantial memory access overhead. To address these challenges, we propose HierMine, a novel AGPM system designed to overcome the aforementioned drawbacks. The core contributions of HierMine include: (i) a hierarchical sampling strategy that minimizes redundant computation and accelerates convergence; (ii) a dynamic sampling adjustment mechanism that partitions the sampling space online and reallocates samplers adaptively to enhance the hierarchical strategy; (iii) an online grouping-based convergence detection technique that enables the fine-grained dynamic sampling adjustment mechanism; and (iv) a hierarchical data layout optimized for memory access efficiency. Extensive experiments show that HierMine delivers an average speedup of up to 28.9× over state-of-the-art AGPM systems such as ScaleGPM, while maintaining strong theoretical guarantees on estimation quality. It also consistently outperforms exact graph pattern mining systems, highlighting the practical benefits of our approximate approach. Xinbiao Gan, Songzhu Mei, Zhengbin Pang, Hongxu Jin |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | LLM-based Rumor Detection via Influence Guided Sample Selection and Game-based Perspective AnalysisabstractZhiliang Tian, Jingyuan Huang, Zejiang He, Zhen Huang, Menglong Lu, Linbo Qiao, Songzhu Mei, Yijie Wang, Dongsheng Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhiliang Tian, Zejiang He, Zhen Huang 0006, Menglong Lu, Linbo Qiao, Songzhu Mei, Yijie Wang 0001 |
ACL (1) | 7 |
| 2025 | Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inferenceabstractWith the continuous advancement in the performance of large language models (LLMs), their demand for computational resources and memory has significantly increased, which poses major challenges for efficient inference on consumer-grade devices and legacy servers.These devices typically feature relatively weaker GPUs and stronger CPUs.Although techniques such as parameter offloading and partial offloading can alleviate GPU memory pressure to some extent, their effectiveness is limited due to communication latency and suboptimal hardware resource utilization.To address this issue, we propose Dovetail 1 , a lossless inference acceleration method that leverages the complementary characteristics of heterogeneous devices and the advantages of speculative decoding.Dovetail deploys a draft model on the GPU to perform preliminary predictions, while a target model running on the CPU validates these outputs.By reducing the granularity of data transfer, Dovetail significantly minimizes communication overhead.To further improve efficiency, we optimize the draft model specifically for heterogeneous hardware environments by reducing the number of draft tokens to lower parallel verification latency, increasing model depth to enhance predictive capabilities, and introducing a Dynamic Gating Fusion (DGF) mechanism to improve the integration of feature and embedding information.We conduct comprehensive evaluations of Dovetail across various consumergrade GPUs, covering multiple tasks and mainstream models.Experimental results on 13B models demonstrate that Dovetail achieves inference speedups ranging from 1.79x to 10.1x across different devices, while maintaining consistency and stability in the distribution of generated texts.* indicates equal contribution.† indicates corresponding authors. 1 Xubaizhou, Zhiliang Tian, Songzhu Mei |
EMNLP | 6 |
| 2025 | GradPS: Resolving Futile Neurons in Parameter Sharing Network for Multi-Agent Reinforcement LearningabstractParameter-sharing (PS) techniques have been widely adopted in cooperative Multi-Agent Reinforcement Learning (MARL). In PS, all the agents share a policy network with identical parameters, which enjoys good sample efficiency. However, PS could lead to homogeneous policies that limit MARL performance. We tackle this problem from the angle of gradient conflict among agents. We find that the existence of futile neurons whose update is canceled out by gradient conflicts among agents leads to poor learning efficiency and diversity. To address this deficiency, we propose GradPS, a gradient-based PS method. It dynamically creates multiple clones for each futile neuron. For each clone, a group of agents with low gradient-conflict shares the neuron's parameters.
Our method can enjoy good sample efficiency by sharing the gradients among agents of the same clone neuron. Moreover, it can encourage diverse behaviors through independently updating an exclusive clone neuron. Through extensive experiments, we show that GradPS can learn diverse policies with promising performance. The source code for GradPS is available in \url{https://github.com/xmu-rl-3dv/GradPS}. Haoyuan Qin, Zhengzhu Liu, Chenxing Lin, Chennan Ma, Songzhu Mei, Cheng Wang 0003 |
ICML | 5 |
| 2025 | PlanU: Large Language Model Reasoning through Planning under UncertaintyabstractLarge Language Models (LLMs) are increasingly being explored across a range of reasoning tasks. However, LLMs sometimes struggle with reasoning tasks under uncertainty that are relatively easy for humans, such as planning actions in stochastic environments. The adoption of LLMs for reasoning is impeded by uncertainty challenges, such as LLM uncertainty and environmental uncertainty. LLM uncertainty arises from the stochastic sampling process inherent to LLMs. Most LLM-based Decision-Making (LDM) approaches address LLM uncertainty through multiple reasoning chains or search trees. However, these approaches overlook environmental uncertainty, which leads to poor performance in environments with stochastic state transitions.
Some recent LDM approaches deal with uncertainty by forecasting the probability of unknown variables. However, they are not designed for multi-step reasoning tasks that require interaction with the environment. To address uncertainty in LLM decision-making, we introduce PlanU, an LLM-based planning method that captures uncertainty within Monte Carlo Tree Search (MCTS). PlanU models the return of each node in the MCTS as a quantile distribution, which uses a set of quantiles to represent the return distribution. To balance exploration and exploitation during tree search, PlanU introduces an Upper Confidence Bounds with Curiosity (UCC) score which estimates the uncertainty of MCTS nodes. Through extensive experiments, we demonstrate the effectiveness of PlanU in LLM-based reasoning tasks under uncertainty. Ziwei Deng, Mian Deng, Chenjing Liang, Zeming Gao, Chennan Ma, Chenxing Lin, Songzhu Mei, Cheng Wang 0003 |
NeurIPS | 8 |
| 2024 | The Dormant Neuron Phenomenon in Multi-Agent Reinforcement Learning Value FactorizationabstractIn this work, we study the dormant neuron phenomenon in multi-agent reinforcement learning value factorization, where the mixing network suffers from reduced network expressivity caused by an increasing number of inactive neurons. We demonstrate the presence of the dormant neuron phenomenon across multiple environments and algorithms, and show that this phenomenon negatively affects the learning process. We show that dormant neurons correlates with the existence of over-active neurons, which have large activation scores. To address the dormant neuron issue, we propose ReBorn, a simple but effective method that transfers the weights from over-active neurons to dormant neurons. We theoretically show that this method can ensure the learned action preferences are not forgotten after the weight-transferring procedure, which increases learning effectiveness. Our extensive experiments reveal that ReBorn achieves promising results across various environments and improves the performance of multiple popular value factorization approaches. The source code of ReBorn is available in \url{https://github.com/xmu-rl-3dv/ReBorn}. Haoyuan Qin, Chennan Ma, Mian Deng, Zhengzhu Liu, Songzhu Mei, Cheng Wang 0003 |
NeurIPS | 5 |
| 2024 | High performance dilated convolutions on multi-core DSPs
Xiangdong Pei, Songzhu Mei, Rongchun Li, Jie Liu 0002 |
CCF Trans. High Perform. Comput. | 4 |
| 2023 | Optimizing Pointwise Convolutions on Multi-core DSPs
Xiangdong Pei, Songzhu Mei, Jie Liu 0002 |
ICA3PP (7) | 4 |
| 2023 | GraphMedia: Communication-balanced Graph Searching for Billion-scale Social Media AccessabstractThe graph has recently enabled substantial advances in big data analysis. As graphs are increasing from billions to trillions, efficient graph processing requires large-scale distributed clusters, which have up to thousands of nodes. For big data applications of which the computation is relatively simple, while the communication, especially for imbalanced communication is the bottleneck on distributed clusters, where huge numbers of small messages are transferred through 2D-topology networks. Graph partitioning is the dominant factor to affect the performance of large-scale distributed graph processing. Current graph partitioning policies have paid extensive attention to the utilization of the power law of big graphs but failed to exploit the advanced architectural benefits of 2D topology. To address such a problem, this paper presents GraphMedia, a communication-balanced graph partitioning for distributed search at scale. The key idea of GraphMedia is a communication-balanced partitioning to balance communication based on hardware/software co-design, in which the power law of graphs would be explored to average communication among nodes, and communication would be balanced between row and column by leveraging advanced 2D-topology knowledge. We use both benchmarks and real-world graphs to validate GraphMedia. Specially, GraphMedia-based Graph500 tests on the Tianhe supercomputer are superior to the fastest systems in the latest Graph500 lists (June 2022). We finally apply GraphMedia to real-world graphs for online graph media access, which outperforms the state-of-the-art graph partitioning and graph system by orders of magnitude. Xinbiao Gan, Peilin Guo, Jiaqi Si, Songzhu Mei, Cong Liu 0033 |
ACM Multimedia | 6 |
| 2023 | RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value FactorizationabstractMulti-agent systems are characterized by environmental uncertainty, varying policies of agents, and partial observability, which result in significant risks. In the context of Multi-Agent Reinforcement Learning (MARL), learning coordinated and decentralized policies that are sensitive to risk is challenging. To formulate the coordination requirements in risk-sensitive MARL, we introduce the Risk-sensitive Individual-Global-Max (RIGM) principle as a generalization of the Individual-Global-Max (IGM) and Distributional IGM (DIGM) principles. This principle requires that the collection of risk-sensitive action selections of each agent should be equivalent to the risk-sensitive action selection of the central policy. Current MARL value factorization methods do not satisfy the RIGM principle for common risk metrics such as the Value at Risk (VaR) metric or distorted risk measurements. Therefore, we propose RiskQ to address this limitation, which models the joint return distribution by modeling quantiles of it as weighted quantile mixtures of per-agent return distribution utilities. RiskQ satisfies the RIGM principle for the VaR and distorted risk metrics. We show that RiskQ can obtain promising performance through extensive experiments. The source code of RiskQ is available in https://github.com/xmu-rl-3dv/RiskQ. Chennan Ma, Weiquan Liu, Yongquan Fu, Songzhu Mei, Cheng Wang 0003 |
NeurIPS | 6 |
| 2022 | Optimizing Irregular-Shaped Matrix-Matrix Multiplication on Multi-Core DSPsabstractGeneral Matrix Multiplication (GEMM) has a wide range of applications in scientific simulation and artificial intelligence. Although traditional libraries can achieve high performance on large regular-shaped G EMMs, they often behave not well on irregular-shaped G EMMs, which are often found in new algorithms and applications of high-performance computing (HPC). Due to energy efficiency constraints, low-power multi-core digital signal processors (DSPs) have become an alternative architecture in HPC systems. Targeting multi-core DSPs in FT-m7032, a prototype CPU-DSPs heterogeneous processor for HPC, an efficient implementation-ftIMM - for three types of irregular-shaped GEMMs is proposed. FtIMM supports automatic generation of assembly micro-kernels, two parallelization strategies, and auto-tuning of block sizes and parallelization strategies. The experiments show that ftIMM can get better performance than the traditional GEMM implementations on multi-core DSPs in FT-m7032, yielding on up to 7.2x performance improvement, when performing on irregular-shaped GEMMs. And ftIMM on multi-core DSPs can also far outperform the open source library on multi-core CPUs in FT-m7032, delivering up to 3.1 x higher efficiency. Shangfei Yin, Ruochen Hao, Tianyang Zhou, Songzhu Mei, Jie Liu 0002 |
CLUSTER | 5 |
| 2022 | Optimizing Depthwise Convolutions on ARMv8 Architecture
Ruochen Hao, Shangfei Yin, Tianyang Zhou, Qingyang Zhang 0009, Songzhu Mei, Jie Liu 0002 |
PDCAT | 6 |
| 2021 | Edge Network Routing Protocol Base on Target Tracking ScenarioabstractAbstract Edge computing perfectly integrates cloud computing centers and edge-end devices together, but there are not many related researches on how the edge-end node devices work to form an edge network and what the protocols used to implement the communication among nodes in the edge network. Aiming at the problem of coordinated communication among edge nodes in the current edge computing network architecture, this paper proposes an edge network routing and forwarding protocol based on target tracking scenarios. This protocol can meet the dynamic changes of node locations, and the elastic expansion of node scale. Individual node failures will not affect the overall network, and the network ensures efficient real-time with less communication overhead. The experimental results display that the protocol can effectively reduce the communications volume of the edge network, improve the overall efficiency of the network, and set the optimal sampling period, so as to ensure that the network delay is minimized. Weihua Zhao, Ouhan Huang, Gangyong Jia, Youhuizi Li, Songzhu Mei, Duan Zhao |
Mob. Networks Appl. | 6 |
| 2021 | Evaluating FFT-based algorithms for strided convolutions on ARMv8 architectures
Xiandong Huang, Shuyu Lu, Ruochen Hao, Songzhu Mei, Jie Liu 0002 |
Perform. Evaluation | 5 |
| 2020 | Optimizing FFT-Based Convolution on ARMv8 Multi-core CPUs
Dongsheng Li 0001, Xiandong Huang, Songzhu Mei, Jie Liu 0002 |
Euro-Par | 5 |
| 2020 | Optimizing One by One Direct Convolution on ARMv8 Multi-core CPUsabstractConvolutional layers are ubiquitous in a variety of deep neural networks. Due to the lower computation complexity and the smaller number of parameters, convolutions with small filter sizes are often used, such as one by one convolution. Nevertheless, these small convolution operations are still time-consuming. A common approach to implementing convolutions is to transform them into matrix multiplications, known as GEMM-based convolutions. The approach maybe incurs additional memory overhead and calls matrix multiplication routines, which are not optimized for matrices generated by convolutions. In this paper, we present a new parallel one by one direct convolution implementation on ARMv8 multi-core CPUs, which doesn't incur any additional memory space requirement. Our implementation is verified on two ARMv8 CPUs, Phytium FT-1500A and FT-2000plus. In terms of performance and scalability, our implementation is better than GEMM-based implementations in all the tests on Phytium FT-1500A. On Phytium FT-2000plus, our approach gives much better performance and scalability than GEMM-based approaches in most cases. Dongsheng Li 0001, Songzhu Mei, Xiandong Huang |
JCC | 3 |
| 2019 | Parallel convolution algorithm using implicit matrix multiplication on multi-core CPUsabstractConvolution neural networks (CNNs) have been extensively used in machine learning applications. The most time-consuming part of CNNs are convolution operations. A common approach to implementing convolution operations is to recast them as general matrix multiplication, known as the im2col+GEMM approach. There are two main drawbacks of this approach. One is that large additional memory space is required. The other is the packing on the input elements of convolution operations are not memory-efficient enough. In this paper, we present a new parallel convolution algorithm using implicit matrix multiplication on multi-core CPUs. In comparison with Im2col+GEMM, our new algorithm can reduce the memory footprints and improve the packing efficiency. The experiment results on two ARV8-based multi-core CPUs demonstrate that our new algorithm gives much better performance and scalability than the im2col+GEMM method in most cases. Songzhu Mei, Jie Liu 0002, Chunye Gong |
IJCNN | 2 |
| 2018 | PruX: Communication Pruning of Parallel BFS in the Graph 500 Benchmark
Menghan Jia, Yiming Zhang 0003, Dongsheng Li 0001, Songzhu Mei |
ICA3PP (1) | 4 |
| 2018 | An Overview on the Convergence of High Performance Computing and Big Data ProcessingabstractAt present, with the rapid development of big data processing technology, streaming data processing and real-time data analysis have gradually become new research hotspots. Both the industry and the academia have invested a lot of research into the efficient processing methods of massive data generated in the environment such as the Internet and e-commerce. Meanwhile, high-performance computing technology and supercomputers are also looking for new business growth points. The convergence of big data processing and high performance computing technology is the general trend of big data analysis in the future. This paper will give a brief overview of typical technologies in the fusion process of big data processing and high-performance computing. Songzhu Mei, Hongtao Guan |
ICPADS | 1 |
| 2013 | Efficient revocation in ciphertext-policy attribute-based encryption based cryptographic cloud storageabstractIt is secure for customers to store and share their sensitive data in the cryptographic cloud storage. However, the revocation operation is a sure performance killer in the cryptographic access control system. To optimize the revocation procedure, we present a new efficient revocation scheme which is efficient, secure, and unassisted. In this scheme, the original data are first divided into a number of slices, and then published to the cloud storage. When a revocation occurs, the data owner needs only to retrieve one slice, and re-encrypt and re-publish it. Thus, the revocation process is accelerated by affecting only one slice instead of the whole data. We have applied the efficient revocation scheme to the ciphertext-policy attribute-based encryption (CP-ABE) based cryptographic cloud storage. The security analysis shows that our scheme is computationally secure. The theoretically evaluated and experimentally measured performance results show that the efficient revocation scheme can reduce the data owner’s workload if the revocation occurs frequently. Zhiying Wang 0003, Jun Ma 0015, Jiangjiang Wu, Songzhu Mei, Jiangchun Ren |
J. Zhejiang Univ. Sci. C | 5 |
| 2011 | SWHash: An Efficient Data Integrity Verification Scheme Appropriate for USB Flash DiskabstractData integrity verification is utmost important in trusted computing and Merkle trees are usually employed in implementation. However, the efficiency of data authentication is regarded as the main bottleneck in performance. In this paper, we propose an efficient data authentication protocol appropriate for a USB flash disk, named UTrustDisk (a trust-based intelligent disk). In our scheme, verification is speed up by using WH universal hash function and speculative caching. WH algorithm can hash message into a short digest at a high speed and the collision probability is almost negligible. Speculative caching will cache the potential hot chunks which can reduce the memory bandwidth pollution. In our experiments the success rate of speculation reaches 94.5% because the UTrustDisk's access mode is usually sequentially. The comparative experiment results show that SWHash average write throughput is 44.8% higher than NH scheme and 316% higher than SHA-1 scheme. Zhiying Wang 0003, Jiangjiang Wu, Songzhu Mei, Jiangchun Ren, Jun Ma 0015 |
TrustCom | 4 |
| 2011 | VerFAT: A Transparent and Efficient Multi-versioning Mechanism for FAT File SystemabstractEnsuring data reliability and continuity has played an important role for ensuring the information system still working normally when suffering from attack or other abnormal events. The existing data protection technologies are difficult to meet the fine-grained and precise data protection requirements on the widely used Windows platform. Inspiring from the data organization in the FAT file system, dedicating to reduce the number of direct disk write requests, we have designed a transparent and efficient multi-versioning mechanism for FAT file system, named VerFAT. During the data backup generation, responding to each file updating request, VerFAT generates multi-versioning data blocks. While we have achieved the goal of greatly improving the efficiency of data failure recovery by modifying the linking relationship of data blocks in FAT Table, merging the separate disk write operations on FAT Table and merging the separate disk write operations on directory entry when recovery a protected directory. Also we present the theoretical analysis for failure recovery mechanism in VerFAT. The experiment results on the prototype system have proved that our design is reasonable and efficient. Jiangjiang Wu, Zhiying Wang 0003, Songzhu Mei, Jiangchun Ren |
TrustCom | 3 |