EDBT 2026 Demo / reviewers in the wild / expert
Shaobo Ma
dblp:347/4857
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APT-LLM: Exploiting Arbitrary-Precision Tensor Core Computing for LLM AccelerationabstractLarge language models (LLMs) have revolutionized AI applications, yet their enormous computational demands severely limit deployment and real-time performance. Quantization methods can help reduce computational costs, however, attaining the extreme efficiency associated with ultra-low-bit quantized LLMs at arbitrary precision presents challenges on GPUs. This is primarily due to the limited support for GPU Tensor Cores, inefficient memory management, and inflexible kernel optimizations. To tackle these challenges, we propose a comprehensive acceleration scheme for arbitrary precision LLMs, namely APT-LLM. Firstly, we introduce a novel data format, bipolar-INT, which allows for efficient and lossless conversion with signed INT, while also being more conducive to parallel computation. We also develop a matrix multiplication (MatMul) method allowing for arbitrary precision by dismantling and reassembling matrices at the bit level. This method provides flexible precision and optimizes the utilization of GPU Tensor Cores. In addition, we propose a memory management system focused on data recovery, which strategically employs fast shared memory to substantially increase kernel execution speed and reduce memory access latency. Finally, we develop a kernel mapping method that dynamically selects the optimal configurable hyperparameters of kernels for varying matrix sizes, enabling optimal performance across different LLM architectures and precision settings. In LLM inference, APT-LLM achieves up to a 3.99× speedup compared to FP16 baselines and a 2.16× speedup over NVIDIA CUTLASS INT4 acceleration on RTX 3090. On RTX 4090 and H800, APT-LLM achieves up to 2.44× speedup over FP16 and 1.65× speedup over CUTLASS integer baselines. Shaobo Ma, Chao Fang 0005, Haikuo Shao, Zhongfeng Wang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | Efficient Arbitrary Precision Acceleration for Large Language Models on GPU Tensor CoresabstractLarge language models (LLMs) have been widely applied but face challenges in efficient inference. While quantization methods reduce computational demands, ultra-low bit quantization with arbitrary precision is hindered by limited GPU Tensor Core support and inefficient memory management, leading to suboptimal acceleration. To address these challenges, we propose a comprehensive acceleration scheme for arbitrary precision LLMs. At its core, we introduce a novel bipolar-INT data format that facilitates parallel computing and supports symmetric quantization, effectively reducing data redundancy. Building on this, we implement an arbitrary precision matrix multiplication scheme that decomposes and recovers matrices at the bit level, enabling flexible precision while maximizing GPU Tensor Core utilization. Furthermore, we develop an efficient matrix preprocessing method that optimizes data layout for subsequent computations. Finally, we design a data recovery-oriented memory management system that strategically utilizes fast shared memory, significantly enhancing kernel execution speed and minimizing memory access latency. Experimental results demonstrate our approach's effectiveness, with up to 2.4× speedup in matrix multiplication compared to NVIDIA's CUTLASS. When integrated into LLMs, we achieve up to 6.7× inference acceleration. These improvements significantly enhance LLM inference efficiency, enabling broader and more responsive applications of LLMs. Shaobo Ma, Chao Fang 0005, Haikuo Shao, Zhongfeng Wang 0001 |
ASP-DAC | 1 |
| 2025 | MeHyper: Accelerating Hypergraph Neural Networks by Exploring Implicit DataflowsabstractHypergraph Neural Networks (HGNNs) are increasingly utilized to analyze complex inter-entity relationships. Traditional HGNN systems, based on a hyperedge-centric dataflow model, independently process aggregation tasks for hyperedges and vertices, leading to significant computational redundancy. This redundancy arises from recalculating shared information across different tasks. For the first time, we identify and harness implicit dataflows (i.e., dependencies) within HGNNs, introducing the microedge concept to effectively capture and reuse intricate shared information among aggregation tasks, thereby minimizing redundant computations. We have developed a new microedge-centric dataflow model that processes shared information as fine-grained microedge aggregation tasks. This dataflow model is supported by the Read-Process-Activate-Generate execution model, which aims to optimize parallelism among these tasks. Furthermore, our newly developed MeHyper, a microedge-centric HGNN accelerator, incorporates a decoupled pipeline for improved computational parallelism and a hierarchical feature management strategy to reduce off-chip memory accesses for large volumes of intermediate feature vectors generated. Our evaluation demonstrates that MeHyper substantially outperforms the leading CPUbased system PyG-CPU and the GPU-based system HyperGef, delivering performance improvements of $1,032.23 \times$ and $10.51 \times$, and energy efficiencies of $1,169.03 \times$ and $9.96 \times$, respectively. Wenju Zhao, Pengcheng Yao, Dan Chen 0006, Long Zheng 0003, Xiaofei Liao, Qinggang Wang, Shaobo Ma, Haifeng Liu 0003, Wenjing Xiao, Hai Jin 0001, Jingling Xue |
HPCA | 7 |
| 2024 | Co-Designing Binarized Transformer and Hardware Accelerator for Efficient End-to-End Edge DeploymentabstractTransformer models have revolutionized AI tasks, but their large size hinders real-world deployment on resource-constrained and latency-critical edge devices. While binarized Transformers offer a promising solution by significantly reducing model size, existing approaches suffer from algorithm-hardware mismatches with limited co-design exploration, leading to suboptimal performance on edge devices. Hence, we propose a co-design method for efficient end-to-end edge deployment of Transformers from three aspects: algorithm, hardware, and joint optimization. First, we propose BMT, a novel hardware-friendly binarized Transformer with optimized quantization methods and components, and we further enhance its model accuracy by leveraging the weighted ternary weight splitting training technique. Second, we develop a streaming processor mixed binarized Transformer accelerator, namely BAT, which is equipped with specialized units and scheduling pipelines for efficient inference of binarized Transformers. Finally, we co-optimize the algorithm and hardware through a design space exploration approach to achieve a global trade-off between accuracy, latency, and robustness for real-world deployments. Experimental results show our co-design achieves up to 2.14~49.37× throughput gains and 3.72~88.53× better energy efficiency over state-of-the-art Transformer accelerators, enabling efficient end-to-end edge deployment. Yuhao Ji, Chao Fang 0005, Shaobo Ma, Haikuo Shao, Zhongfeng Wang 0001 |
ICCAD | 3 |
| 2023 | An Investigation on Intelligent Relay assisted Semantic Communication NetworksabstractWith the development of sixth generation (6G) networks, semantic communication is treated as an emerging technology to provide more intelligent communication between two parties for enhance the accuracy and efficiency of communications. Whereas, it demands to merge all physical layer blocks in the conventional communications, which heavily deteriorates the spectrum sparsity problem of wireless communication networks. For this sake, we develop a novel intelligent relay assisted semantic communications, by enhancing communication efficiency, which creates a paradigm for text transmission fusion in 6G networks. In this contribution, the semantic communication system we have designed combines traditional deep learning methods with intelligent semantic relays, based on which, the protocols about minimizing the semantic errors by recovering the meaning of sentences for addressing channel variations issues. Meanwhile, our designed system could be applied in the case of wireless channel deterioration or knowledge background mismatch at the transmitter and receiver, and is also a paradigm for future applications in one-to-many and many-to-many semantic communication. At last, we propose a range of research directions and open challenges for boosting the implementations of relay assisted semantic communications. Shaobo Ma, Wei Liang 0002, Boxuan Zhang 0003, Dawei Wang 0001 |
WCNC | 1 |