EDBT 2026 Demo / reviewers in the wild / expert
Gu Gong
dblp:214/3799
· DBLP profile ↗
4ranked-venue papers
1as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DynSpAttn: Efficient Attention via Dual-Side Dynamic Sparsity on Sparse Tensor CoresabstractThe high computational complexity of the self-attention mechanism constitutes a primary performance bottleneck in LLM inference. Existing sparse attention mechanisms commonly adopt coarse-grained block sparsity to align with FlashAttention’s tiling and rely on dense Tensor Cores, leaving the potential of emerging hardware Sparse Tensor Cores (SpTCs) and semi-structured sparsity largely untapped. We present DynSpAttn, a dynamic sparse attention mechanism co-designed with NVIDIA Sparse Tensor Cores. DynSpAttn introduces a dual-side 2:4 structured sparsity strategy that prunes both the Query and Score matrices, thereby transforming the dominant matrix multiplications in attention into sparse matrix multiplications (SpMMs) executable on SpTCs. To realize this transformation, DynSpAttn incorporates lightweight in-register pruners and a shuffle-free operand remapping scheme within a fully fused, I/O-aware CUDA kernel. Evaluations on RTX 4090 and L20 GPUs show that DynSpAttn achieves up to 1.70 × kernel-level preformance improvement over FlashAttention and 1.58 × end-to-end inference speedup. These results demonstrate that co-designing semi-structured sparsity with hardware support across the full attention pipeline provides a practical and efficient solution for LLM inference. Xiangrui Yu, Ruibo Fan, Weile Luo, Gu Gong, Xiaowen Chu 0001 |
ICS | 5 |
| 2025 | SpInfer: Leveraging Low-Level Sparsity for Efficient Large Language Model Inference on GPUsabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities, but their immense scale poses significant challenges in terms of both memory and computational costs. While unstructured pruning offers promising solutions by introducing sparsity to reduce resource requirements, realizing its benefits in LLM inference remains elusive. This is primarily due to the storage overhead of indexing non-zero elements and the inefficiency of sparse matrix multiplication (SpMM) kernels at low sparsity levels (around 50%). In this paper, we present SpInfer, a high-performance framework tailored for sparsified LLM inference on GPUs. SpInfer introduces Tensor-Core-Aware Bitmap Encoding (TCA-BME), a novel sparse format that minimizes indexing overhead by leveraging efficient bitmap-based indexing, optimized for GPU Tensor Core architectures. Furthermore, SpInfer integrates an optimized SpMM kernel with Shared Memory Bitmap Decoding (SMBD) and asynchronous pipeline design to enhance computational efficiency. Experimental results show that SpInfer significantly outperforms state-of-the-art SpMM implementations (up to 2.14× and 2.27× over Flash-LLM and SparTA, respectively) across a range of sparsity levels (30% to 70%), with substantial improvements in both memory efficiency and end-to-end inference speed (up to 1.58×). SpInfer outperforms highly optimized cuBLAS at sparsity levels as low as 30%, marking the first effective translation of unstructured pruning's theoretical advantages into practical performance gains for LLM inference. Ruibo Fan, Xiangrui Yu, Peijie Dong, Gu Gong, Qiang Wang 0022, Wei Wang 0030, Xiaowen Chu 0001 |
EuroSys | 5 |
| 2025 | Event-Based Photometric Gaussian Mixture Models for Visual ServoingabstractThis paper presents a novel approach for visual servoing using neuromorphic event-based cameras. We extend the photometric Gaussian mixture model framework from frame-based to event-based vision by developing a mathematical formulation that bridges conventional models with the sparse, temporally-precise nature of event data. Our method transforms raw event streams into effective visual features through time surface representations, enabling visual servoing that leverages the microsecond temporal resolution and high dynamic range of event cameras. Evaluation using the N-Caltech101 dataset demonstrates excellent convergence characteristics and a high success rate (96.9%) across diverse object categories. Results confirm that our event-based photometric Gaussian mixture approach effectively exploits the temporal precision of event cameras while providing reliable performance for robot control tasks. Gu Gong, Qiang Wang 0001, David Navarro-Alarcon |
IECON | 1 |
| 2019 | A Multi-View-Based Collective Entity Linking MethodabstractFacing lots of name mentions appearing on the web, entity linking is essential for many information processing applications. To improve linking accuracy, the relations between entities are usually considered in the linking process. This kind of method is called collective entity linking and can obtain high-quality results. There are two kinds of information helpful to reveal the relations between entities, i.e., contextual information and structural information of entities. Most traditional collective entity linking methods consider them separately. In fact, these two kinds of information represent entities from specific and diverse views and can enhance each other, respectively. Besides, if we look into each view closely, it can be separated into sub-views that are more meaningful. For this reason, this article proposes a multi-view–based collective entity linking algorithm, which combines several views of entities into an objective function for entity linking. The importance of each view can be valued and the linking results can be obtained along with resolving this objective function. Experimental results demonstrate that our linking algorithm can acquire higher accuracy than many state-of-the-art entity linking methods. Besides, since we simplify the entity's structure and change the entity linking to a sub-matrix searching problem, our algorithm also obtains high efficiency. Ming Liu 0004, Gu Gong, Bing Qin 0001, Ting Liu 0001 |
ACM Trans. Inf. Syst. | 2 |