VLDB 2026 Research / reviewers in the wild / expert
Dongge Qin
dblp:345/0083
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0003-7679-0387ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Efficient and distributed learning · 50% Language models and text generation · 38% Deep learning architectures and training · 12% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
GPUs and heterogeneous computing · 50% Hardware accelerators and domain-specific architectures · 38% Reconfigurable computing and FPGAs · 12% |
Topics — the 8 heaviest of 8, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › inference efficiency
LLM inference optimization |
1.0 | 1 | 2026 | Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning |
1.0 | 1 | 2026 | Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference · AAAI 2026 |
GPUs and heterogeneous computing
heterogeneous architecture |
1.0 | 1 | 2026 | DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU · ASPLOS (2) 2026 |
Hardware accelerators and domain-specific architectures › efficient inference
speculative decoding |
1.0 | 1 | 2026 | DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU · ASPLOS (2) 2026 |
Machine learning › Efficient and distributed learning
model compression |
0.3 | 1 | 2026 | Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference · AAAI 2026 |
Machine learning › Deep learning architectures and training › mixture of experts
sparse expert activation |
0.3 | 1 | 2026 | Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference · AAAI 2026 |
Reconfigurable computing and FPGAs
FPGA accelerator |
0.3 | 1 | 2026 | DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU · ASPLOS (2) 2026 |
GPUs and heterogeneous computing
GPU computing |
0.3 | 1 | 2026 | DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU · ASPLOS (2) 2026 |
Methods — techniques the papers use, named apart from their topics
tree-based reasoning · 1.0speculative decoding · 1.0mixture of experts · 1.0gating network · 1.0collaborative debate · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM InferenceabstractWe introduce Mixture-of-Trees (MoT), a novel framework that integrates sparse expert activation with structured tree-based reasoning for efficient LLM inference. MoT employs a learned gating mechanism to selectively activate only the most relevant expert reasoning trees for each problem, where experts use models of varying capacities based on task complexity. The framework features three key innovations: (1) sparse expert activation through unified gating networks, (2) specialized expert trees that leverage domain-specific expertise while optimizing the quality-efficiency trade-off, and (3) collaborative debate mechanisms for conflicting solutions. Additionally, MoT includes a shared baseline tree with early stopping—activated experts perform lightweight validation and terminate early when confidence is high. Experiments across five benchmarks (GSM8K, MATH, AIME 2024, MMLU, HotpotQA) show that MoT achieves 2-7 percentage point accuracy improvements while reducing LLM calls by 37-40% compared to existing multi-path methods. Yangbo Wei, Zhen Huang 0007, Shaoqiang Lu, Junhong Qian, Dongge Qin, Ting-Jung Lin, Wei W. Xing, Lei He 0001 |
AAAI | 5 |
| 2026 | dLLM-OPU: An FPGA Overlay Processor for Accelerated Diffusion Large Language ModelsabstractLarge Language Models (LLMs) are achieving unprecedented performance across diverse tasks, benefiting from autoregressive generation. However, this left-to-right decoding paradigm inherently limits contextual understanding quality. Diffusion-based LLMs (dLLMs) offer a promising alternative by iteratively refining sequences via denoising, enabling stronger bidirectional context modeling and improved generation quality. However, dLLMs face two main challenges: redundant computation and memory overhead in multi-step denoising, and excessive inference cost from over-denoising under fixed-step schedules. To address these issues, we propose dLLM-OPU, an FPGA overlay processor to accelerate dLLMs. Our solution features two key innovations: (1) a Region-Adaptive Caching for Dynamic Column Sparsity Framework that exploits temporal locality for selective recomputation without model retraining, and (2) a Token Entropy-based Early Stopping strategy that dynamically terminates the denoising process based on token-level convergence metrics. We implement these innovations through a specialized sparse processing element (PE) array that maximizes top-k sparsity utilization by minimizing idle cycles via row-column concatenation, complemented by an efficient cache management system that reduces memory access latency and a flexible entropybased decoding unit. Implemented on a U200 FPGA, dLLM-OPU achieves $2.2 \times-5.1 \times$ speedup and $7.6 \times-20.3 \times$ energy efficiency over RTX4090 in LLaDA. Yangbo Wei, Shaoqiang Lu, Junhong Qian, Lei He 0001, Dongge Qin, Xiao Shi 0001 |
ASP-DAC | 5 |
| 2026 | DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU
Shaoqiang Lu, Yangbo Wei, Junhong Qian, Dongge Qin, Shiji Gao, Yizhi Ding, Xiao Shi 0001, Lei He 0001 |
ASPLOS (2) | 4 |
| 2026 | Harnessing Spatiotemporal Redundancy for Fast Diffusion Models on FPGA
Dongge Qin, Junhong Qian, Shaoqiang Lu, Yangbo Wei, Ruizhe Deng, Xiao Shi 0001, Longxing Shi, Lei He 0001 |
ISCAS | 1 |
| 2023 | Area and power optimization for Fixed Polarity Reed-Muller logic circuits based on Multi-strategy Multi-objective Artificial Bee Colony algorithmabstractArea and power optimization of Fixed Polarity Reed–Muller (FPRM) circuits has received a lot of attention. Polarity optimization for FPRM circuits is essentially a binary multi-objective optimization problem. However, the existing area and power optimization approaches for FPRM logic circuits rarely produce a frontier and a greater number of Pareto optimal solutions . In this paper, a Multi-strategy Multi-objective Artificial Bee Colony (MMABC) algorithm is proposed to solve the binary multi-objective optimization problem. The main innovation of MMABC can be summarized as follows: a flexible foraging behavior strategy for employed bees is proposed to improve the searching ability of the algorithm; a genetic retention evolution for onlooker bees is proposed to improve the quality of the population; an efficient transform strategy is proposed to help the algorithm to jump out the local optimal and increase convergence speed. Moreover, we propose an area and power optimization approach for FPRM logic circuits, which uses the MMABC to search for the polarities (i.e., Pareto optimal solutions) with smaller area and lower power. Experimental results demonstrated the effectiveness and superiority of our approach in optimizing area and power of FPRM logic circuits. Dongge Qin, Zhenxue He, Xiaojun Zhao, Jia Liu 0054, Fan Zhang 0037, Limin Xiao 0002 |
Eng. Appl. Artif. Intell. | 1 |