Dongge Qin

dblp:345/0083 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0003-7679-0387ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Efficient and distributed learning · 50% Language models and text generation · 38% Deep learning architectures and training · 12%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
GPUs and heterogeneous computing · 50% Hardware accelerators and domain-specific architectures · 38% Reconfigurable computing and FPGAs · 12%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › inference efficiency
LLM inference optimization
1.012026
Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference · AAAI 2026
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning
1.012026
Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference · AAAI 2026
GPUs and heterogeneous computing
heterogeneous architecture
1.012026
DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU · ASPLOS (2) 2026
Hardware accelerators and domain-specific architectures › efficient inference
speculative decoding
1.012026
DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU · ASPLOS (2) 2026
Machine learning › Efficient and distributed learning
model compression
0.312026
Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference · AAAI 2026
Machine learning › Deep learning architectures and training › mixture of experts
sparse expert activation
0.312026
Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference · AAAI 2026
Reconfigurable computing and FPGAs
FPGA accelerator
0.312026
DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU · ASPLOS (2) 2026
GPUs and heterogeneous computing
GPU computing
0.312026
DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU · ASPLOS (2) 2026

Methods — techniques the papers use, named apart from their topics

tree-based reasoning · 1.0speculative decoding · 1.0mixture of experts · 1.0gating network · 1.0collaborative debate · 1.0
YearPublicationVenuePosition
2026 Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference
abstract
We introduce Mixture-of-Trees (MoT), a novel framework that integrates sparse expert activation with structured tree-based reasoning for efficient LLM inference. MoT employs a learned gating mechanism to selectively activate only the most relevant expert reasoning trees for each problem, where experts use models of varying capacities based on task complexity. The framework features three key innovations: (1) sparse expert activation through unified gating networks, (2) specialized expert trees that leverage domain-specific expertise while optimizing the quality-efficiency trade-off, and (3) collaborative debate mechanisms for conflicting solutions. Additionally, MoT includes a shared baseline tree with early stopping—activated experts perform lightweight validation and terminate early when confidence is high. Experiments across five benchmarks (GSM8K, MATH, AIME 2024, MMLU, HotpotQA) show that MoT achieves 2-7 percentage point accuracy improvements while reducing LLM calls by 37-40% compared to existing multi-path methods.
Yangbo Wei, Zhen Huang 0007, Shaoqiang Lu, Junhong Qian, Dongge Qin, Ting-Jung Lin, Wei W. Xing, Lei He 0001
AAAI5
2026 dLLM-OPU: An FPGA Overlay Processor for Accelerated Diffusion Large Language Models
abstract
Large Language Models (LLMs) are achieving unprecedented performance across diverse tasks, benefiting from autoregressive generation. However, this left-to-right decoding paradigm inherently limits contextual understanding quality. Diffusion-based LLMs (dLLMs) offer a promising alternative by iteratively refining sequences via denoising, enabling stronger bidirectional context modeling and improved generation quality. However, dLLMs face two main challenges: redundant computation and memory overhead in multi-step denoising, and excessive inference cost from over-denoising under fixed-step schedules. To address these issues, we propose dLLM-OPU, an FPGA overlay processor to accelerate dLLMs. Our solution features two key innovations: (1) a Region-Adaptive Caching for Dynamic Column Sparsity Framework that exploits temporal locality for selective recomputation without model retraining, and (2) a Token Entropy-based Early Stopping strategy that dynamically terminates the denoising process based on token-level convergence metrics. We implement these innovations through a specialized sparse processing element (PE) array that maximizes top-k sparsity utilization by minimizing idle cycles via row-column concatenation, complemented by an efficient cache management system that reduces memory access latency and a flexible entropybased decoding unit. Implemented on a U200 FPGA, dLLM-OPU achieves $2.2 \times-5.1 \times$ speedup and $7.6 \times-20.3 \times$ energy efficiency over RTX4090 in LLaDA.
Yangbo Wei, Shaoqiang Lu, Junhong Qian, Lei He 0001, Dongge Qin, Xiao Shi 0001
ASP-DAC5
2026 DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU
Shaoqiang Lu, Yangbo Wei, Junhong Qian, Dongge Qin, Shiji Gao, Yizhi Ding, Xiao Shi 0001, Lei He 0001
ASPLOS (2)4
2026 Harnessing Spatiotemporal Redundancy for Fast Diffusion Models on FPGA
Dongge Qin, Junhong Qian, Shaoqiang Lu, Yangbo Wei, Ruizhe Deng, Xiao Shi 0001, Longxing Shi, Lei He 0001
ISCAS1
2023 Area and power optimization for Fixed Polarity Reed-Muller logic circuits based on Multi-strategy Multi-objective Artificial Bee Colony algorithm
abstract
Area and power optimization of Fixed Polarity Reed–Muller (FPRM) circuits has received a lot of attention. Polarity optimization for FPRM circuits is essentially a binary multi-objective optimization problem. However, the existing area and power optimization approaches for FPRM logic circuits rarely produce a frontier and a greater number of Pareto optimal solutions . In this paper, a Multi-strategy Multi-objective Artificial Bee Colony (MMABC) algorithm is proposed to solve the binary multi-objective optimization problem. The main innovation of MMABC can be summarized as follows: a flexible foraging behavior strategy for employed bees is proposed to improve the searching ability of the algorithm; a genetic retention evolution for onlooker bees is proposed to improve the quality of the population; an efficient transform strategy is proposed to help the algorithm to jump out the local optimal and increase convergence speed. Moreover, we propose an area and power optimization approach for FPRM logic circuits, which uses the MMABC to search for the polarities (i.e., Pareto optimal solutions) with smaller area and lower power. Experimental results demonstrated the effectiveness and superiority of our approach in optimizing area and power of FPRM logic circuits.
Dongge Qin, Zhenxue He, Xiaojun Zhao, Jia Liu 0054, Fan Zhang 0037, Limin Xiao 0002
Eng. Appl. Artif. Intell.1