Yangbo Wei

dblp:400/4705 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2026
0009-0000-3678-8942ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Mixture-of-Trees: Learning to Select and Weigh Reasoning Paths for Efficient LLM Inference
abstract
We introduce Mixture-of-Trees (MoT), a novel framework that integrates sparse expert activation with structured tree-based reasoning for efficient LLM inference. MoT employs a learned gating mechanism to selectively activate only the most relevant expert reasoning trees for each problem, where experts use models of varying capacities based on task complexity. The framework features three key innovations: (1) sparse expert activation through unified gating networks, (2) specialized expert trees that leverage domain-specific expertise while optimizing the quality-efficiency trade-off, and (3) collaborative debate mechanisms for conflicting solutions. Additionally, MoT includes a shared baseline tree with early stopping—activated experts perform lightweight validation and terminate early when confidence is high. Experiments across five benchmarks (GSM8K, MATH, AIME 2024, MMLU, HotpotQA) show that MoT achieves 2-7 percentage point accuracy improvements while reducing LLM calls by 37-40% compared to existing multi-path methods.
Yangbo Wei, Zhen Huang 0007, Shaoqiang Lu, Junhong Qian, Dongge Qin, Ting-Jung Lin, Wei W. Xing, Lei He 0001
AAAI1
2026 MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text
abstract
Artificial intelligence has demonstrated significant potential in clinical decision-making; however, developing models capable of adapting to diverse real-world scenarios and performing complex diagnostic reasoning remains a major challenge. Existing medical multi-modal benchmarks are typically limited to single-image, single-turn tasks, lacking multi-modal medical image integration and failing to capture the longitudinal and multi-modal interactive nature inherent to clinical practice. To address this gap, we introduce MedAtlas, a novel benchmark framework designed to evaluate large language models on realistic medical reasoning tasks. MedAtlas is characterized by four key features: multi-round visual question answering (VQA), Joint reasoning of multiple modalities of medical images, multi-task integration, and high clinical fidelity. It supports four core tasks: open-ended multi-round VQA, closed-ended multi-round VQA, multi-image joint reasoning, and comprehensive disease diagnosis. Each case is derived from real diagnostic workflows and incorporates temporal interactions between textual medical histories and multiple imaging modalities, including CT, MRI, PET, ultrasound, X-ray, etc., requiring models to perform deep integrative reasoning across images and clinical texts. MedAtlas provides expert-annotated gold standards for all tasks. Furthermore, we propose two novel evaluation metrics: Stage Chain Accuracy (SCA) and Error Propagation Suppression Coefficient (EPSC). Benchmark results with existing multi-modal models reveal substantial performance gaps in multi-stage clinical reasoning. MedAtlas establishes a challenging evaluation platform to advance the development of robust and trustworthy medical AI.
Ronghao Xu, Zhen Huang 0007, Yangbo Wei, Xiaoqian Zhou, Zihang Jiang, Shaohua Kevin Zhou
AAAI3
2026 VFlow: Discovering Optimal Agentic Workflows for Verilog Generation
Yangbo Wei, Zhen Huang 0007, Lei He 0001, Ting-Jung Lin, Wei W. Xing
ASP-DAC1
2026 dLLM-OPU: An FPGA Overlay Processor for Accelerated Diffusion Large Language Models
abstract
Large Language Models (LLMs) are achieving unprecedented performance across diverse tasks, benefiting from autoregressive generation. However, this left-to-right decoding paradigm inherently limits contextual understanding quality. Diffusion-based LLMs (dLLMs) offer a promising alternative by iteratively refining sequences via denoising, enabling stronger bidirectional context modeling and improved generation quality. However, dLLMs face two main challenges: redundant computation and memory overhead in multi-step denoising, and excessive inference cost from over-denoising under fixed-step schedules. To address these issues, we propose dLLM-OPU, an FPGA overlay processor to accelerate dLLMs. Our solution features two key innovations: (1) a Region-Adaptive Caching for Dynamic Column Sparsity Framework that exploits temporal locality for selective recomputation without model retraining, and (2) a Token Entropy-based Early Stopping strategy that dynamically terminates the denoising process based on token-level convergence metrics. We implement these innovations through a specialized sparse processing element (PE) array that maximizes top-k sparsity utilization by minimizing idle cycles via row-column concatenation, complemented by an efficient cache management system that reduces memory access latency and a flexible entropybased decoding unit. Implemented on a U200 FPGA, dLLM-OPU achieves $2.2 \times-5.1 \times$ speedup and $7.6 \times-20.3 \times$ energy efficiency over RTX4090 in LLaDA.
Yangbo Wei, Shaoqiang Lu, Junhong Qian, Lei He 0001, Dongge Qin, Xiao Shi 0001
ASP-DAC1
2026 DFVG: A Heterogeneous Architecture for Speculative Decoding with Draft-on-FPGA and Verify-on-GPU
Shaoqiang Lu, Yangbo Wei, Junhong Qian, Dongge Qin, Shiji Gao, Yizhi Ding, Xiao Shi 0001, Lei He 0001
ASPLOS (2)2
2026 TinyDenseUNet: Efficient Dense Feature Reuse for Lightweight Medical Image Segmentation
Suhua Wang, Guizhou Ding, Yangbo Wei, Guangzong Si, Xiaoxin Sun
ICIC (30)3
2026 Harnessing Spatiotemporal Redundancy for Fast Diffusion Models on FPGA
Dongge Qin, Junhong Qian, Shaoqiang Lu, Yangbo Wei, Ruizhe Deng, Xiao Shi 0001, Longxing Shi, Lei He 0001
ISCAS4
2025 H3DE-Net: Efficient and Accurate 3D Landmark Detection in Medical Imaging
abstract
Landmark detection is essential in medical image analysis, aiding tasks like surgical navigation, diagnosis, and treatment planning. However, it remains challenging due to the need for fine-grained local detail and long-range spatial dependency modeling in high-dimensional volumetric data. Existing approaches struggle to balance accuracy, efficiency, and robustness, especially in cases of sparse landmark distribution, anatomical variability, and noisy or incomplete scans. We propose H3DE-Net, a hybrid framework combining CNNs for local feature extraction and a lightweight transformer-based attention module for global context. A volumetric bi-level routing attention mechanism reduces computational overhead while preserving long-range dependencies, and multi-scale feature fusion enhances precision and robustness. This design integrates both global and local representations, overcoming the limitations of CNN-only or transformer-only models. Extensive experiments on a public CT dataset show that H3DE-Net achieves state-of-the-art performance, significantly improving mean radial error (MRE) and success detection rate (SDR) compared to existing methods. The model is robust in challenging scenarios with missing landmarks or anatomical variations, demonstrating its applicability in real-world clinical settings. All code, pretrained weights, and data processing scripts are publicly available for reproducibility and further research.
Ronghao Xu, Yangbo Wei, Wenkai Yang, Suhua Wang, Xiaoxin Sun, Qingsong Yao
BIBM4
2025 MoE-OPU: An FPGA Overlay Processor Leveraging Expert Parallelism for MoE-based Large Language Models
abstract
The advent of Large Language Models (LLMs) like DeepSeek, empowered by the Mixture-of-Experts (MoE) architecture, has driven significant advancements across diverse applications. However, a critical challenge arises during inference: Only a small fraction of experts are activated, causing severe token allocation imbalances among experts. This inefficiency poses substantial storage and computational burdens on resource-constrained devices, exacerbated by the lack of optimization strategies that integrate expert usage-aware parameter pruning and parallel scheduling, ultimately leading to suboptimal resource utilization. To address these limitations, we propose MoE-OPU, an FPGA-based overlay processor that optimizes parallel MoE inference through three key innovations. First, we introduce N:M sparsity (1:4/2:4/4:8/6:8/8:8) in the MLP layers and mixed-precision quantization (BF16/FP8/INT4) guided by expert activation frequency, reducing the parameter size by up to 2.76× while maintaining model accuracy (only 1.53% average drop after fine-tuning). Second, a lightweight prediction network dynamically predicts next-layer "hot" experts by analyzing historical activation patterns and current hidden states, achieving an average prediction hit rate of 83.4%. Third, a reconfigurable multi-core architecture maximizes the utilization of HBM bandwidth via a systolic array that natively supports sparse and mixed-precision computations, coupled with parallel concatenation to balance compute and memory efficiency. Experimental results on a Xilinx V80 FPGA with the DeepSeek-V2-lite model demonstrate that MoE-OPU outperforms the NVIDIA A100 GPU, delivering a 6.78× higher token throughput. Compared to RTX 4090 and U200 FPGA, MoE-OPU achieves 13.37× and 7.85× improvements, respectively. These advancements highlight the potential of algorithm-hardware co-design for scalable deployment of MoE-based LLMs on edge devices.
Shaoqiang Lu, Yangbo Wei, Junhong Qian, Xiao Shi 0001, Lei He 0001
ICCAD2
2025 ModelGen: Automating Semiconductor Parameter Extraction with Large Language Model Agents
abstract
Device models require large numbers of parameters to characterize complex physical effects. Although the latest advancements in machine learning and automated tools have drastically improved efficiency over the classic methods, they still demand a considerable amount of human intervention in the loop to gain accuracy. This drastically limits further automation. Inspired by the success of Multimodal Large Language Models (MLLMs) in addressing tasks across diverse fields, we propose ModelGen, the first in-depth study to leverage MLLMs with RAG (Retrieval-Augmented Generation) to significantly reduce human effort in parameter extraction for compact model. Our contributions include (1) Automated Agentic Workflow Construction that learns to build and refine extraction workflows through iterative optimization, (2) MLLM Judge, a visual scoring mechanism that evaluates fitting quality using actual device characteristic plots rather than simple numerical metrics, and (3) Model-specific RAG for providing relevant domain knowledge during the extraction process. Experimental results demonstrate that ModelGen achieves a 26.8%–33.1% improvement in pass@1,3,5 compared to base LLM methods. The system completes complex model extractions for BSIMs and ASM-HEMT in hours (up to 168× faster) rather than days or weeks, making parameter extraction more accessible to non-experts while maintaining professional engineer-level accuracy.
Yangbo Wei, Zhanfei Chen, Jinlong Yan, Ting-Jung Lin, Zhen Huang 0007, Wei W. Xing, Lei He 0001
ACM Trans. Design Autom. Electr. Syst.1