VLDB 2026 Research / reviewers in the wild / expert
Yubin Qin
dblp:295/6981
· DBLP profile ↗
10ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0001-5530-5416ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts
Yushu Zhao, Yubin Qin, Yang Wang 0089, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ASP-DAC | 2 |
| 2026 | An Energy-Efficient Transformer Fine-Tuning Processor for Personalized Edge ApplicationsabstractTransformer models have achieved remarkable success in various domains. Given concerns about user privacy, there is an urgent need for on-device fine-tuning of Transformer models at the edge. Transformer fine-tuning faces three key challenges: 1)$O(n^{3})$re-computations during BP/WG save only$O(n^{2})$storage, limiting batch size for fine-tuning speedup. 2)Weakly related tokens account for 87.9% of computations but contribute only 8.7% to accuracy. 3)89.1% of multiplications in matrix multiplications (MM) involve dual near-zero operands, leading to a$1.9\times $increase in logic toggling energy due to frequent exponent/mantissa variations near zero. This paper proposes a Transformer-based processor supporting energy-efficient fine-tuning with three key features to tackle the above challenges. 1)An exponent-stationary re-computing scheduler (ESRS) reduces 44.2% of the storage requirement for each batch. 2)An aggressive linear fitting unit (ALFU) saves 47.4% of the computations in each iteration. 3)A logarithmic domain processing element (LDPE) decreases 36.3% of energy for MM in fine-tuning. Fabricated with 22nm technology, the proposed processor has an area of 6.4 mm2. The proposed Transformer processor achieves a peak energy efficiency of 54.94 TFLOPS/W. It reduces fine-tuning energy by$4.27\times $and offers$3.57\times $speedup for GPT-2. Yang Wang 0089, Yubin Qin, Wende Xu, Zhiheng Yue, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video UnderstandingabstractVision Language models (VLMs) have achieved remarkable success in video understanding tasks. Yet, a key question remains: Do they comprehend visual information or merely learn superficial mappings between visual and textual patterns?
Understanding visual cues, particularly those related to physics and common sense, is crucial for AI systems interacting with the physical world. However, existing VLM evaluations primarily rely on positive-control tests using real-world videos that resemble training distributions. While VLMs perform well on such benchmarks, it is unclear whether they grasp underlying visual and contextual signals or simply exploit visual-language correlations. To fill this gap, we propose incorporating negative-control tests, i.e., videos depicting physically impossible or logically inconsistent scenarios, and evaluating whether models can recognize these violations. True visual understanding should evince comparable performance across both positive and negative tests. Since such content is rare in the real world, we introduce VideoHallu, a synthetic video dataset featuring physics- and commonsense-violating scenes generated using state-of-the-art tools such as Veo2, Sora, and Kling. The dataset includes expert-annotated question-answer pairs spanning four categories of physical and commonsense violations, designed to be straightforward for human reasoning. We evaluate several leading VLMs, including Qwen-2.5-VL, Video-R1, and VideoChat-R1. Despite their strong performance on real-world benchmarks (e.g., MVBench, MMVU), these models hallucinate or fail to detect physical or logical violations, revealing fundamental weaknesses in visual understanding. Finally, we explore reinforcement learning-based post-training on our negative dataset: fine-tuning improves performance on VideoHallu without degrading results on standard benchmarks, indicating enhanced visual reasoning in VLMs. Our data is available at https://github.com/zli12321/VideoHallu.git. Zongxia Li, Xiyang Wu, Guangyao Shi, Yubin Qin, Hongyang Du 0002, Tianyi Zhou 0001, Dinesh Manocha, Jordan L. Boyd-Graber |
NeurIPS | 4 |
| 2024 | FQP: A Fibonacci Quantization Processor with Multiplication-Free Computing and Topological-Order RoutingabstractWith the continuous advancement of artificial intelligence, neural networks exhibit an escalating parameter size, demanding increased computational power and excessive memory access. Low bit-width quantization emerges as a viable solution to address this challenge. However, conventional low bit-width uniform quantization suffers from a mismatch with the weight and activation data distribution in neural networks, resulting in accuracy degradation. Yang Wang 0089, Yubin Qin, Jiachen Wang 0010, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
DAC | 3 |
| 2024 | MECLA: Memory-Compute-Efficient LLM Accelerator with Scaling Sub-matrix PartitionabstractLarge language models (LLMs) have been showing surprising performance in processing language tasks, bringing a new prevalence to deploy LLM from cloud to edge. However, being a scaling auto-regressive Transformer with a huge parameter amount and generating output one by one, LLM introduces overwhelming memory footprints and computation during its inference, especially from its linear layers. For example, generating 32 output tokens with LLaMA-7B LLM requires 14GB of weight data and performs over 400 billion operations (98% from linear layers), which is far beyond the capability of consumer-level GPU and traditional accelerators. To solve these issues, we propose a memory-compute-efficient LLM accelerator, MECLA, with a parameter-efficient scaling sub-matrix partition method (SSMP). It decomposes large weight matrices into several tiny-scale source sub-matrices (SS) and derived sub-matrices (DS). Each DS can be obtained by scaling the corresponding SS with a scalar. For memory issues, SSMP avoids accessing the full weight matrix but only requires small SS and DS scaling scalars. For computation issues, the proposed MECLA processor fully exploits the intermediate data reuse of matrix multiplication via on-chip matrix regrouping, inner-product multiplication re-association, and outer-product partial sum reuse. Experiments on 20 benchmarks show that MECLA reduces memory access and computation by 83.6% and 72.2%. It achieves an energy efficiency of 7088GOPS/W. Compared to V100 GPU and state-of-the-art Transformer accelerator SpAtten and FACT, MECLA saves 113.14×, 12.99×, and 1.62× higher energy efficiency. Yubin Qin, Yang Wang 0089, Zhiren Zhao, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ISCA | 1 |
| 2024 | Exploiting Similarity Opportunities of Emerging Vision AI Models on Hybrid Bonding ArchitectureabstractWhile extensive research has focused on optimizing performance and efficiency in vision-based AI accelerators, an unexplored phenomenon, Clustering Similarity Effect, presents a significant opportunity for further improvement. This effect reveals that clusters of neighboring data points exhibit similar values, enabling the potential to skip redundant computations.To fully capitalize on the potential of the Clustering Similarity Effect (CSE), this work integrates hybrid bonding DRAM technology. We conduct a comprehensive analysis of the associated design considerations and integration overhead. Leveraging these insights, we propose a novel CSE-aware architecture specifically tailored for hybrid bonding memory. This architecture facilitates similarity detection and adapts to the inherent data characteristics associated with CSE.Compared with state-of-the-art 2D/2.5D AI accelerators, the hybrid bonding baseline demonstrates an average energy efficiency improvement of $2.89 \times \sim 14.28 \times$ and an area efficiency improvement of $2.67 \times \sim 7.68 \times$. Incorporating the similarity optimizations further enhances energy efficiency and area efficiency improvement to $5.69 \times \sim 28.13 \times$ and $3.82 \times \sim 10.98 \times$, respectively. Zhiheng Yue, Huizheng Wang, Jiahao Fang, Jinyi Deng, Guangyang Lu, Fengbin Tu, Yubin Qin, Yang Wang 0089, Chao Li 0009, Huiming Han, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ISCA | 9 |
| 2024 | SOFA: A Compute-Memory Optimized Sparsity Accelerator via Cross-Stage Coordinated TilingabstractBenefiting from the self-attention mechanism, Transformer models have attained impressive contextual comprehension capabilities for lengthy texts. The requirements of high-throughput inference arise as the large language models (LLMs) become increasingly prevalent, which calls for large-scale token parallel processing (LTPP). However, existing dynamic sparse accelerators struggle to effectively handle LTPP, as they solely focus on separate stage optimization, and with most efforts confined to computational enhancements. By re-examining the end-to-end flow of dynamic sparse acceleration, we pinpoint an ever-overlooked opportunity that the LTPP can exploit the intrinsic coordination among stages to avoid excessive memory access and redundant computation. Motivated by our observation, we present SOFA, a cross-stage compute-memory efficient algorithm-hardware co-design, which is tailored to tackle the challenges posed by LTPP of Transformer inference effectively. We first propose a novel leading zero computing paradigm, which predicts attention sparsity by using log-based add-only operations to avoid the significant overhead of prediction. Then, a distributed sorting and a sorted updating FlashAttention mechanism are proposed with cross-stage coordinated tiling principle, which enables fine-grained and lightweight coordination among stages, helping optimize memory access and latency. Further, we propose a SOFA accelerator to support these optimizations efficiently. Extensive experiments on 20 benchmarks show that SOFA achieves$9.5\times$speed up and$71.5\times$higher energy efficiency than Nvidia A100 GPU. Compared to eight SOTA accelerators, SOFA achieves an average$15.8\times$energy efficiency,$10.3\times$area efficiency and$9.3\times$speed up, respectively. Huizheng Wang, Jiahao Fang, Xinru Tang, Zhiheng Yue, Yubin Qin, Sihan Guan, Qinze Yang, Yang Wang 0089, Chao Li 0009, Yang Hu 0001, Shouyi Yin |
MICRO | 6 |
| 2023 | FACT: FFN-Attention Co-optimized Transformer Architecture with Eager Correlation PredictionabstractTransformer model is becoming prevalent in various AI applications with its outstanding performance. However, the high cost of computation and memory footprint make its inference inefficient. We discover that among the three main computation modules in a Transformer model (QKV generation, attention computation, FFN), it is the QKV generation and FFN that contribute to the most power cost. While the attention computation, focused by most previous works, only has decent power share when dealing with extremely long inputs. Therefore, in this paper, we propose FACT, an efficient algorithm-hardware co-design optimizing all three modules of Transformer. We first propose an eager prediction algorithm which predicts the attention matrix before QKV generation. It further detects the unnecessary computation in QKV generation and assigns mixed-precision FFN with the predicted attention, which helps improve the throughput. Further, we propose FACT accelerator to efficiently support eager prediction with three designs. It avoids the large overhead of prediction by using log-based add-only operations for prediction. It eliminates the latency of prediction through an out-of-order scheduler that makes the eager prediction and computation work in full pipeline. It additionally avoids memory access conflict in the mixed-precision FFN with a novel diagonal storage pattern. Experiments on 22 benchmarks show that our FACT improves the throughput of the whole Transformer by 3.59× on the geomean average. It achieves an enviable 47.64× and 278.1× energy saving when computing attention, compared to previous attention-optimization-only SOTA works ELSA and Sanger. Further, FACT achieves an energy efficiency of 4388 GOPS/W performing the whole Transformer layer on average, which is 94.98× higher than Nvidia V100 GPU. Yubin Qin, Yang Wang 0089, Dazheng Deng, Zhiren Zhao, Leibo Liu, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
ISCA | 1 |
| 2022 | SWPU: A 126.04 TFLOPS/W Edge-Device Sparse DNN Training Processor With Dynamic Sub-Structured Weight PruningabstractWhen deploying deep neural networks (DNNs), edge devices training is practical to improve model adaptivity for various user-specific scenarios while avoiding privacy disclosure. However, the training computation is intolerable for edge devices. It inspires sparse DNN training (SDT) into the limelight, which reduces training computation by dynamic weight pruning. Generally, SDT has two strategies based on the pruning granularity: the structured or the unstructured. Unfortunately, both of them suffer from limited training efficiency due to the gap between pruning granularity and hardware implementation. The former is hardware-friendly but has a low pruning ratio, indicating limited computation reduction. The latter has a high pruning ratio, but the unbalanced workload decreases utilization and irregular sparsity distribution causes considerable sparsity processing overhead. This paper proposes a software-hardware co- design to bridge the gap for improving the efficiency of SDT. On the algorithm side, a sub-structured pruning method, achieved with hybrid shape-wise and line-wise pruning, generates a high sparsity ratio and keeps the hardware-friendly property. On the hardware side, a sub-structured weight processing unit (SWPU) effectively handles the hybrid sparsity with three techniques. First, SWPU dynamically reorders the computation sequence with hamming-distance-based clustering, balancing the irregular workload. Second, SWPU performs runtime scheduling by exploiting the feature of sub-structured sparse convolution through a detect-before-load controller, which skips redundant memory access and sparsity processing. Third, SWPU performs sparse convolution by compressing operands with spatial disconnect log-based routing and recovers their location with bi-directional switching, avoiding the power-consumed routing logic. Synthesized with 28nm CMOS technology, SWPU can enable 0.56V-to-1.0V supply voltage with a maximum frequency of 675 MHz. It achieves a 50.1% higher pruning ratio than structured pruning and$1.53\times $higher energy efficiency than unstructured pruning. The peak energy efficiency of SWPU is 126.04TFLOPS/W, outperforming the state-of-the-art training processor by$1.67\times $. When training a ResNet-18 model, SWPU reduces$3.72\times $energy and offers$4.69\times $speedup than previous sparse training processors. Yang Wang 0089, Yubin Qin, Leibo Liu, Shaojun Wei, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | BR-CIM: An Efficient Binary Representation Computation-In-Memory DesignabstractDeep neural network (DNN) has recently attracted tremendous attention in various fields. But the computing operation requirement and the memory bottleneck limit the energy efficiency of hardware implementations. Binary quantization is proposed to relieve the pressure of hardware design. And the Computing-In-Memory (CIM) is regarded as a promising method to resolve the memory wall challenge. However, the binary computing paradigm is mismatched with the CIM scheme, which incurs complex circuits and peripheral to realize binary operation in previous works. To overcome previous issues, this work presents Binary Representation Computation-In-Memory (BR-CIM) with several key features. (1) A lightweight computation unit is realized within the 6T SRAM array to accelerate binary computing and enlarge signal margin; (2) The reconfigurable computing scheme and mapping method support extendable bit precision to satisfy the accuracy requirement of various applications; (3) Simultaneous computing and weight loading is supported by column circuitry, which shortens the data loading latency; Several experiments are conducted to estimate algorithm accuracy, the computing latency, and power consumption. The energy efficiency reaches up to 1280 TOPs/W for binary representation. And the algorithm accuracy achieves 97.82%/76.4% on MNIST/CIFAR-100 dataset. Zhiheng Yue, Yubin Qin, Leibo Liu, Shaojun Wei, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |