Haoxuan Shan

dblp:310/3711 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
8since 2021 · last 2026
0009-0000-9671-6713ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Platinum: Path-Adaptable LUT-Based Accelerator Tailored for Low-Bit Weight Matrix Multiplication
abstract
The rapid scaling of large language models demands more efficient hardware. Quantization offers a promising trade-off between efficiency and performance. With ultra-low-bit quantization, there are abundant opportunities for results reuse, and thus it can be boosted with lookup tables (LUTs) based acceleration. However, existing LUT-based methods suffer from computation and hardware overheads for LUT construction, and rely solely on bit-serial computation, which is suboptimal for ternary-weight networks. We propose Platinum, a lightweight ASIC accelerator for integer weight mixed-precision matrix multiplication (mpGEMM) using LUTs. Platinum reduces LUT construction overhead via offline-generated construction paths and supports both general bit-serial and optimized ternaryweight execution through adaptive path switching. On BitNet b1.58-3B, Platinum achieves up to $73.6 \times, 4.09 \times$, and $2.15 \times$ speedups over SpikingEyeriss, Prosperity, and 16-thread T-MAC (CPU), respectively, along with energy reductions of $32.4 \times, 3.23 \times$, and $20.9 \times$, all within a $0.96 \mathrm{~mm}^{2}$ chip area. This demonstrates the potential of LUT-based ASICs as efficient, scalable solutions for ultra-low-bit neural networks on edge platforms.
Haoxuan Shan, Cong Guo 0003, Chiyue Wei, Junyao Zhang 0003, Hai Li 0001, Yiran Chen 0001
ASP-DAC1
2026 Frame Skipping Architecture for Video-Language Model Acceleration
Haoxuan Shan, Chiyue Wei, Cong Guo 0003, Yuzhe Fu, Hai Li 0001, Yiran Chen 0001
ACM Great Lakes Symposium on VLSI1
2026 Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
abstract
Vision-Language Models (VLMs) have demonstrated strong performance on tasks such as video captioning and visual question answering. However, their growing scale and video-level inputs lead to significant computational and memory overhead, posing challenges for real-time deployment on hardware accelerators. While prior work attempts to reduce redundancy via token pruning or merging, these methods typically operate at coarse granularity and incur high runtime overhead due to global token-level operations. In this study, we propose Focus, a Streaming Concentration Architecture that efficiently accelerates VLM inference through progressive, fine-grained redundancy elimination. Focus introduces a multilevel concentration paradigm that hierarchically compresses vision-language inputs at three levels: (1) semantic-guided token pruning based on textual prompts, (2) spatial-temporal blocklevel concentration using localized comparisons, and (3) vectorlevel redundancy removal via motion-aware matching. All concentration steps are tightly co-designed with the architecture to support streaming-friendly, on-chip execution. Focus leverages GEMM tiling, convolution-style layout, and cross-modal attention to minimize off-chip access while enabling high throughput. Implemented as a modular unit within a systolic-array accelerator, Focus achieves$2.4 \times$speedup and$3.3 \times$reduction in energy, significantly outperforming state-of-the-art accelerator in both performance and energy efficiency. Full-stack implementation of Focus is open-sourced at https://github.com/dubcyfor3/Focus.
Chiyue Wei, Cong Guo 0003, Junyao Zhang 0003, Haoxuan Shan, Qinsi Wang, Changchun Zhou 0001, Hai Li 0001, Yiran Chen 0001
HPCA4
2026 EVA: Accelerating LLM Decoding via an Efficient Vector Quantization Architecture
Bowen Duan 0003, Cong Guo 0003, Chiyue Wei, Haoxuan Shan, Yuzhe Fu, Xinhua Chen, Changchun Zhou 0001, Hai Li 0001, Yiran Chen 0001
ISCA4
2026 Research on Precise User Intent Recognition Algorithm for Power Grid Enterprise Trade Union Domain Based on Deep Residual Attention Network
abstract
ABSTRACT With the deep development of digital transformation in power grid enterprise trade unions, accurate identification of employee consultation intent has become a key technology for improving trade‐union service quality. To address the limitations of existing intent recognition methods in the trade union domain regarding domain adaptability, short‐text discrimination and statistically robust deployment validation, this paper presents a domain‐adaptive intent recognition solution that couples a purpose‐built trade‐union corpus with a Deep Residual Attention Fusion Network (DRAFNet). The corpus comprises 12,486 annotated consultations across seven intent categories, constructed through dual‐independent annotation plus expert adjudication with an overall Cohen's of 0.885 and a 30‐day‐later blind verification of 0.901 on a stratified random subset, and is processed under a leakage‐free split‐before‐augment data protocol. DRAFNet integrates domain‐lexicon‐enhanced encoding, a BERT backbone, four convolutional residual blocks with skip connections, and a gated fusion of local window attention and global multi‐head self‐attention, jointly optimized via cross‐entropy and focal loss. Over five independent runs on the self‐constructed test set, DRAFNet attains 94.52% 0.18% accuracy and 93.88% 0.21% Macro‐F1, statistically outperforming the strongest pretrained baseline ERNIE 3.0 by 2.21 and 2.25 percentage points respectively at < 0.001 in paired ‐tests. Cross‐corpus generalization is further verified on the public SMP2017‐ECDT Chinese intent benchmark, and a side‐by‐side comparison with three frontier large language models (GPT‐4o, Gemini 1.5 Pro, Claude 3.5 Sonnet) under zero‐shot and five‐shot prompting shows that DRAFNet outperforms the strongest LLM configuration by 6.36 percentage points in accuracy at over two orders of magnitude lower per‐sample latency. Ablation experiments verify the contribution of each module, and a distilled‐and‐quantized variant retains 92.15% Macro‐F1 while reducing storage footprint by 86% and CPU latency by 68%, demonstrating a favourable accuracy‐efficiency trade‐off for industrial deployment. The proposed integration provides a statistically robust and reproducibly evaluated solution for intent recognition in vertical service domains.
Caihua Song, Zhaoxiong Guan, Haoxuan Shan
Expert Syst. J. Knowl. Eng.4
2025 Designing and Training Neural Networks for Analog In-Sensor Deployment: A Hardware-Aware Analysis
abstract
Edge AI and IoT applications demand ultra-low latency and energy efficiency, but these goals are often undermined by the costs of digitizing and transmitting data. Analog in-sensor (AIS) hardware architectures address this bottleneck by enabling analog processing directly within the sensor, minimizing digitization and data movement. However, AIS deployments face key challenges including stringent power, performance, and area constraints, susceptibility to hardware-induced noise and variations, and accuracy degradation from operating on unprocessed sensor outputs rather than refined image data. We address these challenges through a software-driven, hardware-aware analysis that distills actionable design guidance for AIS-optimized convolutional neural networks (CNNs). Drawing on prior literature and our own empirical studies, we derive design recommendations for AIS-friendly network topologies, training recipes that jointly improve noise and quantization robustness, and strategies for effective learning from emulated raw sensor data without a digital image signal processing (ISP) pipeline. This analysis provides insight into hardware-aware software-based techniques that complement cutting-edge circuit and architecture-level approaches, helping advance the limits of high-performance AIS systems.
Mark Horton, Haoxuan Shan, James Kiessling, Huanrui Yang, Yiran Chen 0001, Hai Li 0001
ICCAD2
2024 Semantic-Aware Synthesis Network for Dental Caries Image Generation from CBCT to Micro-CT
abstract
Dental caries is a common oral disease, and accurate imaging is vital for diagnosing, especially for assessing carious lesions and the pulp area. Cone-Beam Computed Tomography (CBCT) is a widely used imaging technology for the clinical diagnosis of caries, but its low resolution and blurred boundaries hinder clear visualization of the pulp and lesion extent. Micro Computed Tomography (Micro-CT) provides higher resolution and precision, but it’s only used for ex vivo imaging. Recently, image synthesis techniques have been extensively applied to enhance the quality of medical images. However, existing synthesis methods often focus on pixel-level correspondence, neglecting the semantic information within the images. This results in limitations in reconstructing structural edges and accurately simulating anatomical regions. To address these challenges, we propose a novel model that generates Micro-CT images from CBCT images. The proposed Semantic-Aware Synthesis Network (SASN) employs a multi-task learning strategy, integrating a segmentation task to enhance the information learning capability of the shared encoder, thereby facilitating synthesis. To achieve better semantic-based synthesis, we employ a Semantic-Guided Attention Module (SGAM) to facilitate feature fusion between branches. Additionally, we introduce a semantic alignment loss to ensure semantic consistency, thereby further enhancing network performance. Experimental results demonstrate that SASN outperforms other existing methods, achieving 27.05 (± 0.18) in PSNR, 84.41[%] (± 0.97) in SSIM, and 45.47 [×1e-3] (± 1.37) in RMSE.
Haoxuan Shan, Wei Liu 0303, Shuai Qi
BIBM1
2024 ModSRAM: Algorithm-Hardware Co-Design for Large Number Modular Multiplication in SRAM
abstract
Elliptic curve cryptography (ECC) is widely used in security applications such as public key cryptography (PKC) and zero-knowledge proofs (ZKP). ECC is composed of modular arithmetic, where modular multiplication takes most of the processing time. Computational complexity and memory constraints of ECC limit the performance. Therefore, hardware acceleration on ECC is an active field of research. Processing-in-memory (PIM) is a promising approach to tackle this problem. In this work, we design ModSRAM, the first 8T SRAM PIM architecture to compute large-number modular multiplication efficiently. In addition, we propose R4CSA-LUT, a new algorithm that reduces the cycles for an interleaved algorithm and eliminates carry propagation for addition based on look-up tables (LUT). ModSRAM is co-designed with R4CSA-LUT to support modular multiplication and data reuse in memory with 52% cycle reduction compared to prior works with only 32% area overhead.
Jonathan Hao-Cheng Ku, Junyao Zhang 0003, Haoxuan Shan, Saichand Samudrala, Jiawen Wu 0006, Qilin Zheng, Ziru Li, Jeyavijayan Rajendran, Yiran Chen 0001
DAC3