Siyuan Wei

dblp:126/0868 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Scheduling Training-Inference Co-Location in Demand Response for Sustainable Edge AI
abstract
In the pursuit of data privacy and reduced latency, the adoption of edge intelligence has surged. Meanwhile, the enormous increase in AI has resulted in significant energy consumption. Edge intelligence plays a crucial role in Energy Demand Response (EDR). However, existing edge intelligence falls short of meeting the demands of co-locating training and inference tasks while satisfying EDR. Specifically, the intertwinement between balancing energy consumption, system delay and model accuracy, and uncertain future inputs adds to the challenge of designing an online sustainable system for co-located training and inference tasks. To address these challenges, we propose a novel two-timescale system for co-locating training and inference EDR. Our approach satisfies EDR by strategically planning training schedules on macro-timescales and migrating inference requests between heterogeneous edges on micro-timescales while minimizing long-term cost. We introduce a novel online polynomial time algorithm that first breaks down the problem into two subproblems, which are subsequently solved using an online-learning-based fractional algorithm and a randomized roun ding algorithm, respectively. Rigorous analysis demonstrates that our approach achieves both sublinear dynamic regret and sublinear dynamic fit. Extensive trace-driven evaluations validate the practical superiority of our approach over multiple existing methods, highlighting its effectiveness in real-world scenarios.
Konglin Zhu, Siyuan Wei, Xuan'er Wu, Lei Jiao 0002, Jin Dong 0004, Lin Zhang 0013
IEEE Trans. Serv. Comput.2
2025 Towards Uyghur Lip-Reading: Dataset Development and Attention-Enhanced Recognition with ECA-S Module
Siyuan Wei, Zilong Xing, Mutallip Mamut, Kurban Ubul
PRCV (15)1
2025 A 75.6 Gb/s 22-bit Floating-Point Coarse-Grained Versatile DSP Embedding 26K-Point Baseband Signal Processing and 2048 × 256-Point Complex FFT
abstract
As millimeter-wave radar technology advances, modern domain-specific digital signal processors (DSPs) struggle to balance versatility and processing scale, while suffering from low data precision and data throughput. To address these challenges, this paper presents a novel coarse-grained versatile DSP (CVDSP). The CVDSP introduces an architecture based on multi-level finite state machines and a custom instruction set to support various algorithms through flexible dataflow. The efficient large-scale PEs are designed with 22-bit floating-point precision to handle 26K-point baseband signal processing and$2048\times 256$-point complex fast Fourier transform. Cooperating with a high-bandwidth instruction-free memory access network, the CVDSP achieves relatively high data throughput. Fabricated in a 65-nm CMOS process, experimental results show that the peak energy efficiency and data throughput of the CVDSP are 299.5 GMACs/J and 75.6 Gb/s. The CVDSP demonstrates fully on-chip implementation of the FMCW radar, Pulse-Doppler radar, and spectrometer algorithms.
Xuanzhe Xu, Xianjun Liu, Siyuan Wei, Shuangming Yu, Runjiang Dou, Xu Yang 0017, Jian Liu 0021, Nanjian Wu
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 A Real-Time 2D/3D Perception Visual Vector Processor for 1920 × 1080 High-Resolution High-Speed Intelligent Vision Chips
abstract
Edge computing of reliable multimodal (2D RGB/3D RGB-Depth) data has a wide range of applications. However, many of currently reported visual processors cannot flexibly handle multimodal data, e.g., the visual streams of RGB-Depth data. The key challenge exists that these prior visual processors do not come with efficient and unified instruction set architecture (ISA) for both conventional and intelligent cognition on the 2D/3D multimodal sensory data. To fill such a gap, this paper proposes a programmable intelligent visual vector processor compatible with multimodal 2D/3D visual data processing ($1920\times 1080$-pixel resolution). The processor consists of a reconfigurable processing element (PE) array, a memory access network flexibly configurable to be fine- or coarse-grained, and a high throughput I/O interface. The vectorial PE array with neighbor PE access increases the data reuse rate and parallel computation efficiency, and can implement both convolutional neural networks (CNNs) and conventional image processing algorithms. The proposed ISA is customized and optimally tailored targeting 2D/3D image processing from RGB/Time-of-Flight(ToF) raw data to intelligent inference results. The chip is fabricated in a 55-nm CMOS process. The experimental results showed that the area efficiency, peak performance, and peak throughput of our chip attained as high as 14.41GOPS/mm2, 409.6GOPS, and 9.6Gbps at 200MHz, respectively. The measured processing speeds of this chip on ToF depth reconstruction is 87fps ($480\times 270$) or 31 fps($1920\times 1080$),on 3D object classification is 219fps ($256\times 256$), and on CNN-based 2D object tracking is 36fps ($256\times 256$).
Siyuan Wei, Lei Kang 0006, Xuemin Zheng, Mingxin Zhao, Mengmeng Xu 0005, Xuanzhe Xu, Runjiang Dou, Shuangming Yu, Xu Yang 0017, Jian Liu 0021, Cong Shi 0003, Nanjian Wu
IEEE Trans. Circuits Syst. I Regul. Pap.1
2023 Joint Token Pruning and Squeezing Towards More Aggressive Compression of Vision Transformers
abstract
Although vision transformers (ViTs) have shown promising results in various computer vision tasks recently, their high computational cost limits their practical applications. Previous approaches that prune redundant tokens have demonstrated a good trade-off between performance and computation costs. Nevertheless, errors caused by pruning strategies can lead to significant information loss. Our quantitative experiments reveal that the impact of pruned tokens on performance should be noticeable. To address this issue, we propose a novel joint Token Pruning & Squeezing module (TPS) for compressing vision transformers with higher efficiency. Firstly, TPS adopts pruning to get the reserved and pruned subsets. Secondly, TPS squeezes the information of pruned tokens into partial reserved tokens via the unidirectional nearest-neighbor matching and similarity-based fusing steps. Compared to state-of-the-art methods, our approach outperforms them under all token pruning intensities. Especially while shrinking DeiT-tiny&small computational budgets to 35%, it improves the accuracy by 1%-6% compared with baselines on ImageNet classification. The proposed method can accelerate the throughput of DeiT-small beyond DeiT-tiny, while its accuracy surpasses DeiT-tiny by 4.78%. Experiments on various transformers demonstrate the effectiveness of our method, while analysis experiments prove our higher robustness to the errors of the token pruning policy. Code is available at https://github.com/megvii-research/TPS-CVPR2023.
Siyuan Wei, Tianzhu Ye, Jiajun Liang
CVPR1
2021 Front-End Architecture Design for Low-Complexity 3-D Ultrasound Imaging Based on Synthetic Aperture Sequential Beamforming
abstract
The 3-D ultrasound imaging provides distinct advantages over its 2-D counterpart leading to a more accurate analysis of tumors and cysts. However, the front end of a 3-D system must receive and process data at prodigious rates, making it impractical for power-constrained portable systems. Synthetic aperture sequential beamforming (SASB) is an ultrasound beamforming technique that splits the computation into two stages, such that the computation in Stage 1 can be completed in the power-constrained front end while the remaining computation can be done elsewhere. In this article, we present several algorithmic and architectural techniques to enable efficient computation of Stage 1 processing without compromising imaging quality. Specifically, we present algorithmic techniques that reduce the computational complexity in Stage 1 by 17× through a systematic reduction in the number of apodization coefficients. We propose a 3-D die stacked architecture where the signals received by 961 active transducers are digitized, routed by a network-onchip, and processed in parallel. This architecture does not require the explicit storage of incoming data samples. We synthesize the architecture using TSMC 28-nm technology node. The front-end power consumption is around 1.5 W, making it suitable for portable applications.
Jian Zhou 0012, Sumit K. Mandal, Brendan L. West, Siyuan Wei, Ümit Y. Ogras, Oliver Kripfgans, J. Brian Fowlkes, Thomas F. Wenisch, Chaitali Chakrabarti
IEEE Trans. Very Large Scale Integr. Syst.4
2013 Sonic Millip3De: A massively parallel 3D-stacked accelerator for 3D ultrasound
abstract
Three-dimensional (3D) ultrasound is becoming common for non-invasive medical imaging because of its high accuracy, safety, and ease of use. Unlike other modalities, ultrasound transducers require little power, which makes hand-held imaging platforms possible, and several low-resolution 2D devices are commercially available today. However, the extreme computational requirements (and associated power requirements) of 3D ultrasound image formation has, to date, precluded hand-held 3D capable devices. We describe the Sonic Millip3De, a new system architecture and accelerator for 3D ultrasound beamformation-the most computationally intensive aspect of image formation. Our three-layer die-stacked design features a custom beamsum accelerator that employs massive data parallelism and a streaming transform-select-reduce pipeline architecture enabled by our new iterative beamsum delay calculation algorithm. Based on RTL-level design and floorplanning for an industrial 45nm process, we show Sonic Millip3De can enable 3D ultrasound with a fully sampled 128×96 transducer array within a 16W full-system power budget (400× less than a conventional DSP solution) and will meet a 5W safe power target by the 11nm node.
Richard Sampson, Ming Yang 0004, Siyuan Wei, Chaitali Chakrabarti, Thomas F. Wenisch
HPCA3