VLDB 2026 Research / reviewers in the wild / expert
Yandong He
dblp:236/6751
· DBLP profile ↗
14ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Power-efficient 5x Compressive Sensing Readout IC for 3D-stacked CMOS Image Sensor
Jiajia Cui, Kwok Cheong Li, Yandong He, Linxiao Shen |
ISCAS | 4 |
| 2026 | An Area-Efficient Noise-Shaping SAR ADC Utilizing Dynamic Common-Gate Amplifier With Charge-Boosted Amplification and Dynamic-Bulk-SwitchingabstractThis paper presents an area-efficient$2{^{\text {nd}}}$-order noise-shaping (NS) SAR ADC leveraging a dynamic common-gate amplifier. By reconfiguring a simple switch into a dynamic common-gate (DCG) amplifier through adjusting the pulse height applied to a transistor’s gate, voltage gain is achieved prior to the passive loop filter, thereby enhancing noise transfer function (NTF) while maintaining area- and power-efficient loop filtering. To implement a$2{^{\text {nd}}}$-order loop filter, two key techniques are proposed: First, a charge-boosted amplification scheme doubles the charges transferred to the residue capacitors, enabling realization of$2{^{\text {nd}}}$-order noise shaping; Second, a dynamic bulk-switching mechanism triggers a second charge transfer by switching the amplifier’s bulk, eliminating the need for additional amplification stages. Furthermore, to ensure robust noise shaping across process-voltage-temperature (PVT) variations, a PVT-tracking pulse generator is introduced to maintain stable amplifier oper ation. With these techniques, the prototype ADC achieves a 76-dB signal-to-noise-and-distortion ratio (SNDR) with a compact active area of 0.0045 mm2. Operating at 5MS/s sample rate with a 312.5-kHz bandwidth, it consumes 31.5uW, yielding a Schreier Figure-of-Merit (FoM) of 176 dB and a Walden FoM of 9.8fJ/conv.step. Jiajia Cui, Jihang Gao, Xinhang Xu, Yandong He, Linxiao Shen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | Tasa: Thermal-aware 3D-Stacked Architecture Design with Bandwidth Sharing for LLM InferenceabstractThe autoregressive decoding in LLMs is the major inference bottleneck due to the memory-intensive operations and limited hardware bandwidth. 3D-stacked architecture is a promising solution with significantly improved memory bandwidth, which vertically stacked multi DRAM dies on top of logic die. However, our experiments also show the 3D-stacked architecture faces severer thermal issues compared to 2D architecture, in terms of thermal temperature, gradient and scalability. To better exploit the potential of 3D-stacked architecture, we present Tasa, a heterogeneous architecture with cross-stack thermal optimizations to balance the temperature distribution and maximize the performance under the thermal constraints. High-performance core is designed for compute-intensive operations, while high-efficiency core is used for memory-intensive operators, e.g. attention layers. Furthermore, we propose a bandwidth sharing scheduling to improve the bandwidth utilization in such heterogeneous architecture. Extensive thermal experiments show that our Tasa architecture demonstrates greater scalability compared with the homogeneous 3D-stacked architecture, i.e. up to 5.55 °C, 9.37 °C, and 7.91 °C peak temperature reduction for 48, 60, and 72 core configurations. Our experimental for Llama-65B and GPT-3 66B inferences also demonstrate 2.85× and 2.21× speedup are obtained over the GPU baselines and state-of-the-art heterogeneous PIM-based LLM accelerator. Peiran Yan, Yandong He, Youwei Zhuo |
ICCAD | 3 |
| 2025 | LP-Spec: Leveraging LPDDR PIM for Efficient LLM Mobile Speculative Inference with Architecture-Dataflow Co-OptimizationabstractLLM inference on mobile devices faces extraneous challenges due to limited memory bandwidth and computational resources. To address these issues, speculative inference and processing-in-memory (PIM) techniques have been explored at the algorithmic and hardware levels. However, speculative inference results in more compute-intensive GEMM operations, creating new design trade-offs for existing GEMV-accelerated PIM architectures. Furthermore, there exists a significant amount of redundant draft tokens in tree-based speculative inference, necessitating efficient token management schemes to minimize energy consumption. In this work, we present LP-Spec, an architecture-dataflow co-design leveraging hybrid LPDDR5 performance-enhanced PIM architecture with draft token pruning and dynamic workload scheduling to accelerate LLM speculative inference. A near-data memory controller is proposed to enable data reallocation between DRAM and PIM banks. Furthermore, a data allocation unit based on the hardware-aware draft token pruner is developed to minimize energy consumption and fully exploit parallel execution opportunities. Compared to end-to-end LLM inference on other mobile solutions such as mobile NPUs or GEMV-accelerated PIMs, our LP-Spec achieves 13.21× , 7.56 ×, and 99.87× improvements in performance, energy efficiency, and energy-delay-product (EDP). Compared with prior AttAcc PIM and RTX 3090 GPU, LP-Spec can obtain 12.83× and 415.31× EDP reduction benefits. Zhantong Zhu, Yandong He |
ICCAD | 3 |
| 2025 | A Bit-Partitioned Floating-Point 6T SRAM Computing-in-Memory Macro Based on Dual-Edge Time-Domain StructureabstractIn the computing-in-memory (CIM) field, floating-point (FP) CIM is afflicted with high computing latency and energy consumption due to the intricate procedures involved in exponent computation and processing. In this work, an 8Kb FP time-domain (TD) static-random-access-memory (SRAM) CIM macro is presented. Fabricated with a 180nm process, this macro exhibits low computational latency and high energy efficiency. A novel FP computing architecture is proposed, which is capable of concurrently executing exponent summation, maximum value finding, difference generation, and mantissa shifting. This architecture effectively reduces the overall delay in exponent computation and processing, thereby enhancing the throughput. Furthermore, a bit-partitioned computing concept and an exponent sparsity scheme are introduced. In this scheme, sparsity judgment is made solely by processing the high 4 bits of the exponent, which significantly reduces power consumption in the remaining exponent computation and processing steps. Additionally, based on the bit-partitioned concept, a dual-edge TD exponent summation and mantissa multiplication-and-accumulation (MAC) circuit is devised. This circuit not only suppresses nonlinear errors during multi-bit computation but also exploits both the rising and falling edges of pulses for computation, thus accelerating the macro’s operation speed. Compared to previous approaches, an extra 24% power reduction is achieved. At a sparsity level of 90%, a normalized energy efficiency of 14.418 TFLOPS/W and a normalized area efficiency of 0.041 TFLOPS/mm2are attained. When this work is applied to the ResNet-18 model with BF16 format for input, weight, and output, the accuracy loss on the CIFAR-100 dataset is merely −0.16%. Chang Xue, Youming Yang 0002, Gang Du, Yuan Wang 0001, Yandong He |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | OLSATM: Online Learning Based State-Aware Task Migration on S-NUCA Many-CoresabstractTask migration maximizes performance while maintaining thermal safety in many-cores systems. Existing techniques exploit offline learning which requires tremendous training data and fixed-cycle migration which causes threads to miss the optimal migration timing. This paper presents Online Learning based State-Aware Task Migration (OLSATM). It pretrains a neural network (NN) with a small set of data and updates the model online to substitute the laborious data collection and model training of offline learning. It is state-aware and detects the timing when migration is needed, overcoming the shortcomings of periodical migration. Experimental results show that OLSATM enhances the performance by 3.7% and reduces the number of migration judgments by 26 % on average compared to the state-of-the-art task migration. Yandong He, Guangda Zhang, Hengzhu Liu, Renzhi Chen |
ICCD | 1 |
| 2024 | A Dual-Mode CMOS Image Sensor Based on in-Pixel Frame DifferencingabstractIn this article, we introduce a dual-mode CMOS image sensor designed for pixel-level motion detection. A compact computational pixel structure is designed with a pitch of 8 μm and a fill factor of 22.1%. We have implemented a current-mode motion detection module to generate motion signals and control the operation of the column-level analog-todigital converter (ADC) to optimize power efficiency. The data rate can be adjusted by tuning the threshold currents. A 128 × 128 image sensor is fabricated in a 0.18-μm CMOS process. Test results indicate that the imager can output full image at 60 frames per second (fps). Of significant note are the substantial power consumption reductions achieved in both of the two motion detection modes, amounting to 30.7% and 60%, respectively. In data compression mode, a 3.02 pJ/pixel/frame Figure of Merit (FoM) is achieved. Xu Ren, Liqiao Liu, Yandong He, Gang Du |
ISCAS | 3 |
| 2023 | Workload-Aware Cache Replacement Policy Based on Bayesian InferenceabstractThe replacement policy contributes to enhancing the cache hit ratio, affecting the performance of the processor indirectly. Prior proposed static replacement policies are limited to certain classes of workload types, failing to achieve high hit rates in various benchmarks. In this work, we propose a self-adaptive replacement algorithm. The algorithm detects the drop points of the hit ratio online based on Bayesian inference and selects a new policy from a policy pool containing multiple policies according to its weight at the change points. Besides, the algorithm updates the weights based on the performance of the selected policy. Compared to the static replacement policy, our algorithm is able to apply to more access patterns and achieve a high hit ratio. We choose 15 benchmarks from DPC3 and concatenate them to generate a total of 13 composite benchmarks. We run our algorithm using a 2MB last-level cache (LLC) and show that our algorithm improves the hit rate by 2.6% over LRU and 10.6% over Random. Yandong He, Zhong Wan, Renzhi Chen |
SMC | 1 |
| 2022 | Reliability-Improved Read Circuit and Self-Terminating Write Circuit for STT-MRAM in 16 nm FinFETabstractHigh power consumption is usually required in a spin-torque-transfer magnetoresistive random access memory (STT-MRAM) array’s peripheral circuits for reliable operations. In read, power needs to be spent for the low absolute resistance in the magnetic tunnel junctions (MTJ), and a limited high-state-to-low-state resistance ratio calls for high currents for the same detectable readout voltage under accuracy requirements. In write, the random programming time poses challenges for energy efficient write operations within an acceptable write error rate. To address the issues mentioned above, in this work, we propose a reliability-improved read circuit that consumes only 92.09 fJ/bit read energy while ensuring correct readout values under 4.5 sigma resistance variance, and a self-terminating write peripheral circuit achieving an energy reduction of 82.3% at 1 part-per-million write error rate (WER) under 20 ns write period. Chang Xue, Yihan Zhang 0002, Mingwei Zhu, Tianqiao Wu, Meng Wu 0005, Yandong He, Le Ye |
ISCAS | 7 |
| 2022 | Proton radiation effects on high-speed silicon Mach-Zehnder modulators for space application
Changhao Han, Zhaoyi Hu, Yuansheng Tao, En-Gang Fu, Yandong He, Fenghe Yang |
Sci. China Inf. Sci. | 5 |
| 2022 | A low-fabrication-temperature, high-gain chip-scale waveguide amplifier
Peiqi Zhou, Yandong He |
Sci. China Inf. Sci. | 4 |
| 2019 | Improved turn-on behavior in a diode-triggered silicon-controlled rectifier for high-speed electrostatic discharge protection
Lizhong Zhang, Yuan Wang 0001, Yize Wang, Xing Zhang 0002, Yandong He |
Sci. China Inf. Sci. | 5 |
| 2013 | A monitoring circuit for NBTI degradation at 65nm technology nodeabstractThe paper introduces a new monitoring circuit to quantify the change in performance of devices undergoing NBTI stress at 65nm technology node. The proposed solution consists of a pMOS device experiencing accelerated NBTI stress, a Capacitive Switch, Relaxation Oscillator structure, serving as a monitoring unit, which allows to dynamically track the NBTI-induced degradation with time. The circuit measures the change of the relaxation frequency which is attributed to the saturation current shift in pMOSFET due to NBTI stress. The proposed circuit with counting unit produces a digital output which makes it easier to collect. The capacitive relaxation oscillator modeling has been established and verified by the experiments. Meanwhile, the circuit is easily to be integrated with the digital logic system. Based on the device matrix, the effect of initial saturation current distribution and NBTI-induced time-dependent variability can also be obtained. Yandong He, Ganggang Zhang, Xing Zhang 0002 |
ISCAS | 1 |
| 2011 | Process optimization of plasma nitridation SiON for 65 nm node gate dielectrics
Yandong He, Yangyuan Wang |
Sci. China Inf. Sci. | 1 |