EDBT 2026 Demo / reviewers in the wild / expert
Heng You
dblp:231/8288
· DBLP profile ↗
10ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0002-9386-8030ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 5 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Approximate Digital Compute-In-Memory Macro with Reconfigurable Computational Precision for Neural Network Acceleration
Heng You, Zixiao Zhan, Guanghua Zhao, Shushan Qiao, Jia Yuan |
ISCAS | 1 |
| 2026 | A Booth-Based Digital Compute-in-Memory Macro With Bitwise-Efficient Multiphase AccumulationabstractThis brief presents a Booth-based all-digital SRAM compute-in-memory (CIM) macro designed for high-efficiency multiply-and-accumulate (MAC) operations in artificial intelligence applications. The architecture incorporates three key innovations: 1) a pre-encoding (PENC) scheme utilizing a Booth-based consolidated truth table, designed to reduce encoder complexity and computing unit overhead; 2) a minimalist bitwise processing technique, employing only two transistors for bit retention and shifting, aimed to minimize computational complexity within the subcomputing array; and 3) a multiphase product summation approach, integrating unified negative compensation with global sign bit expansion, intended to lower power consumption during partial product accumulation and final addition. The proposed architecture is readily adaptable to multiplications across diverse bit widths and accommodates flexible CIM array sizes, ensuring high scalability. Measurement results from a 55-nm-CMOS process demonstrate that the proposed CIM macro achieves an energy efficiency of 16.8 TOPS/W under 8b/8b MAC operations. Yiying Jiang, Heng You, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | A Charge Domain SRAM Computing-in-Memory Macro With Quantized Interval-Optimized ADC and Input Bit-Level Sparsity-Optimized P2O-DAC for 8-b MAC OperationabstractComputing-in-memory (CIM) has recently gained significant attention as it achieves high energy efficiency and throughput for deep convolutional neural networks (DCNNs). In this brief, we present a static random access memory (SRAM) CIM macro aimed at improving the energy efficiency of edge devices when performing 8-b multiply-and-accumulate (MAC) operations. The proposed architecture implements the following: 1) a successive approximation register analog-to-digital converter (SAR ADC) readout circuit based on a weight-flip-store (WFS) coding scheme, where energy efficiency is improved by optimizing the quantized interval; 2) an input-relevant partial power-off digital-to-analog converter (P2O-DAC) using input bit-level sparsity to reduce power consumption; and 3) a pipeline structure for interleaving MAC computation and readout operation to minimize the redundancy when loading input data into the CIM array. Our proposed CIM macro is implemented in TSMC 40-nm CMOS technology. Postlayout simulation results show an average macro energy efficiency of 16.8 TOPS/W without input and weight value sparsity. Shukao Dou, Zupei Gu, Heng You, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | A Transfer Approach Using Graph Neural Networks in Deep Reinforcement LearningabstractTransfer learning (TL) has shown great potential to improve Reinforcement Learning (RL) efficiency by leveraging prior knowledge in new tasks. However, much of the existing TL research focuses on transferring knowledge between tasks that share the same state-action spaces. Further, transfer from multiple source tasks that have different state-action spaces is more challenging and needs to be solved urgently to improve the generalization and practicality of the method in real-world scenarios. This paper proposes TURRET (Transfer Using gRaph neuRal nETworks), to utilize the generalization capabilities of Graph Neural Networks (GNNs) to facilitate efficient and effective multi-source policy transfer learning in the state-action mismatch setting. TURRET learns a semantic representation by accounting for the intrinsic property of the agent through GNNs, which leads to a unified state embedding space for all tasks. As a result, TURRET achieves more efficient transfer with strong generalization ability between different tasks and can be easily combined with existing Deep RL algorithms. Experimental results show that TURRET significantly outperforms other TL methods on multiple continuous action control tasks, successfully transferring across robots with different state-action spaces. Tianpei Yang, Heng You, Jianye Hao, Yan Zheng 0002, Matthew E. Taylor |
AAAI | 2 |
| 2024 | A 409mV, Sub-10nW Power-on Reset Circuit Using Adaptive Accuracy Adjustment for Low Voltage ApplicationsabstractA low power power-on reset (POR) circuit with low temperature coefficient, low quiescent current and small area is proposed in this paper. The POR circuit samples the power supply voltage through a high threshold transistor and then converts the voltage to a current, which will be compared with a native NMOS based current reference to obtain the reset signal. In order to reduce the power consumption in steady state, an adaptive accuracy adjustment mechanism is employed in the proposed POR circuit. The POR circuit uses a low accuracy but energy efficient structure to monitor the supply voltage in steady state, and when a voltage drop is observed, the POR circuit quickly switches to a high accuracy mode to get an accurate brown-out detection trip voltage. The POR circuit is implemented in 55nm CMOS process and the active area is just 107μm2. Post-layout simulation results show that the POR circuit has a POR trip voltage of 409mV, a static power of 9.11nW at a supply voltage of 0.45V. Besides, the temperature coefficient of the proposed POR circuit is only 31.76μV/°C over a temperature range of -40°C to 125°C. Heng You, Dashan Shi, Delong Shang, Shushan Qiao |
ISCAS | 1 |
| 2024 | A 1-8b Reconfigurable Digital SRAM Compute-in-Memory Macro for Processing Neural NetworksabstractThis work presents a 1-8b reconfigurable digital SRAM compute-in-memory (CIM) macro, which significantly improves array utilization and energy efficiency under different input and weight configurations compared to previous works. To ensure the array utilization under different configurations, a row-based bitwise-summation-first digital CIM architecture is proposed. In addition, to realize flexible switching between signed and unsigned operations, a complete 2’s complement encoding method is adopted, which makes the computation of the sign bits consistent with that of the magnitude bits when performing signed operations, thus ensuring that each row of the CIM array can store the sign of the weight. Due to the support of reconfigurable bit width, the proposed CIM macro can be widely used in various neural networks for optimal efficiency. In order to better apply the CIM macro to binarized neural networks, a configurable bitwise multiplier is presented, which supports both AND and XNOR operations. Moreover, since the power consumption of the adder tree occupies a major part of the digital CIM macro, a 4–2 compressor based adder tree is presented to further improve the energy efficiency. Measurement results based on 55nm CMOS process show that the proposed CIM macro achieves an energy efficiency of up to 2238TOPS/W at 1b/1b and 44.82TOPS/W at 4b/4b MAC operations. Heng You, Delong Shang, Shushan Qiao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2023 | Transfer Reinforcement Learning Based Negotiating Agent Framework
Siqi Chen 0001, Tianpei Yang, Heng You, Jianye Hao, Gerhard Weiss 0001 |
PAKDD (2) | 3 |
| 2022 | Cross-domain adaptive transfer reinforcement learning based on state-action correspondenceabstractDespite the impressive success achieved in various domains, deep reinforcement learning (DRL) is still faced with the sample inefficiency problem. Transfer learning (TL), which leverages prior knowledge from different but related tasks to accelerate the target task learning, has emerged as a promising direction to improve RL efficiency. The majority of prior work considers TL across tasks with the same state-action spaces, while transferring across domains with different state-action spaces is relatively unexplored. Furthermore, such existing cross-domain transfer approaches only enable transfer from a single source policy, leaving open the important question of how to best transfer from multiple source policies. This paper proposes a novel framework called Cross-domain Adaptive Transfer (CAT) to accelerate DRL. CAT learns the state-action correspondence from each source task to the target task and adaptively transfers knowledge from multiple source task policies to the target policy. CAT can be easily combined with existing DRL algorithms and experimental results show that CAT significantly accelerates learning and outperforms other cross-domain transfer methods on multiple continuous action control tasks. Heng You, Tianpei Yang, Yan Zheng 0002, Jianye Hao, Matthew E. Taylor |
UAI | 1 |
| 2021 | A 0.5V 36nW 10-Transistor Power-on-Reset Circuit with High AccuracyabstractIn this paper, a low voltage high accuracy 10- transistor power-on-reset circuit with brown-out-reset function is proposed. A native NMOS current reference based architecture is proposed to get high accuracy trip-voltage with a small area and power consumption. By adjusting the number of native NMOS transistors, a stable hysteresis window is obtained. Post-layout simulation results based on SMIC 55nm CMOS process show that the trip-voltage deviation of the proposed power-on-reset circuit is only 34mV under different temperature and process corners. Also, the trip-voltage of the proposed power-on-reset circuit shows great robustness to supply ramp time. The power consumption of the proposed circuit is as low as 36nW at 0.5V. Since the proposed power-on-reset circuit consists of only 10 transistors, the area is as low as 67.5μm2. Heng You, Jia Yuan, Zenghui Yu, Shushan Qiao |
ISCAS | 1 |
| 2021 | Low-Power Retentive True Single-Phase-Clocked Flip-Flop With Redundant-Precharge-Free OperationabstractAs basic components, optimizing power consumption of flip-flops (FFs) can significantly reduce the power of digital systems. In this article, an energy-efficient retentive true-single-phase-clocked (TSPC) FF is proposed. With the employment of input-aware precharge scheme, the proposed TSPC FF precharges only when necessary. In addition, floating node analysis and transistor level optimization are employed to further ensure the high energy efficiency of the FF without significantly increasing the area. Postlayout simulations based on SMIC 55-nm CMOS technology show that at a supply voltage of 1.2 V, the power consumption of the proposed FF is 84.37% lower than that of conventional transmission-gate flip-flop (TGFF) at 10% data activity. The reduction rate is increased to 98.53% as the data activity goes down to 0%. When the supply voltage decreases to 0.6 V, the proposed FF consumes only 0.411 fJ/cycle at 10% data activity, which is 84.23% lower than TGFF. Measurement results of ten test chips demonstrate the great energy efficiency of the proposed FF. Furthermore, the CK-to-Q delay of the proposed FF is 26.18% lower than that of TGFF at a supply voltage of 1.2 V. Heng You, Jia Yuan, Zenghui Yu, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |