EDBT 2026 Demo / reviewers in the wild / expert
Deming Zhang
dblp:29/8053
· DBLP profile ↗
19ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0001-7261-371XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 15 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3Artificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 74.5-dB DR 1.25-MHz BW Continuous-Time Delta-Sigma ADC Using Low-Noise Feedback DAC with Optimized Barrel Shifter
Zhigang Fu, Renjie Fu, Linxu Shi, Deming Zhang |
ISCAS | 5 |
| 2025 | A High-Speed, Low-Power, High-Reliability and Fully Single Event Double Node Upset Tolerant Design for Magnetic Random Access MemoryabstractMagnetic Random Access Memory (MRAM) has enormous application potential in the aerospace field due to its nonvolatile, high speed, low power, and inherent radiation resistance characteristics. Due to its high sensing reliability, pre-charge differential sense amplifier (PCDSA) has been proposed and widely used in MRAM products. However, such PCDSA is based on traditional CMOS technology, and as the size of CMOS technology continues to shrink, its sensing result is easily affected by single event upset (SEU) or even the single event double node upset (SEDU). Recently, a TSC-PCDSA has been proposed to fully tolerate SEDU. However, it still suffers from slow speed, high power consumption and low reliability during normal sense operation. To address these issues, this paper proposes a novel PCDSA circuit that uses 6 three-input approximate C-elements (TACs) and 2 three-input standard C-elements (TSCs) to provide SEDU-tolerance. By reducing the number of transistors on the discharge path and increasing the difference in discharge current, the proposed PCDSA can achieve high speed, low power and high reliability. By using a physics-based STT-MTJ compact model and a commercial CMOS 40 nm design kit, hybrid simulations have been performed to demonstrate its functionality and evaluate its performance. Simulation results show that when the TMR is 150%, the width of N1-N12 is 480 nm and the$\text {V}_{\text {DD}}$is 1.1 V, the proposed PCDSA sensing error rate (SER) is close to 0% during normal sense operation, achieving a high sense speed of 123.6 ps and a low sense energy of 1.6533 fJ. Compared with the previously proposed TSC-PCDSA, the sense reliability is greatly improved, and the sense time and sense energy are reduced by 1.84 times and 1.27 times, respectively. Moreover, the proposed PCDSA can fully tolerate SEDU by optimizing the layout design. In the worst case where deposited charge$Q_{\text {inj}}$is 2 pC, it can achieve a shorter recover time of 1.28244 ns and a lower recover energy dissipation of 2.1604 pJ than the previously proposed TSC-PCDSA. Shixuan Wang, Yue Zhang 0010, Weisheng Zhao 0001, Lang Zeng, Deming Zhang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | A 32 kb 55 nm Radiation-Hardened SRAM Chip With SEU ≤1.1 E-11 Upsets/Bit-Day, SEL >107.1 MeV ⋅ cm²/mg, and TID >100 Krad(Si) for Space ApplicationsabstractIn this paper, a 32kb radiation-hardened (RH) static random access memory (SRAM) chip, named BH55RHSRAM32K, is proposed and fabricated for space applications. The chip is hardened from the view of the circuit level, layout level, and system level and is fabricated using a 55 nm CMOS process design kit with an RH cell library. At the circuit level, the proposed RH-14T SRAM cell and radiation-hardened pre-charged sense amplifier (RH-PCSA) cell adopt a polarity hardening method, making them fully tolerant of single event upset (SEU). At the layout level, the sensitive nodes in the proposed RH-14T SRAM cell and RH-PCSA cell layouts are isolated. Furthermore, the proposed RH-14T SRAM array adopts a bit-interleaved design, effectively reducing single event double upsets (SEDU). At the system level, an error correction coding (ECC) circuit is implemented to enhance SEU tolerance. Experimental results show that the proposed 32kb RH-SRAM chip can not only obtains superior radiation tolerance, i.e., the SEU ≤ 1.1E-11 upsets/bit-day, the SEL > 107.1 MeV⋅cm2/mg, and the TID > 100 Krad(Si), but also a faster access speed of < 10 ns and a lower write power consumption of 14.664 mW in comparison with the related products. Deming Zhang, Dingyi Luo, Lang Zeng, Bi Wang 0002, Yue Zhang 0010, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2024 | HGT: Transformer Architecture for Arbitrary Hypergraphs Based on P-LaplacianabstractIn the real world, many complex relationship networks are better represented as hypergraphs, such as social networks and EDA netlists. Traditional graph neural networks struggle to capture the high-order relationships expressed by hyperedges, resulting in a loss of structure information. Transformer models can be adapted to run on fully connected graph neural networks. Unfortunately, this architecture ignores the inherent connectivity induction bias of hypergraphs and fails to capture the relative positional relationships between hyper-nodes and hyperedges. Therefore, the challenge in hypergraph representation learning is how to capture the more complex and unique structural information of hypergraphs while ensuring the stability of the model. To address this issue, we propose the Hypergraph Transformer model (HGT), a general Transformer model capable of extracting hypergraph node embeddings while preserving rich structural information. In HGT, we introduce p*-Laplacian operator and HSPD algorithm. We assign a p*-Laplacian eigenvector as the positional encoding for each node by a neat way and design the HSPD algorithm to extract the hyperedge structural information in advance. To enhance the hypergraph representation capability of HGT, we incorporate the shortest path feature through the hyperedge into the computation of attention coefficients between nodes. Finally, we conduct extensive multidimensional experiments on HGT. The results validate that our approach does not increase the time complexity compared with the traditional p-Laplacian operator, while surpassing it across multiple indicators in downstream tasks, such as node classification, improving accuracy by 3% to 5%. Deming Zhang |
IJCNN | 3 |
| 2024 | In-Memory Wallace Tree Multipliers Based on Majority Gates Within Voltage-Gated SOT-MRAM Crossbar ArraysabstractIn-memory computing represents an efficient paradigm for high-performance computing using crossbar arrays of emerging nonvolatile devices. While various techniques have emerged to implement Boolean logic in memory, the latency of arithmetic circuits, particularly multipliers, significantly increases with bit-width. In this work, we introduce an in-memory Wallace tree multiplier based on majority gates within voltage-gated spin-orbit torque (SOT) magnetoresistive random access memory (MRAM) crossbar arrays. By utilizing a resistance sum, the majority gate is implemented during READ operations in voltage-gated SOT-MRAM crossbar arrays, resulting in reduced read currents and improved energy efficiency. We employ a series of READ and WRITE operations to perform multiplier calculations, leveraging the fast READ and WRITE speeds of voltage-gated SOT-MRAM devices. Furthermore, the use of five-input majority gates simplifies multiplication by employing uniform logic gates and reducing logic depth, thereby lowering the operation’s complexity and the total number of occupied cells. Our experimental results demonstrate that the proposed in-memory Wallace tree multipliers consume three times less energy for in-memory operations than previously reported$4\times 4$multipliers. Moreover, the proposed method reduces the delay overhead from O ($n^{2}$) to O ($\log _{2}{n}$), where$\mathit {n}$represents the number of bits. Yajuan Hui, Qingzhen Li, Leimin Wang, Cheng Liu 0008, Deming Zhang, Xiangshui Miao |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | Design of Ultracompact Content Addressable Memory Exploiting 1T-1MTJ CellabstractContent addressable memories (CAMs) are a promising category of computing-in-memory (CiM) elements that can perform highly parallel and efficient search operations for routers, pattern matching, and other data-intensive applications. Various magnetic tunnel junction (MTJ)-based CAM designs have been proposed to realize zero standby power and high-performance search. However, due to the relatively small tunnel magneto-resistance (TMR) ratio, MTJ-based CAMs require extra transistors and differential MTJ branches to distinguish between the parallel and anti-parallel resistance states, resulting in significant area and energy overhead. In this article, we propose a device-circuit co-design approach for an ultracompact CAM design by only exploiting a 1T-1MTJ structure in each cell. We propose a 2-step search scheme to enable the parallel in-memory search operation across the proposed CAM array and demonstrate the sufficient sensing margin of the array in a successful search operation. Evaluation results suggest that our proposed 1T-1MTJ-based CAM design improves$179\times /301\times $area efficiency compared with the state-of-the-art 15T-4MTJ/20T-6MTJ CAM design. Application benchmarking on hyperdimensional computing (HDC) inference shows a$54.6\times /12.8\times $speedup compared with GPU/20T-6MTJ CAM-based approaches. Cheng Zhuo, Kai Ni 0004, Mohsen Imani, Yuxuan Luo 0001, Shaodi Wang, Deming Zhang, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | A Novel 9T1C-SRAM Compute-In-Memory Macro With Count-Less Pulse-Width Modulation Input and ADC-Less Charge-Integration-Count OutputabstractThis paper presents a novel compute-in-memory (CIM) macro, which mainly consists of three modules: input generator, 9T1C-SRAM CIM array and charge-integration-count output (CICO) circuit. For the input generator, it can achieve the pulse-width modulation mapping scheme without counts, leading to a small area overhead. For the CIM array, one row-cascade current mirror circuit instead of a bias voltage source is shared by all CIM cells in one row. And in each CIM cell, its multiply result is characterized by the charge on its capacitor. Based on the charge-sharing principle, the accumulated result is represented by the charge ($\text{Q}_{\text {CBL}}$) on the charge-bit-line (CBL). In this way, the voltage on the CBL is limited regardless of the number of rows of the CIM array, allowing the large-scale CIM array. For the CICO circuit, it is proposed to quantify the$\text{Q}_{\text {CBL}}$without the ADC, aiming to achieve high area efficiency. With the 14nm FinFET design kit, the design specification of the proposed CIM macro is introduced in detail and its performance is evaluated. Simulation results show that the proposed CIM macro can achieve 4–1370 TOPS/W energy efficiency with IN/W/OUT precision of 6/1/6b and 98.48%/84.56% test accuracy on MNIST and CIFAR-10. Deming Zhang, Zhipeng Guo 0006, You Wang 0002, Yue Zhang 0010, Lang Zeng |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2021 | A Reconfigurable Arbiter PUF Based on STT-MRAMabstractWith the rapid development of the Internet of Things (IoT) infrastructure, electronic devices are becoming ubiquitous, in which authentication and secure communication are required. As a result, novel hardware security primitives have been developed to overcome the deficiencies of conventional security methods and address the growing security issues. Physical unclonable function (PUF) is an emerging hardware security primitive that plays an important role in authenticity and reliability of integrated circuits (ICs). Spin-transfer torque magne- toresistive random access memory (STT-MRAM) is a promising technology that is dense, fast, non-volatile, highly endurant and energy-efficient. STT-MRAM is considered a promising primitive as it has several intrinsic randomness sources, such as stochastic switching, process variations and statistical read/write failures. This paper proposes a novel hybrid STT-MRAM/complementary metal-oxide semiconductor (CMOS) based reconfigurable arbiter PUF. The functionality of the design is validated by a 28nm CMOS technology and a compact magnetic tunnel junction (MTJ) model. Simulation results show that the proposed PUF has a mean intra-hamming distance (HD) of 0.24%, a mean inter-HD of 51.1% and passes the National Institute of Standards and Technology (NIST) statistical tests. You Wang 0002, Zhengyi Hou, Deming Zhang, Erya Deng, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2021 | Spin-Orbit Torque Nonvolatile Flip-Flop DesignsabstractFlip-flops (FFs) are basic units in electronic circuits. Recently, nonvolatile FFs (NVFFs) have attracted great interests for power-gating applications and a variety of NVFFs have been proposed by integrating nonvolatile memory devices. Among them, magnetic tunnel junction (MTJ) based NVFFs show considerable potential in terms of zero static power consumption and high endurance. Nevertheless, the mainstream spin transfer torque (STT) effect based MTJ switching approach for data storing still consumes much dynamic power and long delay, limiting the system performance and data reliability. The spin-orbit torque (SOT) effect provides an alternative approach for high-speed and low-power MTJ switching, therefore rather promising for NVFF design. In this work, we propose four NVFF designs based on the FF architectures (either DFF or SRFF) and perpendicular MTJ (pMTJ). The circuit structures and operations are investigated, and the performance is evaluated and compared at the 40 nm process technology node. Simulation results show that the proposed NVFFs can achieve high read speed (<; 200 ps), low read power consumption (<; 10 fJ) and area efficiency. Erya Deng, Wang Kang 0001, Weisheng Zhao 0001, Shaoqian Wei, You Wang 0002, Deming Zhang |
ISCAS | 6 |
| 2021 | SpinSim: A Computer Architecture-Level Variation Aware STT-MRAM Performance Evaluation FrameworkabstractWith low power consumption, fast access speed, high scalability and infinite endurance, spin-transfer torque magnetoresistive random access memory (STT-MRAM) is considered as one of the most promising alternatives to SRAM. However, The performance of STT-MRAM is significantly influenced by several reliability issues, such as process variations and stochastic switching. Most of the reliability analysis of relative circuits are performed at bit-cell and memory level, while that at computer-system level is missing. This paper proposes an efficient framework for performance evaluation of STT-MRAM on computer architecture-level implemented by GEM5+NVMain co-simulator in consideration of the reliability issues. The results show that the overall average latency and energy of STT-MRAM can be up to 5.996% and 20.65% larger than that of the nominal cases in a computer system-level memory architecture taking reliability issues into account. Because reliability issues are considered during the design phase, our framework can provide more accurate performance evaluation and contribute to a higher yield of STT-MRAM based computer systems. You Wang 0002, Zhengyi Hou, Deming Zhang, Erya Deng, Gefei Wang, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2021 | Fully Single Event Double Node Upset Tolerant Design for Magnetic Random Access MemoryabstractBenefitting from its non-volatility, high speed, low power and inherent radiation hardened characteristic, magnetic random access memory (MRAM) has been used in aerospace and avionic electronics. Owing to its high sensing reliability, precharge differential sense amplifier (PCDSA) has been proposed and widely used in MRAM products. However, such PCDSA is based on the conventional CMOS technology and its sensing result is prone to be affected by the single event upset (SEU) and even the single event double node upset (SEDU) when the CMOS technology node shrinks into the nanometer scale. In this paper, we propose a novel PCDSA to tolerate the SEDU, in which the special three-input C-element that behaves as an inverter when its inputs have the same logic value and holds its previous value when its inputs have the different logic values is employed. By using a physics-based STT-MTJ compact model and a commercial CMOS 40 nm design kit, hybrid simulations have been performed to demonstrate its functionality and evaluate its performance. Simulation results show that it can fully tolerate the SEDU when the amount of the deposited charge (Qinj) reaches up to 2 pC. In the worst case where the Qinjis 2 pC, it can achieve a small recover time of 1.3368 ns and low recover energy dissipation of 1.967 pJ with the optimized VDDof 1 V. Deming Zhang, Lang Zeng, You Wang 0002, Bi Wang 0002, Erya Deng, Chuanjie Wang, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 1 |
| 2020 | A Modeling Attack Resilient Physical Unclonable Function Based on STT-MRAMabstractPhysical unclonable function (PUF) is considered as a promising hardware security primitive for a variety of applications. Recently, with the rapid development of integrated circuit (IC), the requirement for low complexity, high power efficiency and high performance PUFs become urgent. Moreover, a variety of powerful attack approaches have been carried out to counterfeit PUFs. This paper proposes a novel PUF design by utilizing the spin transfer torque magnetic random-access memory (STT-MRAM). The intrinsic process variation of STT-MRAM is exploited as an entropy source for generating PUF response. The primary performance metrics in terms of reliability, uniformity, uniqueness, and diffuseness of our proposed PUF have been verified, which validate its functionality. In addition, machine learning based modeling attacks are employed to evaluate the security level of proposed STT-MRAM based PUF (MPUF). The statistical results show that MPUF is much more immune to modeling attacks compared with the traditional Arbiter PUF. Zhengyi Hou, You Wang 0002, Deming Zhang, Hao Cai 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | Voltage-Gated Spin-Hall Effect Based Magnetic Non-Volatile Flip-Flop for High Speed, Low Power and Compact Cell AreaabstractIn this paper, we present a novel magnetic nonvolatile flip-flop (MNV-FF) for fast and low-power backup operation with a compact cell area. It employs perpendicular magnetic tunnel junctions (p-MTJs) as its non-volatile data backup storage units and exploits the voltage-gated spin-hall effect (VGSHE) for data backup operation. Benefitting from the assistance of the voltage-controlled magnetic anisotropy (VCMA) effect, the critical write current for 1-ns backup operation can be reduced to 3 μA or even lower, thus resulting in high speed and low power consumption. Moreover, such small write current allows to be driven by the cross-coupled inverters in the master latch, instead of a dedicated write driver, leading to a low cell area overhead. Additionally, by using an antiferromagnetic (AFM) metal that can provide both an exchange bias and the SHE instead of the heavy metal, no external magnetic field is required, making it suitable for practical applications. Our simulation results show that our proposed VGSHE-based MNV-FF can achieve 58.2× less backup energy, 1.85× less backup delay and 1.625× less cell area overhead than the previous SHE-based MNV-FF. Deming Zhang, Lang Zeng, Weisheng Zhao 0001 |
ISCAS | 2 |
| 2019 | Modulation and Demodulation of Digital Frequency Shift Keying System Based on Spin Torque Nano Oscillator with Voltage Controlled Magnetic Anisotropy EffectabstractIn this work, a spin torque nano oscillator (STNO) device whose frequency can be tuned by Voltage Controlled Magnetic Anisotropy effect (VCMA) is proposed. The requirement of magnetic bias field in previous STNO devices is eliminated by the introduction of VCMA effect. Based on VCMA-STNO, a novel architecture is proposed which can compose of a modulation/demodulation digital frequency shift keying (DFSK) communication system. The proposed architecture utilizes VCMA-STNO as core devices and is much simpler comparing with its CMOS counterpart. The proposed VCMA-STNO modulation/demodulation architecture will help to design next generation spintronics DFSK communication system. Lang Zeng, Zuodong Zhang, Haoxuan Chen, Tianqi Gao, Deming Zhang, Mingzhi Long, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2016 | Spin wave based synapse and neuron for ultra low power neuromorphic computation systemabstractIn this work, we have proposed that the neural synapses and neurons can be realized by utilizing spin waves (SWs) as information carrier. The SWs is excited by spin torque nano-oscillator (STNO), and detected with several different physical mechanisms: 1) tunneling magnetic-resistance 2) spin pumping and 3) inverse spin hall effect. The proposed SWs based synapses and neurons can be further combined together to form a neuromorphic computation system with crossbar structure. Possible ultra low power consumption and ultra high speed are the advantage of our proposed SWs based synapses and neurons. Lang Zeng, Deming Zhang, Youguang Zhang, Fanghui Gong, Tianqi Gao, Sa Tu, Haiming Yu, Weisheng Zhao 0001 |
ISCAS | 2 |
| 2015 | Energy-efficient neuromorphic computation based on compound spin synapse with stochastic learningabstractRecently, magnetic tunnel junction with in-plane magnetization (i-MTJ) has been exploited to behave as a binary stochastic synapse. However, it suffers from its limited level of synaptic weight, resulting in an inaccurate learning. In this work, a compound synapse that employs multiple perpendicular MTJs (p-MTJs) in series is proposed. It possesses an analog-like synaptic weight under weak programming conditions, which leads to a stochastic learning rule and low power consumption per synaptic event. By performing system-level simulations on the MNIST database, it has been demonstrated that such compound spin synapses can realize stochastic neuromorphic computation with high accuracy and low energy consumption. Deming Zhang, Lang Zeng, Yuanzhuo Qu, Youguang Zhang, Mengxing Wang 0001, Weisheng Zhao 0001, Tianqi Tang 0001, Yu Wang 0002 |
ISCAS | 1 |
| 2011 | Gosset lattice spherical vector quantizationwith lowcomplexityabstractThis paper introduces a novel and highly efficient realization of a spherical vector quantizer (SVQ), the "Gosset Low Complexity Vector Quantizer" (GLCVQ). The GLCVQ codebook is composed of vectors that are located on spherical shells of the Gosset lattice E8. A high encoding efficiency is achieved by representing the spherical vector codebook as aggregated permutation codes. Compared to previous algorithms, the computational complexity and memory consumption is further reduced by exploiting the properties of so called classleader root vectors and by a novel approach for the codevector-to-index-mapping. The GLCVQ concept can be generalized to vector dimensions that are multiples of eight. In particular, GLCVQ for 16-dimensional vectors is used in Amd. 6 to ITU-T Rec. G.729.1. Hauke Krüger, Bernd Geiser, Peter Vary, Hai Ting Li, Deming Zhang |
ICASSP | 5 |
| 2010 | Superwideband extension of g.718 and g.729.1 speech codecsabstractThis communication presents the recently standardized superwideband (SWB) extensions of ITU-T G.718 and G.729.1. These extensions were standardized as G.718 annex B and G.729.1 annex E. The SWB functionality is implemented using embedded scalable layers on top of the wideband (WB) core codecs, and it extends the bit rate of the codecs to 48 and 64 kbit/s for the G.718 and G.729.1, respectively. The main technology is a two-mode SWB coding method of the high frequencies. In addition, the G.729.1 SWB extension enhances the lower frequency range. The codec performance is illustrated with some listening test results extracted from the ITU-T Characterization phase. Lasse Laaksonen, Mikko Tammi, Vladimir Malenovsky, Tommy Vaillancourt, Mi Suk Lee, Tomofumi Yamanashi, Masahiro Oshikiri, Claude Lamblin, Balázs Kövesi, Lei Miao 0004, Deming Zhang, Jon Gibbs, Holly Francois |
INTERSPEECH | 11 |
| 2009 | Candidate proposal for ITU-T super-wideband speech and audio codingabstractThis paper describes the speech and audio codec that has been submitted to ITU-T by Huawei and ETRI as a candidate for the upcoming super-wideband and stereo extensions of Rec. G.729.1 and G.718. The core codec in the current implementation is G.729.1 and the encoded frequency range is increased from 7 kHz to 14 kHz. Therefore, the maximum bit rate is raised from 32 kbit/s to 64 kbit/s by adding five bitstream layers. A comprehensive overview of the codec is presented with a focus on the mono coding components. The results of the listening tests that have been conducted during the ITU-T qualification phase are summarized. The proposed codec passes all quality requirements for mono input signals. Bernd Geiser, Hauke Krüger, Heinrich W. Löllmann, Peter Vary, Deming Zhang, Hualin Wan, Hai Ting Li |
ICASSP | 5 |