Wang Ye

dblp:50/6156 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0002-7636-0227ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 An energy-efficient FeFET-based computing-in-memory macro using BEOL-integrated HZO ferroelectric capacitors
Weizeng Li, Zhidao Zhou, Linfang Wang, Junyu Zhu, Junzhe Shen, Hongyang Hu, Baihan Wang, Zhi Li 0062, Wang Ye, Zhongze Han, Hanghang Gao, Chunmeng Dou
Sci. China Inf. Sci.9
2025 PKNet: Infrared Small Target Detection via Parallel Interactive Kolmogorov-Arnold Network
abstract
Infrared small target detection (IRSTD) remains challenging due to low signal-to-clutter ratios, unpredictable environmental interference, and weak spatial features. Convolutional neural network (CNN)-based methods excel at extracting local details but often fail to capture long-range dependencies, while Transformer-based approaches model global interactions through self-attention mechanisms but suffer from high computational costs. To overcome these challenges, we introduce a novel parallel interactive kolmogorov–arnold network, termed PKNet. In this architecture, the CNN branch is designed to extract fine-grained local features, while the KAN branch leverages learnable univariate functions to model contextual dependencies. To further strengthen global feature representation, we design a multi-grained KAN (MG-KAN) Block, which enhances context modeling by promoting nonlinear interactions across multiple token dimensions, enabling efficient feature extraction while preserving long-range dependencies. Moreover, we develop a cyclic interactive fusion module that facilitates bidirectional information refinement between the CNN and KAN branches. This module dynamically aligns and integrates multi-scale local and global features, significantly improving the network’s ability to distinguish small targets from complex backgrounds. Extensive experiments on three public benchmarks demonstrate that PKNet achieves superior performance in terms of both detection accuracy and efficiency, substantially outperforming existing state-of-the-art methods. The code will be available at:https://github.com/Zhishe-Wang/PKNet.
Xiaomei Yan, Wang Ye, Chunfa Wang, Chaoqun Xia, Jiawei Xu 0004, Zhishe Wang
IEEE Trans. Geosci. Remote. Sens.2
2025 An RRAM Digital Computing-in-Memory Macro With Dual-Mode Multiplication and Maximum Value Rounding Adder Tree
abstract
Implementing digital computing-in-memory (DCIM) based on resistive memory (RRAM) faces several critical challenges due to the small signal margin, large device variations, and large energy- and area-overhead induced by the digital adder tree (AT). To address these issues, we propose an RRAM DCIM macro based on the standard foundry one-transistor-one-resistor (1T1R) cell array featuring: 1) dual-mode MAC operation for efficiency- or accuracy-oriented optimization; 2) margin-enhanced digitized unit (MEDU) to amplify the signal ratio; and 3) maximum value rounding AT (MVR-AT) to reduce its power- and area-overhead. A test chip is demonstrated using a 180 nm CMOS process to verify the concept. It achieves a peak energy efficiency (EF) of 63.08 TOPS/W in the efficiency-oriented mode and a minimum error rate of 1.58% in the accuracy-oriented mode. Their combination can meet the requirements of different workloads in AI computing tasks to optimize the overall power consumption with negligible accuracy loss.
Wang Ye, Hanghang Gao, Zhidao Zhou, Linfang Wang, Weizeng Li, Zhi Li 0062, Jinshan Yue, Xiaoxin Xu, Hongyang Hu, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.1
2024 A 2T P-Channel Logic Flash Cell for Reconfigurable Interconnection in Chiplet-Based Computing-In-Memory Accelerators
abstract
In this work, we propose a two-transistor (2T) p-type channel (p-channel) logic-compatible flash cell. Compared to the previous designs, the proposed structure features reduced area-cost and enhanced ability to pass through the logic ‘1’. Due to these advantages, we explore its application as the reconfigurable interconnections in the chiplet-based system. By integrating them into the silicon interposer, the 2T p-channel flash cells can potentially lead to the dense and flexible interconnection between multiple computing-in-memory (CIM) chiplets, resulting in highly reconfigurable and scalable chiplet-based CIM accelerators. A 180nm 1Kb 2T p-channel flash cell array is fabricated and characterized. The characterization results show the 2T p-channel flash cells exhibit a signal ratio >103over 1000 program/erase (P/E) cycles and the device-to-device variations are less than 21.07%. Their typical behaviors as routers are also confirmed by circuit simulations.
Weizeng Li, Linfang Wang, Zhi Li 0062, Wang Ye, Zhidao Zhou, Haiyang Zhou, Hanghang Gao, Jinshan Yue, Hongyang Hu, Fengman Liu, Chunmeng Dou
ISCAS4
2024 Write-Verify-Free MLC RRAM Using Nonbinary Encoding for AI Weight Storage at the Edge
abstract
High-density and reliable multilevel-cell (MLC) resistive random access memory (RRAM) is expected to meet the ever-increasing demand for on-chip weight storages in the intelligent edge devices. However, due to the device variations, many write-and-verify (WAV) iterations are usually required to program the RRAM cell, which causes high power consumption, long latency, and degradation on the memory lifetime. To address this issue, we propose a write–verify-free MLC RRAM macro for weight storage with 1) a cascode-current-mirror multibit write (CCM-MW) driver and 2) a nonbinary programming scheme (NB-PS) with a radix not greater than 2. A 180-nm 400-Kb RRAM test chip is demonstrated in silicon. For 2-bit-per-cell MLC storage, the value error rates can be reduced by 24.13% after introducing two redundant bits (RBDs). In addition, compared to the single-level cell (SLC) storage scheme, a 37.50% reduction in the number of cells can be achieved to store the ResNet-8 model with a 0.79% loss in inference accuracy without the need for WAV iterations.
Junjie An, Zhidao Zhou, Linfang Wang, Wang Ye, Weizeng Li, Hanghang Gao, Zhi Li 0062, Jinghui Tian, Hongyang Hu, Jinshan Yue, Lingyan Fan, Shibing Long, Qi Liu 0010, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.4
2023 FABRIKv: A Fast, Iterative Inverse Kinematics Solver for Surgical Continuum Robot with Variable Curvature Model
abstract
Due to the advantages of high flexibility, large workspace, and good human-body compatibility, flexible tendon-driven surgical continuum robots have attracted a lot of attention in robot-assisted minimally invasive surgery. However, due to the coupling of the position and angle of the continuum robot, and the easy deformation of the external force, its inverse kinematics solution has always been a challenge. This paper proposes a fast inverse kinematics solver for surgical continuum robots with a variable curvature model. Firstly, the deformation of the continuum robot is analyzed, and a representation method of the variable curvature model is proposed. Next, to solve the inverse kinematics problem when the continuum robot deforms under load, FABRIKv is proposed by improving the Forward And Backward Reaching Inverse Kinematics (FABRIK). During the inverse kinematics solution, the algorithm preserves the real-time nature of FABRIK and corrects for deformation effects caused by the load. Finally, the experiment verifies the rationality and effectiveness of the variable curvature model representation method, as well as the fastness and accuracy of the FARIKv solver.
Wang Ye, Xiaoyang Kang 0001, Jingjing Luo, Xiuhong Tang
IROS2
2021 Sparsity-Aware Clamping Readout Scheme for High Parallelism and Low Power Nonvolatile Computing-in-Memory Based on Resistive Memory
abstract
The input parallelism of resistive memory (RRAM) based nonvolatile computing-in-memory (nvCIM) structure is limited by the signal margin as well as the readout precision. In this work, we propose a sparsity-aware clamping (SAC) scheme and its circuit implementation for nvCIM by co-design of circuit and algorithm. It can adaptively tune the quantized range and resolution of the readout circuit according to the degree of sparsity in neural network models. As a result, the SAC scheme can effectively increase the input parallelism of nvCIMs without incurring degradation on the signal margin or increasing the hardware cost for analogue readout. A case study on processing a multi-layer perceptron (MLP) model with the proposed nvCIM structure shows that the SAC scheme can improve the throughput by 2 times and increase the energy efficiency by 25.35% with negligible inference accuracy loss.
Linfang Wang, Wang Ye, Junjie An, Chunmeng Dou, Qi Liu 0010, Meng-Fan Chang, Ming Liu 0022
ISCAS2