VLDB 2026 Research / reviewers in the wild / expert
Kejie Huang
dblp:05/10461
· DBLP profile ↗
48ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0003-3722-9979ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 19 · 5 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 15 since 2021Artificial intelligence and machine learning · 16 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flow Augmentation and Knowledge Distillation for Lightweight Face Presentation Attack Detection
Muhammad Shahid Jabbar, Muhammad Sohail Ibrahim, Taha Hasan Masood Siddique, Kejie Huang, Shujaat Khan |
FG | 4 |
| 2026 | HTCNN: High-Throughput Batch CNN Inference With Homomorphic EncryptionabstractHomomorphic Encryption (HE) technology allows for processing encrypted data, breaking through data isolation barriers and providing a promising solution for privacy-preserving computation. The integration of HE technology into Convolutional Neural Network (CNN) inference shows potential in addressing privacy issues in identity verification, medical imaging diagnosis, and various other applications. The CKKS HE algorithm stands out as a popular option for homomorphic CNN inference due to its capability to handle real number computations. However, challenges such as computational delays and resource overhead present significant obstacles to the practical implementation of homomorphic CNN inference, largely due to the complex nature of HE operations. In addition, current methods for speeding up homomorphic CNN inference primarily address individual images or large batches of input images, lacking a solution for efficiently processing a moderate number of input images with fast homomorphic inference capabilities. In response to these challenges, we introduce a novel leveled homomorphic CNN inference scheme aimed at reducing latency and improving throughput using the CKKS scheme. Our proposed inference strategy involves mapping multiple inputs to a set of ciphertext by exploiting the sliding window properties of convolutions to utilize CKKS's inherent Single-Instruction-Multiple-Data (SIMD) capability. To mitigate the delay associated with homomorphic CNN inference, we introduce optimization techniques, including mask-weight merging, rotation multiplexing, stride convolution segmentation, and folding rotations. The efficacy of our homomorphic inference scheme is demonstrated through evaluations carried out on the MNIST and CIFAR-10 datasets. Specifically, results from the MNIST dataset on a single CPU thread show that inference for 163 images can be completed in 10.4 seconds with an accuracy of 98.95%, which is a 6.9× throughput improvement over state-of-the-art works. Comparative analysis with existing methodologies highlights the superior performance of our proposed inference scheme in terms of latency, throughput, communication overhead, and memory utilization. Zewen Ye, Tianyu Wang 0037, Tianshun Huang, Yonggen Li, Chengxuan Wang, Ray C. C. Cheung, Kejie Huang |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2026 | A Bit-Level Loosely Coupled Spiking Neural Network Accelerator With Fast Inference and Hybrid Early TerminationabstractDeep spiking neural networks (SNNs) tend to suffer from long inference latency. While existing encoding schemes are insufficient, this work applies a bit-level loosely coupled (BLC) SNN in which output bits are generated sequentially from the current input bit. Consequently, the time steps can be further reduced with minor accuracy loss. The BLC SNN is optimized for hardware implementation and configured based on quantized artificial neural networks (ANNs). A hybrid early termination (ET) scheme is employed to skip redundant computation cycles without accuracy degradation. A pipelined digital accelerator architecture is designed to implement the BLC SNN, in which line buffers are utilized to maximize data reuse. Simulation results show that the proposed hybrid ET scheme reduces computation cycles by 28.31% on LeNet-5. Simulated in SMIC 28 nm technology, the accelerator consumes$0.10~\mu $J/img at 500 MHz with 30.49 TSOPS/W and is adaptable to 8/4 time steps of inference. Junchuan Gu, Yiwen Gu, Haibin Shen, Kejie Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video EditingabstractText-to-video diffusion models have made remarkable advancements. Driven by their ability to generate temporally coherent videos, research on zero-shot video editing using these fundamental models has expanded rapidly. To enhance editing quality, structural controls are frequently employed in video editing. Among these techniques, cross-attention mask control stands out for its effectiveness and efficiency. However, when cross-attention masks are naively applied to video editing, they can introduce artifacts such as blurring and flickering. Our experiments uncover a critical factor overlooked in previous video editing research: cross-attention masks are not consistently clear but vary with model structure and denoising timestep. To address this issue, we propose the metric Mask Matching Cost (MMC) that quantifies this variability and propose FreeMask, a method for selecting optimal masks tailored to specific video editing tasks. Using MMC-selected masks, we further improve the masked fusion mechanism within comprehensive attention features, e.g., temp, cross, and self-attention modules. Our approach can be seamlessly integrated into existing zero-shot video editing frameworks with better performance, requiring no control assistance or parameter fine-tuning but enabling adaptive decoupling of unedited semantic layouts with mask precision control. Extensive experiments demonstrate that FreeMask achieves superior semantic fidelity, temporal consistency, and editing quality compared to state-of-the-art methods. Lingling Cai, Hangjie Yuan, Yingya Zhang, Shiwei Zhang 0001, Kejie Huang |
AAAI | 6 |
| 2025 | VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion TransformersabstractThe Diffusion Transformers Models (DiTs) have transitioned the network architecture from traditional UNets to transformers, demonstrating exceptional capabilities in image generation. Although DiTs have been widely applied to high-definition video generation tasks, their large parameter size hinders inference on edge devices. Vector quantization (VQ) can decompose model weight into a codebook and assignments, allowing extreme weight quantization and significantly reducing memory usage. In this paper, we propose VQ4DiT, a fast post-training vector quantization method for DiTs. We found that traditional VQ methods calibrate only the codebook without calibrating the assignments. This leads to weight sub-vectors being incorrectly assigned to the same assignment, providing inconsistent gradients to the codebook and resulting in a suboptimal result. To address this challenge, VQ4DiT calculates the candidate assignment set for each weight sub-vector based on Euclidean distance and reconstructs the sub-vector based on the weighted average. Then, using the zero-data and block-wise calibration method, the optimal assignment from the set is efficiently selected while calibrating the codebook. VQ4DiT quantizes a DiT XL/2 model on a single NVIDIA A100 GPU within 20 minutes to 5 hours depending on the different quantization settings. Experiments show that VQ4DiT establishes a new state-of-the-art in model size and performance trade-offs, quantizing weights to 2-bit precision while retaining acceptable image generation quality. Juncan Deng, Shuaiting Li, Zeyu Wang 0010, Kedong Xu, Kejie Huang |
AAAI | 6 |
| 2025 | MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector QuantizationabstractVector quantization(VQ) is a hardware-friendly DNN compression method that can reduce the storage cost and weight-loading datawidth of hardware accelerators. However, conventional VQ techniques lead to significant accuracy loss because the important weights are not well preserved. To tackle this problem, a novel approach called MVQ is proposed, which aims at better approximating important weights with a limited number of codewords. At the algorithm level, our approach removes the less important weights through N:M pruning and then minimizes the vector clustering error between the remaining weights and codewords by the masked k-means algorithm. Only distances between the unpruned weights and the codewords are computed, which are then used to update the codewords. At the architecture level, our accelerator implements vector quantization on an EWS (Enhanced weight stationary) CNN accelerator and proposes a sparse systolic array design to maximize the benefits brought by masked vector quantization. Shuaiting Li, Chengxuan Wang, Juncan Deng, Zeyu Wang 0010, Zewen Ye, Zongsheng Wang, Haibin Shen, Kejie Huang |
ASPLOS (1) | 8 |
| 2025 | A FeFET-Based Compute-in-Memory Architecture on FPGA for Neural Network InferenceabstractImplementing compute-in-memory (CIM) architectures on FPGA offers an effective solution to the von Neumann bottleneck by enabling fast configuration and computation directly within memory. Traditional custom solutions rely on the modification of block RAM (BRAM) to implement memory computing. However, single-word-line activation of BRAM results in low parallelism, and the need for additional adder trees to accumulate partial sums further limits efficiency. To overcome these limitations, we propose a CIM core based on a 2T1C structure as a replacement for BRAM units. This core utilizes a charge redistribution mechanism and reuse of ADC capacitors, achieving high parallelism, low power consumption, and a compact area. By incorporating computational capabilities within a single cell, our design enables dual parallelism, further enhancing performance and efficiency. In addition, we present an automated deployment and mapping tool for deep neural networks (DNNs) on FPGA, allowing users to rapidly develop FPGA-based solutions for different network architectures. Compared to state-of-the-art solutions, our design achieves a peak throughput improvement of 4.5× and a reduction in area by 53%. Minghan Jiang, Yonggen Li, Rui Xiao 0003, Haibin Shen, Kejie Huang |
FCCM | 5 |
| 2025 | ViM-VQ: Efficient Post-Training Vector Quantization for Visual MambaabstractVisual Mamba networks (ViMs) extend the selective state space model (Mamba) to various vision tasks and demonstrate significant potential. As a promising compression technique, vector quantization (VQ) decomposes network weights into codebooks and assignments, significantly reducing memory usage and computational latency, thereby enabling the deployment of ViMs on edge devices. Although existing VQ methods have achieved extremely low-bit quantization (e.g., 3-bit, 2-bit, and 1-bit) in convolutional neural networks and Transformer-based networks, directly applying these methods to ViMs results in unsatisfactory accuracy. We identify several key challenges: 1) The weights of Mamba-based blocks in ViMs contain numerous outliers, significantly amplifying quantization errors. 2) When applied to ViMs, the latest VQ methods suffer from excessive memory consumption, lengthy calibration procedures, and suboptimal performance in the search for optimal codewords. In this paper, we propose ViM-VQ, an efficient post-training vector quantization method tailored for ViMs. ViM-VQ consists of two innovative components: 1) a fast convex combination optimization algorithm that efficiently updates both the convex combinations and the convex hulls to search for optimal codewords, and 2) an incremental vector quantization strategy that incrementally confirms optimal codewords to mitigate truncation errors. Experimental results demonstrate that ViM-VQ achieves state-of-the-art performance in low-bit quantization across various visual tasks. Juncan Deng, Shuaiting Li, Zeyu Wang 0010, Kedong Xu, Kejie Huang |
ICCV | 6 |
| 2025 | SSVQ: Unleashing the Potential of Vector Quantization with Sign-SplittingabstractVector Quantization (VQ) has emerged as a prominent weight compression technique, showcasing substantially lower quantization errors than uniform quantization across diverse models, particularly in extreme compression scenarios. However, its efficacy during fine-tuning is limited by the constraint of the compression format, where weight vectors assigned to the same codeword are restricted to updates in the same direction. Consequently, many quantized weights are compelled to move in directions contrary to their local gradient information. To mitigate this issue, we introduce a novel VQ paradigm, Sign-Splitting VQ (SSVQ), which decouples the sign bit of weights from the codebook. Our approach involves extracting the sign bits of uncompressed weights and performing clustering and compression on all-positive weights. We then introduce latent variables for the sign bit and jointly optimize both the signs and the codebook. Additionally, we implement a progressive freezing strategy for the learnable sign to ensure training stability. Extensive experiments on various modern models and tasks demonstrate that SSVQ achieves a significantly superior compression-accuracy trade-off compared to conventional VQ. Furthermore, we validate our algorithm on a hardware accelerator, showing that SSVQ achieves a 3$\times$ speedup over the 8-bit compressed model by reducing memory access. Our code is available at https://github.com/list0830/SSVQ. Shuaiting Li, Juncan Deng, Chengxuan Wang, Kedong Xu, Rongtao Deng, Haibin Shen, Kejie Huang |
ICCV | 8 |
| 2025 | A 1FeFET-1T-1C based Compute-in-Memory Macro with Capacitor Reused Pipeline SAR ADCabstractComputing-in-memory (CIM) significantly reduces latency and power consumption by combining computation and memory, typically utilizing non-volatile memories (NVM). However, device manufacturing non-uniformity on NVMs can cause output deviations. Additionally, the necessity for bit-shifting circuits and Analog-to-Digital Converters (ADC) increases the area and power overhead. To tackle these challenges, we propose a high-density 1FeFET-1T-1C based CIM macro, integrated with a pipeline Successive-Approximation-Register (SAR) ADC. The design introduces a capacitor structure that counters the non-uniformity issues inherent in FeFET devices. Also, the capacitor array is reused as charge-redistribution and ADCs, substantially minimizing the area and power overhead. Moreover, the pipeline architecture accelerates the conversion process, achieving high speed and high precision. The design is implemented using SMIC 55nm PDK. The energy efficiency (EF) and area efficiency (AF) of the proposed macro are 80.9 TOPS/W and 1.161 TOPS/mm2, respectively. The inference accuracy reaches 91.2% on the CIFAR-10 dataset. Minghan Jiang, Rui Xiao 0003, Shuaiting Li, Yishu Zhang, Haibin Shen, Kejie Huang |
ISCAS | 8 |
| 2025 | Knowledge distillation with predicted depth for robust and lightweight face presentation attack detection
Muhammad Shahid Jabbar, Taha Hasan Masood Siddique, Kejie Huang, Shujaat Khan |
Knowl. Based Syst. | 3 |
| 2025 | Spatio-temporal deep learning for improved face presentation attack detection
Shujaat Khan, Taha Hasan Masood Siddique, Muhammad Sohail Ibrahim, Abdul Jabbar Siddiqui, Kejie Huang |
Knowl. Based Syst. | 5 |
| 2025 | AdvSpoofGuard: Optimal transport driven robust face presentation attack detection system
Taha Hasan Masood Siddique, Shujaat Khan, Zeyu Wang 0010, Kejie Huang |
Knowl. Based Syst. | 4 |
| 2025 | Cross-Modal Adaptation for Object Detection in Infrared Remote Sensing ImageryabstractModern Thermal InfraRed (TIR) technology has been proven highly significant in Remote Sensing Imagery (RSI). Currently, multimodal RSI object detection based on RGB-TIR image pairs has attracted widespread research. However, capturing features in the TIR domain poses a challenge, as existing object detectors heavily focus on chromatic information in the RGB domain. Furthermore, the quality of RGB images can be influenced by complex environmental conditions, limiting the practicality of multimodal detection. In this paper, we introduce Cross-Modal-YOLO (CM-YOLO), a lightweight yet effective object detector specifically designed for TIR remote sensing images. CM-YOLO employs cross-modal adaptation to enhance the awareness of TIR-RGB modality translation. Specifically, we leverage a Prior Modality Translator (PMT) to learn the InfraRed-Visible (IV) features, which are incorporated into the detection backbone using our IV-Gate modules. Experimental results on the VEDAI dataset demonstrate that CM-YOLO significantly outperforms conventional methods. Moreover, CM-YOLO exhibits a strong generalization ability for TIR-based object detection in urban scenes on the FLIR dataset. Zeyu Wang 0010, Shuaiting Li, Kejie Huang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2025 | PQNTRU: Acceleration of NTRU-Based Schemes via Customized Post-Quantum ProcessorabstractPost-quantum cryptography (PQC) has rapidly evolved in response to the emergence of quantum computers, with the US National Institute of Standards and Technology (NIST) selecting four finalist algorithms for PQC standardization in 2022, including the Falcon digital signature scheme. Hawk is currently the only lattice-based candidate in NIST Round 2 additional signatures. Falcon and Hawk are based on the NTRU lattice, offering compact signatures, fast generation, and verification suitable for deployment on resource-constrained Internet-of-Things (IoT) devices. Despite the popularity of ML-DSA and ML-KEM, research on NTRU-based schemes has been limited due to their complex algorithms and operations. Falcon and Hawk's performance remains constrained by the lack of parallel execution in crucial operations like the Number Theoretic Transform (NTT) and Fast Fourier Transform (FFT), with data dependency being a significant bottleneck. This paper enhances NTRU-based schemes Falcon and Hawk through hardware/software co-design on a customized Single-Instruction-Multiple-Data (SIMD) processor, proposing new SIMD hardware units and instructions to expedite these schemes along with software optimizations to boost performance. Our NTT optimization includes a novel layer merging technique for SIMD architecture to reduce memory accesses, and the use of modular algorithms (Signed Montgomery and Improved Plantard) targets various modulus data widths to enhance performance. We explore applying layer merging to accelerate fixed-point FFT at the SIMD instruction level and devise a dual-issue parser to streamline assembly code organization to maximize dual-issue utilization. A System-on-chip (SoC) architecture is devised to improve the practical application of the processor in real-world scenarios. Evaluation on 28$nm$technology and field programmable gate array (FPGA) platform shows that our design and optimizations can increase the performance of Hawk signature generation and verification by over 7$\times$. Zewen Ye, Junhao Huang 0001, Tianshun Huang, Yudan Bai, Guangyan Li, Donald Donglong Chen, Ray C. C. Cheung, Kejie Huang |
IEEE Trans. Computers | 10 |
| 2025 | A Robust Computing-in-Memory Macro With 2T1R1C Cells and Reused Capacitors for Successive-Approximation ADCabstractComputing-in-memory (CIM) has emerged as a practical paradigm to bypass the von Neumann bottleneck. However, traditional CIM schemes face challenges due to the nonideal characteristics of nonvolatile memory (NVM). To address this issue, this work provides a resistive random access memory (RRAM)-based CIM macro employing two-transistor-one-RRAM–one-capacitor (2T1R1C) cells, with capacitors reused for the successive-approximation analog-to-digital converter (SAR ADC). Single-level RRAM is utilized to mitigate resistance variation. The multiply-accumulate (MAC) operation is performed via the charge and discharge of capacitors, enhancing robustness across different process, voltage, and temperature (PVT) corners. The capacitors in 2T1R1C cells are repurposed as sampling capacitors to integrate the ADC with the array. A precision-adjustable SAR (PA-SAR) logic is proposed to generate partial sums at varying precision levels aligned with different input bits, optimizing energy efficiency while maintaining reliability. Our proposed 2T1R1C array features an average area of$3.403~\mu $m2 for each cell, which accounts for 87.46% of the total macro area. The total macro area is 1.020 mm2 with a capacity of 256 Kb, achieving an energy density of 0.201 TOPS/mm2. The PA-SAR logic boosts energy efficiency to 44.71 TOPS/W, marking a 38.55% improvement over conventional full-precision schemes. Rui Xiao 0003, Minghan Jiang, Haibin Shen, Kejie Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | RVSLH: Acceleration of Postquantum Standard SLH-DSA With Customized RISC-V ProcessorabstractPostquantum cryptography (PQC) has developed quickly in response to the rise of quantum computers. The US National Institute of Standards and Technology (NIST) recently released three PQC standards, one of which is the hash-based standard stateless hash-based digital signature standard (SLH-DSA), built on SPHINCS+ selected during the NIST Round 3 submissions. Despite its potential, SLH-DSA’s performance is hindered by inefficient execution in the hash function and extensive memory accesses, with the data dependency of the hash function presenting a notable bottleneck. This brief aims to enhance the efficiency of SHAKE-based SLH-DSA schemes using hardware/software co-design on a customized RISC-V processor. We incorporate tightly coupled hardware units and instructions on RISC-V to expedite SLH-DSA, coupled with memory optimizations to enhance overall performance. The contributions of this brief are twofold. First, our design introduces customized single-instruction-multiple-data (SIMD) instructions and corresponding computation hardware units to accelerate Keccak, the essential operation of SHAKE256. In addition, our design streamlines hash operation memory accesses by reorganizing memory space. The proposed processor’s performance is evaluated on Artix-7 field-programmable gate array (FPGA) and 28-nm technology application-specific integrated circuit (ASIC), demonstrating approximately 15 times acceleration compared with the baseline design. Furthermore, it surpasses recent state-of-the-art works in terms of the performance-power–area product. Zewen Ye, Xin Li 0177, Chuhui Wang, Ray C. C. Cheung, Kejie Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | A Folded Computation-in-Memory Accelerator for Fast Polynomial Multiplication in BIKE
Chuhui Wang, Zewen Ye, Haibin Shen, Kejie Huang |
Euro-Par (2) | 4 |
| 2024 | Ground-Guided Conditional Pixel Synthesizer for Height-based Satellite Imagery Super-ResolutionabstractRemote sensing plays a crucial role in various fields. However, challenges associated with acquiring high-resolution data from satellite cameras significantly limit their practical applications. The high semantic density per pixel in satellite images makes it challenging for existing methods to extract adequate geometric and semantic information from extremely low-resolution inputs for super-resolution reconstruction. This paper introduces a satellite imagery super-resolution architecture guided by ground-view images, framing the problem as neural pixel synthesis with satellite camera height as a variable factor. This approach proposes a hypergraph-based cross-view mapper module that achieves low-order geometric registration and high-order feature fusion by capturing cross-view visual correlations, accompanied by a height-based pixel synthesizer for continuous multi-level super-resolution, conceptualized as neural rendering. Furthermore, we have developed a multi-level resolution satellite image dataset, complete with ground images from corresponding locations. Extensive experiments on diverse datasets validate the effectiveness of our proposed method in a range of application scenarios. Kejie Huang |
IJCNN | 2 |
| 2024 | Tetris-SDK: Efficient Convolution Layer Mapping with Adaptive Windows for Fast In Memory ComputingabstractShifted-and-Duplicated-Kernel (SDK) mapping has emerged as a promising technique for accelerating convolutional layers in Compute-In-Memory (CIM) architectures. While state-of-the-art SDK variants have achieved decent mapping efficiency, optimizations are still desired to enhance CIM utilization and reduce computing cycles. In this work, we propose Tetris-SDK, a novel tool that exploits adaptive windows to further improve the performance of convolution layer mapping. These windows can accommodate a larger number of input channels, increase array utilization at marginal space, and adjust window shapes to minimize compute latency. Our experiments with a 512 × 512 CIM array demonstrate that Tetris-SDK remarkably accelerates CNN layers up to 78.4×, 8×, and 1.3× compared to the baseline mapping algorithms, i.e., img2col, SDK, and VW-SDK, respectively. This shows that Tetris-SDK is a promising design automation solution to map Convolutional Neural Networks in CIM hardware. Kejie Huang, Bo Wang 0020 |
ISCAS | 2 |
| 2024 | Bridging partial-gated convolution with transformer for smooth-variation image inpainting
Zeyu Wang 0010, Haibin Shen, Kejie Huang |
Multim. Tools Appl. | 3 |
| 2024 | An All-digital Compute-in-memory FPGA Architecture for Deep Learning AccelerationabstractField Programmable Gate Array (FPGA) is a versatile and programmable hardware platform, which makes it a promising candidate for accelerating Deep Neural Networks (DNNs). However, FPGA’s computing energy efficiency is low due to the domination of energy consumption by interconnect data movement. In this article, we propose an all-digital Compute-in-memory FPGA architecture for deep learning acceleration. Furthermore, we present a bit-serial computing circuit of the Digital CIM core for accelerating vector-matrix multiplication (VMM) operations. A Network-CIM-deployer ( NCIMD ) is also developed to support automatic deployment and mapping of DNN networks. NCIMD provides a user-friendly API of DNN models in Caffe format. Meanwhile, we introduce a Weight-stationary dataflow and describe the method of mapping a single layer of the network to the CIM array in the architecture. We conduct experimental tests on the proposed FPGA architecture in the field of Deep Learning (DL), as well as in non-DL fields, using different architectural layouts and mapping strategies. We also compare the results with the conventional FPGA architecture. The experimental results show that compared to the conventional FPGA architecture, the energy efficiency can achieve a maximum speedup of 16.1×, while the latency can decrease up to 40% in our proposed CIM FPGA architecture. Yonggen Li, Xin Li 0177, Haibin Shen, Jicong Fan 0002, Yanfeng Xu, Kejie Huang |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2024 | ProgramGalois: A Programmable Generator of Radix-4 Discrete Galois Transformation Architecture for Lattice-Based CryptographyabstractLattice-based cryptography (LBC) has been established as a prominent research field, with particular attention on post-quantum cryptography (PQC) and fully homomorphic encryption (FHE). As the implementing bottleneck of PQC and FHE, number theoretic transform (NTT) has been extensively studied. However, current works struggled with scalability, hindering their adaptation to various parameters, such as bit width and polynomial length. In this article, we proposed a novel Discrete Galois Transformation (DGT) algorithm utilizing the radix-4 variant to achieve a higher level of parallelism to the existing NTT. Furthermore, to implement the efficient radix-4 DGT adapting more LBCs, we proposed a set of scalable building blocks, including a modified Barrett modular multiplier accepting arbitrary modulus with only one integer multiplier, a radix-4 DGT butterfly unit, and a stream permutation network. The proposed modules are implemented on the Xilinx Virtex-7 and U250 FPGA to evaluate resource utilization and performance. Lastly, a design space exploration framework is proposed to generate optimized radix-4 DGT hardware constrained by polynomial and platform parameters. The sensitivity analysis showcases the generated hardware’s performance and scalability. The implementation results on the Xilinx Virtex-7 and U250 FPGA show significant performance improvements over the state-of-the-art works, which reached at least 35%, 192%, and 68% area-time product improvements in terms of LUTs, BRAMs, and DSPs, respectively. Guangyan Li, Zewen Ye, Donald Donglong Chen, Wangchen Dai, Gaoyu Mao, Kejie Huang, Ray C. C. Cheung |
ACM Trans. Reconfigurable Technol. Syst. | 6 |
| 2023 | Thermal Infrared Image Inpainting Via Edge-Aware GuidanceabstractImage inpainting has achieved fundamental advances with deep learning. However, almost all existing inpainting methods aim to process natural images, while few target Thermal Infrared (TIR) images, which have widespread applications. When applied to TIR images, conventional inpainting methods usually generate distorted or blurry content. In this paper, we propose a novel task—Thermal Infrared Image Inpainting, which aims to reconstruct missing regions of TIR images. Crucially, we propose a novel deep-learning-based model TIR-Fill. We adopt the edge generator to complete the canny edges of broken TIR images. The completed edges are projected to the normalization weights and biases to enhance edge awareness of the model. In addition, a refinement network based on gated convolution is employed to improve TIR image consistency. The experiments demonstrate that our method outperforms state-of-the-art image inpainting approaches on FLIR thermal dataset. Zeyu Wang 0010, Haibin Shen, Changyou Men, Kejie Huang |
ICASSP | 5 |
| 2023 | WaveIPT: Joint Attention and Flow Alignment in the Wavelet domain for Pose TransferabstractHuman pose transfer aims to synthesis a new image of the source person in a target pose. Among the various existing methods, attention and flow have emerged as two of the most popular and effective approaches. Attention excels at preserving the semantic structure of the source image, which is more reflected in the low-frequency domain. Contrastively, flow is better at retaining fine-grained texture details in the high-frequency domain. To leverage the advantages of both attention and flow simultaneously, this paper proposes Wavelet-aware Image-based Pose Transfer (WaveIPT) as a novel approach to fuse the attention and flow in the wavelet domain. To improve the fusion effect and avoid interference from irrelevant information across different frequencies, WaveIPT first applies Intra-scale Local Correlation (ILC) to adaptively fuse attention and flow in the same scale according to their strengths in low and high-frequency domains. Subsequently, WaveIPT employs Inter-scale Feature Interaction (IFI) to explore inter-scale frequency features, facilitating effective information transfer across different scales. Furthermore, we introduce Progressive Flow Regularization (PFR), an effective method that alleviates the challenges of flow estimation under large pose differences. The experiments on the DeepFashion dataset demonstrate that WaveIPT achieves a new state-of-the-art in terms of both FID and LPIPS, with improvements of 4.97% and 3.89%, respectively. Tingwei Gao, Haitian Jiang, Haibin Shen, Kejie Huang |
ICCV | 5 |
| 2023 | TIRDet: Mono-Modality Thermal InfraRed Object Detection Based on Prior Thermal-To-Visible TranslationabstractCross-modality images that combine visible-infrared spectra can provide complementary information for object detection. In particular, they are well-suited for autonomous vehicle applications in dark environments with limited illumination. However, it is time-consuming to acquire a large number of pixel-aligned visible-thermal image pairs, and real-time alignment is challenging in practical driving systems. Furthermore, the quality of visible-spectrum images can be adversely affected by complex environmental conditions. In this paper, we propose a novel neural network called TIRDet, which only utilizes Thermal InfraRed (TIR) images for mono-modality object detection. To compensate for the lacked visible-band information, we adopt a prior Thermal-To-Visible (T2V) translation model to obtain the translated visible images and the latent T2V codes. In addition, we introduce a novel attention-based Cross-Modality Aggregation (CMA) module, which can augment the modality-translation awareness of TIRDet by preserving the T2V semantic information. Extensive experiments on FLIR and LLVIP datasets demonstrate that our TIRDet significantly outperforms all mono-modality detection methods based on thermal images, and it even surpasses most State-Of-The-Art (SOTA) multispectral methods using visible-thermal image pairs. Code is available at https://github.com/zeyuwang-zju/TIRDet Zeyu Wang 0010, Fabien Colonnier, Jinghong Zheng 0001, Jyotibdha Acharya, Kejie Huang |
ACM Multimedia | 6 |
| 2023 | OAW-GAN: occlusion-aware warping GAN for unified human video synthesis
Dongxu Wei, Kejie Huang, Jiashen Hua, Baisheng Lai, Haibin Shen |
Appl. Intell. | 2 |
| 2023 | IFGLT: Information fusion guided lightweight Transformer for image denoising
Fengyin Liu, Ziqun Zhou, Changyou Men, Kejie Huang |
J. Vis. Commun. Image Represent. | 5 |
| 2023 | MsVRL: Self-Supervised Multiscale Visual Representation Learning via Cross-Level Consistency for Medical Image SegmentationabstractAutomated medical image segmentation for organs or lesions plays an essential role in clinical diagnoses and treatment plannings. However, training an accurate and robust segmentation model is still a long-standing challenge due to the time-consuming and expertise-intensive annotations for training data, especially 3-D medical images. Recently, self-supervised learning emerges as a promising approach for unsupervised visual representation learning, showing great potential to alleviate the expertise annotations for medical images. Although global representation learning has attained remarkable results on iconic datasets, such as ImageNet, it can not be applied directly to medical image segmentation, because the segmentation task is non-iconic, and the targets always vary in physical scales. To address these problems, we propose a Multi-scale Visual Representation self-supervised Learning (MsVRL) model, to perform finer-grained representation and deal with different target scales. Specifically, a multi-scale representation conception, a canvas matching method, an embedding pre-sampling module, a center-ness branch, and a cross-level consistent loss are introduced to improve the performance. After pre-trained on unlabeled datasets (RibFrac and part of MSD), MsVRL performs downstream segmentation tasks on labeled datasets (BCV, spleen of MSD, and KiTS). Results of the experiments show that MsVRL outperforms other state-of-the-art works on these medical image segmentation tasks. Ruifeng Zheng, Senxiang Yan, Hongcheng Sun, Haibin Shen, Kejie Huang |
IEEE Trans. Medical Imaging | 6 |
| 2023 | FDA-GAN: Flow-Based Dual Attention GAN for Human Pose TransferabstractHuman pose transfer aims at transferring the appearance of the source person to the target pose. Existing methods utilizing flow-based warping for non-rigid human image generation have achieved great success. However, they fail to preserve the appearance details in synthesized images since the spatial correlation between the source and target is not fully exploited. To this end, we propose the Flow-based Dual Attention GAN (FDA-GAN) to apply occlusion- and deformation-aware feature fusion for higher generation quality. Specifically, deformable local attention and flow similarity attention, constituting the dual attention mechanism, can derive the output features responsible for deformable- and occlusion-aware fusion, respectively. Besides, to maintain the pose and global position consistency in transferring, we design a pose normalization network for learning adaptive normalization from the target pose to the source person. Both qualitative and quantitative results show that our method outperforms state-of-the-art models in public iPER and DeepFashion datasets. Kejie Huang, Dongxu Wei, Zhaoyan Ming, Haibin Shen |
IEEE Trans. Multim. | 2 |
| 2023 | A Low-Power In-Memory Multiplication and Accumulation Array With Modified Radix-4 Input and Canonical Signed Digit WeightsabstractData transfer between the processing and storage units has become a significant bottleneck in modern von Neumann computing systems for artificial intelligence (AI) tasks. Computing in memory (CIM) has emerged as a promising candidate for lowering latency and power consumption. However, the conventional analog CIM schemes are suffering from reliability issues, which may significantly degenerate the accuracy of the computation. Recently, digitized input data and weights have been utilized for high-reliable in-memory computing. However, the properties of the digital memory and input data are not fully utilized. This article presents a novel low-power CIM scheme to further reduce the power consumption by using a modified radix-4 (M-RD4) booth algorithm at the input and a modified canonical signed digit (M-CSD) for the network weights. The simulation results show that M-RD4 and M-CSD reduce the number of nonzero activation bits by 24.2% and the number of nonzero weight bits by 36.0% in AlexNet, respectively. The power consumption can be reduced by 41.6% on average. The computing-power ratio at the fixed-point 8 bit is 60.7 tera operations per second per watt (TOPS/W), and the density is 0.177 TOPS/mm2. Rui Xiao 0003, Yewei Zhang, Bo Wang 0020, Yanfeng Xu, Jicong Fan 0002, Haibin Shen, Kejie Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2022 | A Reconfigurable Convolution-in-Pixel CMOS Image Sensor ArchitectureabstractThe separation of the data capture and analysis in modern vision systems has led to a massive amount of data transfer between the end devices and cloud computers, resulting in long latency, slow response, and high power consumption. Efficient hardware architectures are under focused development to enable Artificial Intelligence (AI) at the resource-limited sensing devices. One of the most promising solutions is to enable Processing-in-Pixel (PIP) scheme. However, the conventional schemes suffer from the low fill-factor issue. This paper proposes a PIP based Complementary Metal-Oxide-Semiconductor (CMOS) sensor architecture, which allows convolution operation before the column readout circuit to significantly reduce the overall power consumption while improving the resource utilization of the succeeding deep learning accelerator. The simulation results show that the proposed architecture could support the computing efficiency up to 3.37 TOPS/W at the 8-bit weight configuration, which is four times as high as the conventional schemes after normalization. The transistors required for each pixel are only 3.5T, significantly improving the fill-factor. Ruibing Song, Kejie Huang, Zongsheng Wang, Haibin Shen |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | An 8-Bit in Resistive Memory Computing Core With Regulated Passive Neuron and Bitline Weight MappingabstractThe rapid development of artificial intelligence (AI) and Internet of Things (IoT) increase the requirement for edge computing with low power and relatively high processing speed devices. The computing-in-memory (CIM) schemes based on emerging resistive nonvolatile memory (NVM) show great potential in reducing the power consumption for AI computing. However, the inconsistency of the NVM may significantly degenerate the performance of the neural network. In this article, we propose a low power resistive RAM (RRAM)-based CIM core to not only achieve high computing efficiency but also greatly enhance the robustness by bit line (BL) regulator and BL weight mapping algorithm. The simulation results show that the power consumption of our proposed 8-bit CIM core is only 12.6 mW ($256\times 256$at 8b). The spurious-free dynamic range (SFDR) and signal to noise and distortion ratio (SNDR) of the CIM core achieve 62.64 and 45.92 dB, respectively. The proposed BL weight mapping scheme improves the top-1 accuracy by 2.46% and 3.47% for AlexNet and VGG16 on ImageNet Large Scale Visual Recognition Competition 2012 (ILSVRC 2012) in 8-bit mode, respectively. Yewei Zhang, Kejie Huang, Rui Xiao 0003, Bo Wang 0020, Yanfeng Xu, Jicong Fan 0002, Haibin Shen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2021 | C2F-FWN: Coarse-to-Fine Flow Warping Network for Spatial-Temporal Consistent Motion TransferabstractHuman video motion transfer (HVMT) aims to synthesize videos that one person imitates other persons' actions. Although existing GAN-based HVMT methods have achieved great success, they either fail to preserve appearance details due to the loss of spatial consistency between synthesized and exemplary images, or generate incoherent video results due to the lack of temporal consistency among video frames. In this paper, we propose Coarse-to-Fine Flow Warping Network (C2F-FWN) for spatial-temporal consistent HVMT. Particularly, C2F-FWN utilizes coarse-to-fine flow warping and Layout-Constrained Deformable Convolution (LC-DConv) to improve spatial consistency, and employs Flow Temporal Consistency (FTC) Loss to enhance temporal consistency. In addition, provided with multi-source appearance inputs, C2F-FWN can support appearance attribute editing with great flexibility and efficiency. Besides public datasets, we also collected a large-scale HVMT dataset named SoloDance for evaluation. Extensive experiments conducted on our SoloDance dataset and the iPER dataset show that our approach outperforms state-of-art HVMT methods in terms of both spatial and temporal consistency. Source code and the SoloDance dataset are available at https://github.com/wswdx/C2F-FWN. Dongxu Wei, Xiaowei Xu 0004, Haibin Shen, Kejie Huang |
AAAI | 4 |
| 2021 | DualPathGAN: Facial reenacted emotion synthesisabstractAbstract Facial reenactment has developed rapidly in recent years, but few methods have been built upon reenacted face in videos. Facial‐reenacted emotion synthesis can make the process of facial reenactment more practical. A facial‐reenacted emotion synthesis method is proposed that includes a dual‐path generative adversarial network (GAN) for emotion synthesis and a residual‐mask network to impose structural restrictions to preserve the mouth shape of the source person. To train the dual‐path GAN more effectively, a learning strategy based on separated discriminators is proposed. The method is trained and tested on a very challenging imbalanced dataset to evaluate the ability to deal with complex practical scenarios. Compared with general emotion synthesis methods, the proposed method can generate more realistic facial emotion synthesised images or videos with higher quality while retaining the expression contents of the original videos. The DualPathGAN achieves a Fréchet inception distance (FID) score of 9.20, which is lower than the FID score of 11.37 achieved with state‐of‐the‐art methods. Jiahui Kong, Haibin Shen, Kejie Huang |
IET Comput. Vis. | 3 |
| 2021 | Foreground-Background Parallel Compression With Residual Encoding for Surveillance VideoabstractThe data storage has been one of the bottlenecks in surveillance systems. The conventional video compression schemes such as H.264 and H.265 do not fully utilize the low information density characteristic of the surveillance video, and they attach equal importance to foreground and background when performing compression. In this article, we propose a novel video compression scheme that compresses the foreground and background of the surveillance video separately. The compression ratio is greatly improved by sharing background information among adjacent frames through an adaptive background updating and interpolation module. Besides, we present two different schemes to compress the foreground and compare their performance in the ablation study to show the importance of temporal information for video compression. In the decoding end, a coarse-to-fine two-stage module is applied to achieve the composition of the foreground and background and the enhancements of frame quality. The experimental results show that our proposed method requires 49.75% less bpp (bits per pixel) than the conventional algorithm H.265 to achieve the same PSNR (36 dB) on the HEVC dataset. Lirong Wu, Kejie Huang, Haibin Shen, Lianli Gao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | GAC-GAN: A General Method for Appearance-Controllable Human Video Motion TransferabstractHuman video motion transfer has a wide range of applications in multimedia, computer vision, and graphics. Recently, due to the rapid development of Generative Adversarial Networks (GANs), there has been significant progress in the field. However, almost all existing GAN-based works are prone to address the mapping from human motions to video scenes, with scene appearances encoded individually in the trained models. Therefore, each trained model can only generate videos with a specific scene appearance, and new models are required to be trained to generate new appearances. Besides, existing works lack the capability of appearance control. For example, users have to provide video records of wearing new clothes or performing in new backgrounds to enable clothes or background changing in their synthetic videos, which greatly limits the application flexibility. In this paper, we propose General Appearance-Controllable GAN (GAC-GAN), a general method for appearance-controllable human video motion transfer. To enable general-purpose appearance synthesis, we propose to include appearance information in the conditioning inputs. Thus, once trained, our model can generate new appearances by altering the input appearance information. To achieve appearance control, we first obtain the appearance-controllable conditioning inputs, and then utilize a two-stage GAC-GAN to generate the corresponding appearance-controllable outputs, where we utilize an Appearance-Consistency GAN (ACGAN) loss, and a shadow extraction module for output foreground, and background appearance control respectively. We further build a solo dance dataset containing a large number of dance videos for training, and evaluation. Experimental results on our solo dance dataset, and iPER dataset show that our proposed GAC-GAN can not only support appearance-controllable human video motion transfer but also achieve higher video quality than state-of-art methods. Dongxu Wei, Xiaowei Xu 0004, Haibin Shen, Kejie Huang |
IEEE Trans. Multim. | 4 |
| 2020 | SNEQ: Semi-Supervised Attributed Network Embedding with Attention-Based QuantisationabstractLearning accurate low-dimensional embeddings for a network is a crucial task as it facilitates many network analytics tasks. Moreover, the trained embeddings often require a significant amount of space to store, making storage and processing a challenge, especially as large-scale networks become more prevalent. In this paper, we present a novel semi-supervised network embedding and compression method, SNEQ, that is competitive with state-of-art embedding methods while being far more space- and time-efficient. SNEQ incorporates a novel quantisation method based on a self-attention layer that is trained in an end-to-end fashion, which is able to dramatically compress the size of the trained embeddings, thus reduces storage footprint and accelerates retrieval speed. Our evaluation on four real-world networks of diverse characteristics shows that SNEQ outperforms a number of state-of-the-art embedding methods in link prediction, node classification and node recommendation. Moreover, the quantised embedding shows a great advantage in terms of storage and time compared with continuous embeddings as well as hashing methods. Tao He 0007, Lianli Gao, Jingkuan Song, Xin Wang 0019, Kejie Huang, Yuanfang Li |
AAAI | 5 |
| 2020 | A GAN-based Tunable Image Compression SystemabstractThe method of importance map has been widely adopted in DNN-based lossy image compression to achieve bit allocation according to the importance of image contents. However, insufficient allocation of bits in non-important regions often leads to severe distortion at low bpp (bits per pixel), which hampers the development of efficient content-weighted image compression systems. This paper rethinks content-based compression by using Generative Adversarial Network (GAN) to reconstruct the non-important regions. Moreover, multiscale pyramid decomposition is applied to both the encoder and the discriminator to achieve global compression of high-resolution images. A tunable compression scheme is also proposed in this paper to compress an image to any specific compression ratio without retraining the model. The experimental results show that our proposed method improves MS-SSIM by more than 10.3% compared to the recently reported GAN-based method [3] to achieve the same low bpp (0.05) on the Kodak dataset. Lirong Wu, Kejie Huang, Haibin Shen |
WACV | 2 |
| 2020 | An Efficient Hardware Accelerator for Structured Sparse Convolutional Neural Networks on FPGAsabstractDeep convolutional neural networks (CNNs) have achieved state-of-the-art performance in a wide range of applications. However, deeper CNN models, which are usually computation consuming, are widely required for complex artificial intelligence (AI) tasks. Though recent research progress on network compression, such as pruning, has emerged as a promising direction to mitigate computational burden, existing accelerators are still prevented from completely utilizing the benefits of leveraging sparsity due to the irregularity caused by pruning. On the other hand, field-programmable gate arrays (FPGAs) have been regarded as a promising hardware platform for CNN inference acceleration. However, most existing FPGA accelerators focus on dense CNN and cannot address the irregularity problem. In this article, we propose a sparsewise dataflow to skip the cycles of processing multiply-and-accumulates (MACs) with zero weights and exploit data statistics to minimize energy through zeros gating to avoid unnecessary computations. The proposed sparsewise dataflow leads to a low bandwidth requirement and high data sharing. Then, we design an FPGA accelerator containing a vector generator module (VGM) that can match the index between sparse weights and input activations according to the proposed dataflow. Experimental results demonstrate that our implementation can achieve 987-, 46-, and 57-imag/s performance for AlexNet, VGG-16, and ResNet-50 on Xilinx ZCU102, respectively, which provides 1.5×-6.7× speedup and 2.0×-6.0× energy efficiency over previous CNN FPGA accelerators. Kejie Huang, Shuyuan Yang 0003, Haibin Shen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2018 | Racetrack Memory based hybrid Look-Up Table (LUT) for low power reconfigurable computing
Kejie Huang, Yong Lian 0001 |
J. Parallel Distributed Comput. | 1 |
| 2018 | Accurate iris center localization method using facial landmark, snakuscule, circle fitting and binary connected component
Kejie Huang, Yue Qiu 0003, Haibin Shen |
Multim. Tools Appl. | 2 |
| 2016 | Magnetic Domain-Wall Racetrack Memory-Based Nonvolatile Logic for Low-Power Computing and Fast Run-Time-ReconfigurationabstractThe high power and the long global interconnect delay are two of the major bottlenecks that limit the further scaling down of the process nodes in the VLSI systems. Therefore, new technologies and computer architectures are under focused development to reduce the power consumption and the interconnect delay. Current-induced magnetic domain-wall (DW) racetrack memory (RM) has the advantages of nonvolatility, fast switching speed, and high density. It may offer opportunities to open a new paradigm of circuits and architectures to significantly alleviate the power and delay issues. This paper presents the magnetic DW RM-based new nonvolatile logic designs for the low-power computing and fast run-time-reconfiguration. Both data transfer and reconfiguration are achieved by shifting the magnetic strips. Verify-before-shift approach is used to greatly reduce the shifting energy. Compared with the conventional nonvolatile logic gate, the proposed nonvolatile logic scheme doubles the operating speed with 87% lower operating energy. Moreover, the proposed nonvolatile logic gates can be reconfigured after fabrication, which makes the designs more flexible and robust. The reference reconfiguration and the polarity reconfiguration are presented in this paper, which can be finished in 1 ns with 130 fJ/strip energy and 6 ns with 390 fJ/strip energy, respectively. Kejie Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | High-Density and High-Reliability Nonvolatile Field-Programmable Gate Array With Stacked 1D2R RRAM ArrayabstractThe huge area overhead of the interconnect is one of the critical issues in static random access memory (SRAM)-based field-programmable gate arrays (FPGAs), resulting in high power consumption and slow operation speed. Another critical issue is the volatile feature of the SRAM, which leads to high standby leakage current and long power-ON time. Resistive random access memory (RRAM) with a high resistance ratio and zero standby power possesses great potential in the FPGA applications. The conventional RRAM-based nonvolatile FPGAs (NVFPGAs) may use one-transistor 2-RRAM (1T2R) storage element to replace the SRAM or the one RRAM (1R) cell to replace both nMOS switch and SRAM. However, those NVFPGA schemes may suffer from the issues of low reliability, high configuration power, and high active leakage power. In this paper, we propose a novel element [one-diode two-RRAM (1D2R) cells] to replace the nMOS switch and 6 Transistors (6T) SRAM. Meanwhile, the novel block structures of the logic block, connection block, switch block, and the FPGA architecture based on the 1D2R element are proposed. Compared with the conventional 1T2R-based NVFPGA, our novel structure could improve the operation speed by 53% with a 40.5% lower operation power. Compared with the conventional 1R-based NVFPGA, the proposed scheme could greatly reduce the write error rate by eight orders with more than 20 times lower write power. Kejie Huang, Yong Lian 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2016 | Racetrack Memory-Based Nonvolatile Storage Elements for Multicontext FPGAsabstractA multicontext field-programmable gate array (FPGA) is a solution to achieve fast run-time reconfiguration. However, SRAM-based multicontext FPGAs still suffer from high leakage power during sleep, slow power-ON speed, and excessive large memory area. Racetrack memory is one of the most promising resistive nonvolatile memories, with the advantages of low power, high density, and high speed. In this paper, we propose two racetrack memory-based nonvolatile storage elements (NVSEs) for multicontext FPGAs. One is the shifting-based NVSE (type-1) with the advantages of high density and low power. The other one is the address-based NVSE (type-2) with the advantages of high context switching speed and low context switching power. The versatile place and route simulation results show that the type-1 NVSE-based eight-context FPGA reduces the area, critical path delay, and the power of the SRAM-based eight-context FPGA by more than 68.1%, 22.8%, and 13%, respectively. The proposed type-2 NVSE-based FPGAs allow the contexts to be switched 4.46 times faster than the type-1 NVSE-based FPGAs. Both designs improve the FPGA power-ON speed by more than a million times. Compared with the conventional racetrack memory-based lookup table (LUT), the proposed racetrack memory-based LUT may reduce the total power by more than 25%. Kejie Huang, Yong Lian 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Optimization Scheme to Minimize Reference Resistance Distribution of Spin-Transfer-Torque MRAMabstractSpin-transfer-torque magnetoresistive random access memory (STT-MRAM) is an emerging type of nonvolatile memory with compelling advantages in endurability, scalability, speed, and energy consumption. As the process technology shrinks, STT-MRAM has limited sensing margin due to the decrease in supply voltage and increase in process variation. Furthermore, the relatively smaller resistance difference of two states in STT-MRAM poses challenges for its read/write circuit design to maintain an acceptable sensing margin. The proposed reference circuits optimization scheme solves the reference resistance distribution issue to maximize the sensing margin and minimize the read disturbance, with low power consumption. Simulation results show that the optimization scheme is able to significantly improve the read reliability with the presence of one or few cases of reference cell failure, thus it eliminates the requirement of additional circuits for failure detection of reference cell or referencing to neighboring blocks. Kejie Huang, Ning Ning 0001, Yong Lian 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Artificial neuron with somatic and axonal computation units: Mathematical and neuromorphic models of persistent firing neuronsabstractThe conventional view of axon as a transmission cable has been challenged by continuous progress in neuroscience discoveries, which indicate the rich functional and computational repertoire of the axon. Recent experimental findings of slow integration induced persistent firing in distal axons of interneurons have shown that the slow integration from tens of seconds to minutes in distal axon leads to persistent firing of action potentials lasting for similar duration, suggesting that the axon performs its own integration functions. In this paper, we present an artificial neuron model including both somatic and axonal computation units, which reproduces the neural behavior of persistent firing. Complementary to the classic somatic computational unit which evokes action potentials by integrating dendritic inputs in a short timescale, the axon integrates the soma evoked spikes in a longer timescale of tens of second to minutes. Consequently the persistent firing behavior of the axon is determined through toggling the axon dynamics between passive conduction mode and persistent firing mode based on the integrated axonal potential. We present and discuss in this work the mathematical and neuromorphic models of the artificial neuron, as well as their simulation results. The artificial neuron proposed, being computationally efficient yet bio-plausible, would be useful to construct and simulate the large scale models of animal or human cortex, which provides a neuromorphic platform for further investigation of the possible functions of persistent firing and their roles in animal and human brain, especially their correlations with working memory. Ning Ning 0001, Kejie Huang, Luping Shi |
IJCNN | 2 |
| 2011 | Axonal Slow Integration Induced Persistent Firing Neuron Model
Ning Ning 0001, Kaijun Yi, Kejie Huang, Luping Shi |
ICONIP (1) | 3 |