Jiahao Cai

dblp:225/7886 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021Security and privacy · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FeFET-Based Analog In-Memory Computing With Inherent Shift-and-Add Capability
abstract
In-memory computing (IMC) architecture has emerged as a highly promising approach, enhancing the energy efficiency of multiply-and-accumulate (MAC) operations in deep neural networks (DNNs) by embedding parallel computations directly into memory arrays. However, existing ferroelectric FET (FeFET)-based analog IMC designs are often constrained to cell-level optimizations and struggle to achieve high-precision MAC operations. In contrast, high-precision analog IMC architectures typically perform MAC operations for partial inputs and weights within the array in a single cycle and then accumulate partial results over multiple cycles. During this procedure, circuits that handle weight shift-and-add process, whether in digital or analog form, incur significant overhead. This paper presents energy-efficient high-precision analog IMC designs leveraging FeFET technology, which inherently support a shift-and-add mechanism for weights. Initially, we introduce an IMC array paradigm that performs partial MAC operations within each column, and seamlessly incorporates the shift-and-add process for weights by utilizing the analog storage properties of FeFET-based cells. Building upon this paradigm, we propose single-level cell (SLC) FeFET-based designs, namely CurFe and ChgFe, operating in the current and charge modes, respectively. Additionally, to leverage FeFET’s multi-level cell (MLC) properties, we propose a novel hybrid SLC-MLC FeFET-based design, MulFe, which offers higher storage density and energy efficiency. Comprehensive evaluations are conducted at both the circuit and system levels, and the results indicate that the average energy efficiency of the proposed FeFET-based analog IMC designs is 1.32× to 2.71× higher compared to state-of-the-art (SOTA) IMC designs.
Qingrong Huang, Yu Qian 0002, Jiahao Cai, Kai Ni 0004, Thomas Kämpfe, Zheyu Yan, Xunzhao Yin, Cheng Zhuo
IEEE Trans. Computers4
2025 Automatic Diagnosis of Hip and Knee Osteoarthritis from Medical Text Records
abstract
Osteoarthritis (OA) is a highly frequent musculoskeletal condition defined by the progressive degradation of joint cartilage and the underlying bone, resulting in the manifestation of pain and functional limitations. Identification of symptoms as early as possible for timely intervention is critical for effective pain management and treatment. Knee and hip OA are common in older patients which can greatly affect their mobility, lifestyle, and lead to other health complications. Electronic Medical Records (EMR) in primary care settings contain patients’ structured historical data including unstructured text data in the encounter chart notes. The unstructured notes are often very long, compiled from multiple patient-physician encounters, and contain medical jargon including personal data. The data offers a variety of computational challenges but it contains valuable information for disease diagnosis especially for detecting hip or knee OA. We demonstrate multiple keyword-based strategies to detect the OA-affected bone joints, including a simple rule-based approach and a machine learning based approach. We also provide an ablation study to show the effects of the different natural language text processing methods and validate our results against gold standard data labelled by a human expert. Our Random Forest (RF) model achieved the best result of 74.89% F1-score with OA related paragraph extraction.
Jiahao Cai, Vidhi Kokel, Farhana Zulkernine, John A. Queenan, David Barber
COMPSAC1
2025 FACAM: Design and Optimization of A Compact Energy Efficient FeFET-Based Analog Content Addressable Memory
abstract
Content Addressable Memory (CAM) is known for highly parallel pattern matching capability, which is widely used for data-centric applications and advanced machine learning models that involve associative search tasks. However, most state-of-the-art CAM designs focus on binary/multi-bit CAMs (B/MCAMs) based on CMOS or emerging nonvolatile memories (NVMs), which struggle in scenarios where analog values, rather than discrete levels, need to be stored and searched. Therefore, analog CAMs (ACAMs) offer a promising solution to further increase memory density, improve energy efficiency and extend practical scenarios. Among NVMs, ferroelectric field effect transistors (FeFETs) have emerged as a strong candidate for efficient CAM designs due to the three-terminal structure, high on-off ratio, high OFF resistance and voltage-driven write/read mechanisms. In this paper, we propose FACAM, a compact and energy efficient single-input 2FeFET-1T ACAM cell design, with a two-phase search scheme, that sets the location and width of the matching range through two FeFETs, respectively. We further present a FACAM array which reduces the matchline (ML) voltage swing by shifting ML precharging into the in-cell search operations. We also propose an adaptive scheme to selectively early-terminate second search phase for further search energy optimization. Evaluation results suggest that our proposed FACAM achieves 8.39× and 2.94× energy efficiency compared with the state-of-the-art better 6T-2R ACAM and 2FeFET ACAM. Benchmarking results in deep random forest accelerator show that our approach is 2.14× faster and 7.82× energy efficient than 2FeFET ACAM.
Jiahao Cai, Ann Franchesca Laguna, Thomas Kämpfe, Zheyu Yan, Cheng Zhuo, Xunzhao Yin
ICCAD1
2025 A Scalable 2T-1FeFET-Based Content Addressable Memory Design for Energy Efficient Data Search
abstract
Content addressable memory (CAM) is widely used in advanced machine learning models and data-intensive applications for associative search tasks, thanks to the highly parallel pattern matching capability. Most state-of-the-art CAM designs primarily aim to reduce the CAM cell area by utilizing nonvolatile memories (NVMs). However, there has been limited research on optimizing the design and energy efficiency of NVM-based CAMs for practical deployment in edge devices and AI hardware. This article introduces a general compact and energy efficient CAM design scheme that minimizes design overhead by using only one NVM device per cell. Our proposed CAM design realizes both binary CAM (BCAM) and multibit CAM (MCAM) by leveraging the binary and multilevel storage property of NVM devices without additional cell overheads. Additionally, we propose an adaptive matchline (ML) precharge and discharge scheme to further optimize search energy by significantly reducing the ML voltage swing. Ferroelectric field-effect transistors (FeFETs) serve as representative NVMs in our proposed design, and we present a 2T-1FeFET CAM array incorporating a sense amplifier that implements the proposed ML scheme. Evaluation results show that our proposed 2T-1FeFET BCAM design achieves energy efficiency improvements of$6.64\times $/$4.74\times $/$9.14\times $/$3.02\times $compared to CMOS/ReRAM/STT-MRAM/2FeFET BCAM arrays, while 2T-1FeFET MCAM design achieves$8.25\times $/$5.68\times $/$56.35\times $better-energy efficiency compared to ReRAM/3T-1FeFET/1FeFET-1R MACM arrays. Benchmarking results demonstrate that our BCAM/MCAM approach provides$3.2\times $/$3.7\times $and$2.0\times $/$2.2\times $energy-delay product improvement over the 2T-2R and 2FeFET CAM in accelerating query processing applications.
Jiahao Cai, Hamza Errahmouni Barkam, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2024 Assist-Bot: A Voice-Enabled Assistant for Seniors
abstract
With the increasing elderly population, healthcare systems are facing a crisis in providing 24/7 assistance to seniors. Voice-enabled robots can serve as assistants and companions for seniors. Many voice and chat systems, such as Alexa, Siri, and Google, have emerged and become the top consumer brands. However, a significant gap still exists in making these systems mimic human-like conversation. Most of these systems lack empathy and cannot understand users' expressions and intent in delivering information as expected by users. Studies have shown that seniors struggle to effectively communicate with the existing voice conversation systems, leading to a stigma around technology acceptance. In this pilot project, we explore existing literature to identify the key challenges and the types of care that seniors need, and present a personalized voice communication system that can remind about tasks such as taking medication and doing health checks; answer predefined questions about personalized routines and medications; launch specific applications such as the contactless wellness monitoring application “Veyetals”, “Uber” or a music plugin, and answer basic questions from a pre-defined Frequently Asked Questions (FAQ) database. We extend existing voice bots to develop the Assist-Bot to serve as a voice-enabled assistant for nontechnical seniors. Caregivers can configure medication and check-up schedules, and contact numbers for different needs.
Jiahao Cai, Farhana Zulkernine, Nauman Jaffar, Amina Al-Marzouqi, Nabeel Al-Yateem, Syed Azizur Rahman
COMPSAC2
2024 Bridging the Gap: A Self-Learning Model Using Implicit Knowledge for Chinese Spelling Correction
abstract
Chinese Spelling Correction (CSC) is a challenging and essential task in natural language processing. In this study, we introduces a new method for Chinese Spelling Correction (CSC) that addresses three unattended areas in prior studies. Firstly, we use an Implicit Knowledge Extraction Network to overcome limitations of conventional methods that rely on explicit knowledge alone. Secondly, we use KL divergence to limit the effect of incorrect characters on semantic understanding, ensuring consistent meaning. Finally, we employ a Cor-Det framework rather than the traditional Det-Cor framework, offering more consistent learning objectives. Tests on three SIGHAN benchmarks show this method significantly surpassing baseline models, highlighting the crucial role of implicit knowledge in Chinese Spelling Correction tasks.
Wenyao Cui, Jiahao Cai, Baohua Zhang 0002, Yongyi Huang, Huaping Zhang
ICASSP2
2024 STMAP: A novel semantic text matching model augmented with embedding perturbations
abstract
Semantic text matching models have achieved outstanding performance, but traditional methods may not solve Few-shot learning problems and data augmentation techniques could suffer from semantic deviation. To solve this problem, we propose STMAP, which is implemented from the perspective of data augmentation based on Gaussian noise and Noise Mask signal. We also employ an adaptive optimization network to dynamically optimize the several training targets generated by data augmentation. We evaluated our model on four English datasets: MRPC, SciTail, SICK, and RTE, with achieved scores of 90.3%, 94.2%, 88.9%, and 68.8%, respectively. Our model obtained state-of-the-art (SOTA) results on three of the English datasets. Furthermore, we assessed our approach on three Chinese datasets, and achieved an average improvement of 1.3% over the baseline model. Additionally, in the Few-shot learning experiment, our model outperformed the baseline performance by 5%, especially when the data volume was reduced by around 0.4. Our ablation experiments further validated the effectiveness of STMAP.2
Baohua Zhang 0002, Weikang Liu, Jiahao Cai, Huaping Zhang
Inf. Process. Manag.4
2023 VisPhone: Chinese named entity recognition model enhanced by visual and phonetic features
abstract
Many Chinese NER models only focus on lexical and radical information, ignoring the fact that there are also certain rules for the pronunciation of Chinese entities. In this paper, we propose VisPhone, which incorporates Chinese characters’ Phonetic features into Transformer Encoder along with the Lattice and Visual features. We present the common rules for the pronunciation of Chinese entities and explore the most appropriate method to encode it. VisPhone uses two identical cross transformer encoders to fuse the visual and phonetic features of the input characters with the text embedding. A selective fusion module is used to get the final features. We conducted experiments on four well-known Chinese NER benchmark datasets: OntoNotes4.0, MSRA, Resume, and Weibo, with F1 scores of 82.63%, 96.07%, 96.26%, 70.79% respectively, improving the performance by 0.79%, 0.32%, 0.39%, and 3.47%. Our ablation experiments have also demonstrated the effectiveness of VisPhone.
Baohua Zhang 0002, Jiahao Cai, Huaping Zhang, Jianyun Shang
Inf. Process. Manag.2
2022 Energy efficient data search design and optimization based on a compact ferroelectric FET content addressable memory
abstract
Content Addressable Memory (CAM) is widely used for associative search tasks in advanced machine learning models and data-intensive applications due to the highly parallel pattern matching capability. Most state-of-the-art CAM designs focus on reducing the CAM cell area by exploiting the nonvolatile memories (NVMs). There exists only little research on optimizing the design and energy efficiency of NVM based CAMs for practical deployment in edge devices and AI hardware. In this paper, we propose a general compact and energy efficient CAM design scheme that alleviates the design overhead by employing just one NVM device in the cell. We also propose an adaptive matchline (ML) precharge and discharge scheme that further optimizes the search energy by fully reducing the ML voltage swing. We consider Ferroelectric field effect transistors (FeFETs) as the representative NVM, and present a 2T-1FeFET CAM array including a sense amplifier implementing the proposed ML scheme. Evaluation results suggest that our proposed 2T-1FeFET CAM design achieves 6.64×/4.74×/9.14×/3.02× better energy efficiency compared with CMOS/ReRAM/STT-MRAM/2FeFET CAM arrays. Benchmarking results show that our approach provides 3.3×/2.1× energy-delay product improvement over the 2T-2R/2FeFET CAM in accelerating query processing applications.
Jiahao Cai, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin
DAC1
2019 Node Influence Calculation of Novel Social Networks
abstract
Identifying Influential nodes in complex networks is of great significance for network structure optimization and robustness enhancement. To measure the role of nodes, many classic methods could help identify the influential nodes. In this paper, we take the simple complex network composed of Jin Yong's novels as an example, applying the classic measurement methods for key nodes to analyze them. In the introduction section, we will briefly introduce background knowledge of complex networks, and measurement methods for key nodes in complex networks. In experiments, we will analyze each measurement method in detail for influential node as well as sum up the advantage and disadvantage of each method.
Jiahao Cai, Minyong Shi
ICIS1
2019 Zero-sum Distinguishers for Round-reduced GIMLI Permutation
abstract
GIMLI is a 384-bit permutation proposed by Bernstein et al. at CHES 2017. It is designed with the goal of achieving both high security and high performance across a wide range of hardware and software platforms. Since GIMLI can be used as a building block for many cryptographic schemes, it is important to understand its concrete security. To the best of our knowledge, third party cryptanalysis of GIMLI is limited. In this paper, we identify some zero-sum distinguishers for 14-round GIMLI with the inside-out technique, which are one-round longer than the integral distinguishers presented by the designers. Although we obtain improved cryptanalysis results, these zero-sum distinguishers are far from threatening the full version of GIMLI.
Jiahao Cai, Zihao Wei, Siwei Sun, Lei Hu 0003
ICISSP1
2018 Speeding up MILP Aided Differential Characteristic Search with Matsui's Strategy
Siwei Sun, Jiahao Cai, Lei Hu 0003
ISC3