Ruibin Mao

dblp:253/9848 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-6085-0486ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 EvaCAM: A Circuit-Level Evaluation Tool for General Content Addressable Memories
abstract
Content addressable memories (CAMs) are special-purpose in-memory computing units that support parallel searches directly in memory. There is growing interest in CAMs for data-intensive applications such as machine learning, data mining, and bioinformatics, which has led to a rapidly growing CAM design space. CAM cells can be implemented exclusively by CMOS or with various non-volatile memory (NVM) devices. In addition to traditional binary and ternary CAMs (BCAMs and TCAMs), analog CAM (ACAM) and multi-bit CAM (MCAM) designs have recently been introduced, which could further improve density, and also support unique in-memory distance functions. Furthermore, aside from the widely-used exact match function, CAM-based approximate match functions, such as threshold match and best match, have been proposed to further extend the utility of CAMs to new application spaces. As the CAM design space is large, evaluating different CAM design options for a given application is both crucial and challenging. This paper presents EvaCAM, a circuit-level modeling and evaluation tool for CAMs. EvaCAM supports TCAM, ACAM, and MCAM designs implemented in either CMOS or NVMs, for both exact and approximate match functions. It also allows for the exploration of different CAM designs under various optimization targets. EvaCAM has been validated against measured data from fabricated chips and detailed SPICE simulations. A comprehensive design space exploration for CAMs is provided to illustrate the impact of various design decisions and to demonstrate the use cases of EvaCAM.
Liu Liu 0023, Mohammad Mehdi Sharifi, Kunshi Wang, Ruibin Mao, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2025 High-Performance In-Memory Bayesian Inference With Multi-Bit Ferroelectric FET
abstract
Conventional neural network-based machine learning algorithms often encounter difficulties in data-limited scenarios or where interpretability is critical. Conversely, Bayesian inference-based models excel with reliable uncertainty estimates and explainable predictions. Recently, many in-memory computing (IMC) architectures achieve exceptional computing capacity and efficiency for neural network tasks leveraging emerging nonvolatile memory (NVM) technologies. However, their application in Bayesian inference remains limited because the operations in Bayesian inference differ substantially from those in neural networks. In this article, we introduce a compact in-memory Bayesian inference engine with high efficiency and performance utilizing a multi-bit ferroelectric field-effect transistor (FeFET). This design encodes a Bayesian model within a compact FeFETbased crossbar by mapping quantized probabilities to discrete FeFET states. Consequently, the crossbar’s outputs naturally represent the output posteriors of the Bayesian model. Our design facilitates efficient Bayesian inference, accommodating various input types and probability precisions, without additional calculation circuitry. As the first FeFET-based in-memory Bayesian inference engine, our design demonstrates a notable storage density of 26.32 Mb/mm2and a computing efficiency of 581.40 TOPS/W in a representative Bayesian classification task, indicating a 10.7×/43.4× compactness/efficiency improvement compared to the state-of-the-art alternative. Utilizing the proposed Bayesian inference engine, we develop a feature selection system that efficiently addresses a representative NP-hard optimization problem, showcasing our design’s capability and potential to enhance various Bayesian inference-based applications. Test results suggest that our design identifies the essential features, enhancing the model’s performance while reducing its complexity, surpassing the latest implementation in operation speed and algorithm efficiency by 2.9×/2.0×, respectively.
Chao Li 0065, Xuchu Huang, Ruibin Mao, Thomas Kämpfe, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Cheng Zhuo
IEEE Trans. Computers5
2024 FeBiM: Efficient and Compact Bayesian Inference Engine Empowered with Ferroelectric In-Memory Computing
abstract
In scenarios with limited training data or where explainability is crucial, conventional neural network-based machine learning models often face challenges. In contrast, Bayesian inference-based algorithms excel in providing interpretable predictions and reliable uncertainty estimation in these scenarios. While many state-of-the-art in-memory computing (IMC) architectures leverage emerging non-volatile memory (NVM) technologies to offer unparalleled computing capacity and energy efficiency for neural network workloads, their application in Bayesian inference is limited. This is because the core operations in Bayesian inference, i.e., cumulative multiplications of prior and likelihood probabilities, differ significantly from the multiplication-accumulation (MAC) operations common in neural networks, rendering them generally unsuitable for direct implementation in most existing IMC designs. In this paper, we propose FeBiM, an efficient and compact Bayesian inference engine powered by multi-bit ferroelectric field-effect transistor (FeFET)-based IMC. FeBiM effectively encodes the trained probabilities of a Bayesian inference model within a compact FeFET-based crossbar. It maps quantized logarithmic probabilities to discrete FeFET states. As a result, the accumulated outputs of the crossbar naturally represent the posterior probabilities, i.e., the Bayesian inference model's output given a set of observations. This approach enables efficient in-memory Bayesian inference without the need for additional calculation circuitry. As the first FeFET-based in-memory Bayesian inference engine, FeBiM achieves an impressive storage density of 26.32 Mb/mm2 and a computing efficiency of 581.40 TOPS/W in a representative Bayesian classification task. These results demonstrate 10.7×/43.4× improvement in compactness/efficiency compared to the state-of-the-art hardware implementation of Bayesian inference.
Chao Li 0065, Ruibin Mao, Can Li 0024, Thomas Kämpfe, Kai Ni 0004, Xunzhao Yin
DAC4
2024 FeReX: A Reconfigurable Design of Multi-Bit Ferroelectric Compute-in-Memory for Nearest Neighbor Search
abstract
Rapid advancements in artificial intelligence have given rise to transformative models, profoundly impacting our lives. These models demand massive volumes of data to operate effectively, exacerbating the data-transfer bottleneck inherent in the conventional von-Neumann architecture. Compute-in-memory (CIM), a novel computing paradigm, tackles these issues by seam-lessly embedding in-memory search functions, thereby obviating the need for data transfers. However, existing non-volatile memory (NVM)-based accelerators are application specific. During the similarity based associative search operation, they only support a single, specific distance metric, such as Hamming, Manhattan, or Euclidean distance in measuring the query against the stored data, calling for reconfigurable in-memory solutions adaptable to various applications. To overcome such a limitation, in this paper, we present FeReX, a reconfigurable associative memory (AM) that accommodates various distance metrics including Hamming, Manhattan, and Euclidean distances. Leveraging multi-bit ferroelectric field-effect transistors (FeFETs) as the proxy and a hardware-software co-design approach, we introduce a constrained satisfaction problem (CSP)-based method to automate AM search input voltage and stored voltage configurations for different distance based search functions. Device-circuit co-simulations first validate the effectiveness of the proposed FeReX methodology for reconfigurable search distance functions. Then, we benchmark FeReX in the context of k-nearest neighbor (KNN) and hyperdimensional computing (HDC), which highlights the robustness of FeReX and demonstrates up to 250× speedup and 104energy savings compared with GPU.
Che-Kai Liu, Chao Li 0065, Ruibin Mao, Jianyi Yang 0003, Thomas Kämpfe, Mohsen Imani, Can Li 0024, Cheng Zhuo, Xunzhao Yin
DATE4
2024 ShiftCAM: A Time-Domain Content Addressable Memory Utilizing Shifted Hamming Distance for Robust Genome Analysis
abstract
Fast and efficient genome analysis can have a significant impact in areas such as scientific discovery and personalized medicine. Given the extensive data produced by sequencing machines, in-memory computing is considered a strong candidate to tackle the frequent data movement issue. Previous research has introduced many designs based on Content Addressable Memories (CAM), mainly optimized for tolerating edit distance; however, these systems struggle when there are a few insertion or deletion errors. This limitation presents a significant challenge for genome analysis, as current Third-Generation Sequencing still has high error rates. In this work, we introduce ShiftCAM, a time-domain Content Addressable Memory, designed to accommodate the high error rates in practical scenarios. Utilizing time-domain comparison, ShiftCAM effectively calculates the Shifted Hamming Distance to better approximate the computationally expensive edit distance. Additionally, the Modification to Accidental Match strategy specially designed for hardware implementation is introduced to eliminate accidental matches of single base pairs, further reducing false positives and improving edit distance approximation. Monte Carlo simulations based on physical ReRAM device statistical measurements and commercial PDK are also conducted to validate the robustness of the ShiftCAM design. Our experiments demonstrate that ShiftCAM can achieve an average of 2.1× (from 40.1% to 83.8%) higher F1 score in contamination analysis, 21.3% estimation error in relative abundance analysis, 51.2% reduction in cell area, 29.5× speed up, and 9.4× higher energy efficiency, compared to state-of-the-art in-memory DNA classification accelerators.
Peiyi He, Ruibin Mao, Keyi Shan, Yunwei Tong, Muyuan Peng, Ruibang Luo, Can Li 0024
ICCAD2
2023 ReRAM-based graph attention network with node-centric edge searching and hamming similarity
abstract
The graph attention network (GAT) has demonstrated its advantages via local attention mechanism but suffered from low energy and latency efficiency when implemented on conventional von-Neumann hardware. This work proposes and experimentally demonstrates an algorithm-hardware co-designed GAT that runs efficiently and reliably in ReRAM-based hardware. The neighborhood information is retrieved from trained node embeddings stored on crossbars in a single time step, and attention is implemented by efficient hashing and hamming similarity for higher robustness. Our scaled simulation based on the experimentally-validated model shows only 0.9% accuracy loss with over 35,500x energy improvement on the Cora dataset compared with GPU, and 1.1% accuracy improvement with 2× energy improvement compared with state-of-the-art ReRAM-based GNN accelerator.
Ruibin Mao, Xia Sheng, Catherine Graves, Can Li 0024
DAC1
2023 YOLO-table: disclosure document table detection with involution
Daqian Zhang, Ruibin Mao, Runting Guo
Int. J. Document Anal. Recognit.2
2020 The Design and Construction of a Chinese Sarcasm Dataset
abstract
As a typical multi-layered semi-conscious language phenomenon, sarcasm is widely existed in social media text for enhancing the emotion expression. Thus, the detection and processing of sarcasm is important to social media analysis. However, most existing sarcasm dataset are in English and there is still a lack of authoritative Chinese sarcasm dataset. In this paper, we presents the design and construction of a largest high-quality Chinese sarcasm dataset, which contains 2,486 manual annotated sarcastic texts and 89,296 non-sarcastic texts. Furthermore, a balanced dataset through elaborately sampling the same amount non-sarcastic texts for training sarcasm classifier. Using the dataset as the benchmark, some sarcasm classification methods are evaluated.
Xiaochang Gong, Ruibin Mao
LREC4
2020 Target-based Sentiment Annotation in Chinese Financial News
abstract
This paper presents the design and construction of a large-scale target-based sentiment annotation corpus on Chinese financial news text. Different from the most existing paragraph/document-based annotation corpus, in this study, target-based fine-grained sentiment annotation is performed. The companies, brands and other financial entities are regarded as the targets. The clause reflecting the profitability, loss or other business status of financial entities is regarded as the sentiment expression for determining the polarity. Based on high quality annotation guideline and effective quality control strategy, a corpus with 8,314 target-level sentiment annotation is constructed on 6,336 paragraphs from Chinese financial news text. Based on this corpus, several state-of-the-art sentiment analysis models are evaluated.
Chaofa Yuan, Rongdi Yin, Qinling Zhu, Ruibin Mao
LREC6
2019 A Knowledge Regularized Hierarchical Approach for Emotion Cause Analysis
abstract
Chuang Fan, Hongyu Yan, Jiachen Du, Lin Gui, Lidong Bing, Min Yang, Ruifeng Xu, Ruibin Mao. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Chuang Fan, Hongyu Yan, Jiachen Du, Lin Gui 0003, Lidong Bing, Min Yang 0007, Ruifeng Xu 0001, Ruibin Mao
EMNLP/IJCNLP (1)8