Jialin Cao

dblp:58/4526 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Enhancing Aspect-Based Sentiment Analysis via Augmented Semantic and Syntactic Graph Fusion
abstract
ABSTRACT Recommendation systems are rapidly evolving from static interaction‐driven models to dynamic, knowledge‐augmented architectures. A key challenge in this evolution is accurately capturing users' fine‐grained preferences from unstructured review text, which directly impacts the explainability and personalization of recommendations. As an essential enabling technology, Aspect‐Based Sentiment Analysis (ABSA) extracts aspect‐level sentiment elements that can be explicitly mapped to user preference vectors or product attribute ratings. With the integration of semantic and syntactic information, current works have significantly enhanced the performance of ABSA. However, existing graph‐based approaches that rely on dependency‐tree structures often converge to suboptimal solutions when handling implicit sentiment in natural language. To address this gap, we propose a graph fusion network that leverages augmented semantic and syntactic graphs. Specifically, we explicitly model word‐dependency correlations via contextual augmentation, and incorporate selected part‐of‐speech (POS) features to refine semantic graph construction. Concurrently, a syntactic graph is constructed by pruning the nodes based on the distance to the aspect term. The resulting semantic and syntactic representations are then fused through a dual graph convolutional network block, whereas the gating mechanism is used to regulate information flow during graph construction. Experiments on seven benchmarks demonstrate that our approach outperforms baselines by up to and in Macro‐F1 scores, establishing new state‐of‐the‐art results and providing a more reliable sentiment extraction module for downstream recommendation tasks.
Zhiyuan Ma 0001, Yuze Wang, Jialin Cao, Nan Wang 0003
Expert Syst. J. Knowl. Eng.4
2026 Towards anti-forgetting with masked optimal transport regularization for continual named entity recognition
Zhiyuan Ma 0001, Miaomiao Gu, Nan Wang 0003, Jialin Cao
Neurocomputing4
2024 FLAME: Fully Leveraging MoE Sparsity for Transformer on FPGA
abstract
MoE (Mixture-of-Experts) mechanism has been widely adopted in transformer-based models to facilitate further expansion of model parameter size and enhance generalization capabilities. However, the practical deployment of MoE mechanism for transformer on resource-constrained platforms, such as FPGA, remains challenging due to heavy memory footprints and impractical runtime costs introduced by the MoE mechanism. Diving into the MoE mechanism, we raise two key observations: (1) Expert weights are heavy but cold, making it ideal to leverage expert weight sparsity. (2) There exists highly skewed expert activation paths for MoE layers in transformer-based models, making it feasible to conduct expert prediction and prefetching. Motivated by these two observations, we propose FLAME, the first algorithm-hardware co-optimized MoE accelerating framework designed to fully leverage MoE sparsity for efficient transformer deployment on FPGA. First, to leverage expert weight sparsity, we integrate an N:M pruning algorithm, allowing for the pruning of expert weights without significantly compromising model accuracy. Second, to settle expert activation sparsity, we propose a circular expert prediction (CEPR) strategy. CEPR prefetches expert weights from external storage to on-chip cache before the activated expert index is determined. Last, we co-optimize both MoE sparsity through the introduction of an efficient pruning-aware expert buffering (PA-BUF) mechanism. Experimental results demonstrate that FLAME achieves 84.4% accuracy of expert prediction with merely two expert caches on-chip. In comparison with CPU and GPU, FLAME achieves 4.12× and 1.49× speedup, respectively.
Xuanda Lin, Huinan Tian, Wenxiao Xue, Lanqi Ma, Jialin Cao, Manting Zhang, Jun Yu 0010, Kun Wang 0005
DAC5
2024 FNM-Trans: Efficient FPGA-based Transformer Architecture with Full N: M Sparsity
abstract
Transformer models have become popular in various AI applications due to their exceptional performance. However, their impressive performance comes with significant computing and memory costs, hindering efficient deployment of Transformer-based applications. Many solutions focus on leveraging sparsity in weight matrix and attention computation. However, previous studies fail to exploit unified sparse pattern to accelerate all three modules of Transformer (QKV generation, attention computation and FFN). In this paper, we propose FNM-Trans, an adaptable and efficient algorithm-hardware co-design aimed at optimizing all three modules of the Transformer by fully harnessing N : M sparsity. At the algorithm level, we fully explore the interplay of dynamic pruning with static pruning under high N : M sparsity. At the hardware level, we develop a dedicated hardware architecture featuring a custom computing engine and a softmax module, tailored to support varying levels of N : M sparsity. Experiment results show that, our algorithm optimizes accuracy by 11.03% under 2:16 attention sparsity and 4:16 weight sparsity, compared to other methods. Additionally, FNM-Trans achieves speedups of 27.13× and 21.24× over Intel i9-9900X and NVIDIA RTX 2080 Ti, respectively, and outpaces current FPGA-based Transformers by 1.88× to 36.51×.
Manting Zhang, Jialin Cao, Kejia Shi, Keqing Zhao, Genhao Zhang, Jun Yu 0010, Kun Wang 0005
DAC2
2023 Token Packing for Transformers with Variable-Length Inputs
abstract
Transformer-based models has achieved remarkable success in extensive tasks for natural language processing. To face the variable-length sentences in human language, popular deep learning frameworks rely on zero padding for batch processing, which introduces significant computation and memory overhead. Existing works attempt to eliminate padding redundancy but results in low hardware efficiency due to the mismatch between the variable shape of operations and fixed shape of processing elements (PEs). This paper proposes a reconfigurable systolic array with token packing in three folds to boost hardware efficiency. First, matrix multiplications for different tokens can be packed along the array columns to improve spatial efficiency. Meanwhile, for temporal efficiency, we develop a coarse-grained pipeline for attention, where stages can run on different parts of the array at the same time. We further exploit the masking redundancy in the Transformer decoder with runtime reconfigurable inter-PE connection and buffer switching. Applied to GPT, our FPGA design has achieved 1.16× higher normalized throughput and 1.94× better runtime MAC utilization over the state-of-the-art GPU performance for variable-length input sequences from GLUE and SQuAD dataset.
Tiandong Zhao, Siyuan Miao, Shaoqiang Lu, Jialin Cao, Xiao Shi 0001, Kun Wang 0005, Lei He 0001
FPL4
2023 g-BERT: Enabling Green BERT Deployment on FPGA via Hardware-Aware Hybrid Pruning
abstract
Transformer-based models suffer from large num-ber of parameters and high inference latency, whose deployment are not green due to the potential environmental damage caused by high inference energy consumption. In addition, it is difficult to deploy such models on devices, especially on resource constrained devices such as FPGA. Various model pruning methods are proposed to shrink the model size and resource consumption, so as to fit the models on hardware. However, such methods often introduce floating point of operations (FLOPs) as an agent of hardware performance, which is not accurate. Furthermore, structural pruning methods are always in a single head-wise or layer-wise pattern, which fails to compress the models to the extreme. To resolve the above issues, we propose a green BERT deployment method on FPGA via hardware-aware and hybrid pruning, named g-BERT. Specifically, two hardware-aware metrics are introduced by High Level Synthesis (HLS) to evaluate the latency and power consumption of inference on FPGA, which can be optimized directly while pruning. Moreover, we simultaneously consider pruning of heads and full encoder layers. To efficiently find the optimal structure, g-BERT applies differentiable neural architecture search (NAS) with a special 0–1 loss function. Compared with the BERT-base, g-BERT achieves$2.1\times$speedup,$1.9\times$power consumption reduction and$1.8\times$model size reduction with comparable accuracy, on par with the state-of-the-art methods.
Yueyin Bai, Hao Zhou 0008, Ruiqi Chen 0001, Kuangjie Zou, Jialin Cao, Jianli Chen, Jun Yu 0010, Kun Wang 0005
ICC5
2023 PP-Transformer: Enable Efficient Deployment of Transformers Through Pattern Pruning
abstract
Transformer models have been widely adopted in the field of Natural Language Processing (NLP) and Computer Vision (CV). However, the excellent performance of Transformers comes at the cost of heavy memory footprints and gigantic computing complexity. To deploy Transformers on resource constrained platforms, e.g., FPGA, diverse weight pruning strategies have been proposed. However, pattern pruning, as an alternative pruning method, is not well explored in the context of Transformers. In this paper, we propose PP-Transformer, a framework specifically designed to efficiently deploy Transformer models on FPGA using pattern pruning. At the algorithm level, we leverage pattern pruning, a coarse-grained structured pruning strategy, to reduce parameter storage. Meanwhile, we have developed a dedicated hardware architecture, featuring a custom computing engine tailored to support pattern pruning algorithm. Experimental results demonstrate that our algorithm achieves up to$2.26\times$reduction in parameter storage with acceptable accuracy degradation. Additionally, our hardware implementation exhibits$839.72\times$and$5.72\times$speedup in comparison to CPU and GPU implementations.
Jialin Cao, Xuanda Lin, Manting Zhang, Kejia Shi, Jun Yu 0010, Kun Wang 0005
ICCAD1
2012 A 60mW baseband SoC for CMMB receiver
abstract
This paper describes baseband SoC implementation of China Mobile Multimedia Broadcasting (CMMB) receiver, which integrates analog to digital (ADC), physical layer (PHY) baseband processor and medium access control (MAC) processor in single silicon wafer. MAC functions are fully implemented by firmware on an embedded 32-bit RISC-based processor. In addition, several power management techniques are utilized to reduce the power consumption of baseband SoC. The baseband SoC was successfully fabricated in 0.13µm one-poly six-metal (1P6M) CMOS process. Both analog and digital circuits are integrated on 4.8×4.8 mm2die consuming 60mW total power dissipation under 1.2V and 3.3V supplies. The experiment results reveal the proposed baseband SoC has excellent performance under the multipath channels.
Jialin Cao, Dan Bao, Yun Chen 0001, Xiaoyang Zeng
ASP-DAC2
2006 A Design of Pipelined Carry-dependent Sum Adder With its Self-checking Structure
abstract
In this paper a pipelined carry-dependent sum adder with the self-checking structure is proposed. The adder includes four 8-bit carry-dependent sum adder (CDSA) , a 4-bit block carry look-ahead unit (BCLU) and a parity checker. The necessary area of the proposed adder is only about 3.85% over the traditional ripple carry adders, while the sum of the traditional adders is delayed by 39.2% with respect to the proposed adder for 32-bit implementation
Shiyi Xu, Jialin Cao, Feng Ran, Shiwei Ma
ATS3
2006 Algorithm Analysis and Application Based on Chaotic Neural Network for Cellular Channel Assignment
Xiaojin Zhu 0002, Yanchun Chen, Jialin Cao
ICIC (1)4