Jiale Yan

dblp:228/2666 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Bringing Near Data Processing Into the Low-Bit Floating-Point Era
Tongxin Xie, Yuechen Xi, Bing Li 0017, Mo Guang, Jiale Yan, Kaiwen Long, Xingcheng Zhang, Huazhong Yang, Yuan Xie 0001
ISCA8
2026 PyPIMalDet: A malicious PyPI package detection method combining code features and metadata features
Jiale Yan, Bo Zhao 0023
Neural Networks1
2025 Stealthy Backdoor Attack against Video Recognition Models
abstract
Deep neural networks have achieved great success in various domains, however, it is fragile against backdoor attacks. Currently, the backdoor attack method in video recognition mainly comes from the expansion of the field of image classification, it faces the challenge that the trigger can be perceived by human eyes and easily captured by the defense methods in the existing image field. To solve the above challenges, in this paper, we propose a Stealthy Backdoor Attack method against Video Action Recognition model (SBAVAR), which improves the imperceptibility and effectiveness of backdoor triggers, enhancing their threat. Specifically, we formalize backdoor trigger generation as an optimization problem and utilize a feasible optimization scheme that integrates sparse spatial transformation perturbation and additive perturbation to generate a trigger, ensuring the imperceptibility and effectiveness of backdoor attacks. Our intensive experiments show that SBAVAR can achieve superior imperceptibility and effectiveness compared with baseline method. Additionally, SBAVAR can circumvent the existing backdoor defense methods.
Jiale Yan
ICASSP1
2025 BingoGCN: Towards Scalable and Efficient GNN Acceleration with Fine-Grained Partitioning and SLT
abstract
Graph Neural Networks (GNNs) are increasingly popular due to their wide applicability to tasks requiring the understanding of unstructured graph data, such as those in social network analysis and autonomous driving.However, real-time, large-scale GNN inference faces challenges due to the large size of node features and adjacency matrices, leading to memory communication and buffer size overheads caused by irregular memory access patterns.While graph partitioning can help with localized access patterns and reduction in on-chip buffer size, fine-grained partitioning results in increased inter-partition edges and off-chip memory accesses, negatively impacting overall performance.To overcome these limitations, we propose BingoGCN, a scalable GNN acceleration framework that introduces multidimensional dynamic feature summarization called Cross-Partition Message Quantization (CMQ) for inter-partition message passing.This eliminates irregular off-chip memory access without additional training and accuracy loss, even with fine-grained partitioning.By shifting the bottleneck from memory to computation, BingoGCN allows for further performance optimization through the Strong Lottery Ticket (SLT) theory using randomly generated weights.BingoGCN addresses the challenge of SLT's unstructured sparsity in hardware acceleration with a novel training algorithm and random weight generator designs, enabling fine-grained (FG) sparsity and improved load balancing.We integrated CMQ and FG-SLT into the messagepassing of GNNs and designed an efficient hardware architecture to support this flow.Our FPGA-based implementation achieves a significant reduction in memory accesses while preserving accuracy comparable to the original models.
Jiale Yan, Hiroaki Ito, Yuta Nagahara, Kazushi Kawamura, Masato Motomura, Thiem Van Chu, Daichi Fujiki
ISCA1
2025 EUN: Enhanced unlearnable examples generation approach for privacy protection
Xiaotian Chen, Sicong Zhang, Jiale Yan, Weida Xu, Xinlong He
Comput. Vis. Image Underst.4
2025 DMSA: An Efficient Architecture for Sparse-Sparse Matrix Multiplication Based on Distribute-Merge Product Dataflow
abstract
The sparse–sparse matrix multiplication (SpMSpM) is a fundamental operation in various applications. Existing SpMSpM accelerators based on inner product (IP) and outer product (OP) suffer from low computational efficiency and high memory traffic due to inefficient index matching and merging overheads. Gustavson’s product (GP)-based accelerators mitigate some of these challenges but struggle with workload imbalance and irregular memory access patterns, limiting computational parallelism. To overcome these limitations, we propose a distribute-merge product (DMP), a novel SpMSpM dataflow that evenly distributes workloads across multiple computation streams and merges partial results efficiently. We design and implement DMP-based SpMSpM architecture (DMSA), incorporating four key techniques to fully exploit the parallelism of DMP and efficiently handle irregular memory accesses. Implemented on a Xilinx ZCU106 FPGA, DMSA achieves speedups of up to$3.38\times $and$1.73\times $over two state-of-the-art FPGA-based SpMSpM accelerators while maintaining comparable hardware resource usage. In addition, compared to CPU and GPU implementations on an NVIDIA Jetson AGX Xavier, DMSA is$4.96\times $and$1.53\times $faster while achieving$6.67\times $and$2.33\times $better energy efficiency, respectively.
Yuta Nagahara, Jiale Yan, Kazushi Kawamura, Daichi Fujiki, Masato Motomura, Thiem Van Chu
IEEE Trans. Very Large Scale Integr. Syst.2
2024 Sparse-Sparse Matrix Multiplication Accelerator on FPGA featuring Distribute-Merge Product Dataflow
abstract
Sparse-Sparse matrix multiplication (SpMSpM) is a critical computation in various fields such as computational science and graph analysis. It poses computational challenges for general-purpose CPUs and GPUs due to its requirements for random memory access and the inherently low spatial/temporal locality. Given the increasing importance of SpMSpM, numerous accelerators have been recently proposed. However, they suffer from various issues such as low input utilization, heavy computational load, and excessive memory traffic during the merging process of intermediate results. This paper introduces a novel Distribute-Merge Product (DMP) SpMSpM dataflow and a DMP-based SpMSpM Architecture (DMSA). DMP distributes the workload into balanced streams, generates partial matrices based on these streams, and merges the partial results in a parallel and pipelined fashion. We have designed DMSA as a highly scalable architecture, implemented it on a Xilinx ZCU106 Evaluation Kit, and evaluated it on a set of benchmarks from the SuiteSparse matrix collection. When compared to a latest SpMSpM accelerator with approximately the same amount of hardware resources on the same FPGA platform, DMSA achieves 2.72 × speedup, by facilitating the parallelism of partial matrix generation and merging. The speedup on the same platform reaches 4.80 × when the parallelism explored in the merging process is doubled, evidencing the DMSA’s superb scalability.
Yuta Nagahara, Jiale Yan, Kazushi Kawamura, Masato Motomura, Thiem Van Chu
ASPDAC2
2024 Enhance membership inference attacks in federated learning
Xinlong He, Sicong Zhang, Weida Xu, Jiale Yan
Comput. Secur.5
2018 GNA: Reconfigurable and Efficient Architecture for Generative Network Acceleration
abstract
Generative networks have become ubiquitous in image generation applications like image super-resolution, image to image translation, and text to image synthesis. They are usually composed of convolutional (CONV) layers, convolution-based residual blocks, and deconvolutional (DeCONV) layers. Previous works on neural network acceleration focus too much on optimizing CONV layers computation such as data-reuse or parallel computation, but have low processing element (PE) utilization in computing residual blocks and DeCONV layers: residual blocks require very high memory bandwidth when performing elementwise additions on residual paths; DeCONV layers have imbalanced operation counts for different outputs. In this paper, we propose a dual convolution mapping method for CONV and DeCONV layers to make full use of the available PE resources. A cross-layer scheduling method is also proposed to avoid extra off-chip memory access in residual block processing. Precision-adaptive PEs and buffer bandwidth reconfiguration are used to support flexible bitwidths for both inputs and weights in deep neural networks. We implement a generative network accelerator (GNA) based on intra-PE processing, inter-PE processing, and cross-layer scheduling techniques. Owing to the proposed optimization techniques, GNA achieves energy efficiency of 2.05 TOPS/W with 61% higher PE utilization than traditional methods in generative network acceleration.
Jiale Yan, Shouyi Yin, Fengbin Tu, Leibo Liu, Shaojun Wei
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1