Yanfeng Xu

dblp:249/4280 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Graph Self-Supervised Learning via Learnable View Augmentation for Recommender System
abstract
In the field of recommender systems, graph neural networks (GNNs) have been extensively applied to collaborative filtering to generate personalized recommendations for users. To solve the problem of lack of observed data and contrasting interactions during representation learning, graph contrastive learning as an effective self-supervised learning (SSL) technique is presented to obtain augmented user and item representations. Nevertheless, most self-supervised approaches to generate recommendation either disrupt the graph structure or node embeddings through random augmentations or introduce augmented SSL information from biased data through heuristic methods. To overcome these challenges, we propose a learnable view augmentation model for collaborative filtering (LACF). Specifically, our framework embeds parameterized learnable view generators layer by layer into the automatic augmentation strategy, thus dynamically optimizing the adaptive augmented views of users and items through the backpropagation of weight gradients. In addition, LACF introduces a multiscale learning strategy that guides the view generator with layer-wise aware optimization and graph-level adaptive augmentation, enabling joint learning of representations with topological heterogeneity and semantic similarity from integrated viewpoint, achieving superior view augmentation. Extensive experiments on realworld datasets demonstrate that our LACF outperforms state-of-the-art baselines. In-depth analysis confirms the advantages of LACF in resistance against noise disturbances, alleviating data sparsity, and improving training efficiency.
Hengjing Xiang, Yanfeng Xu, Sen Liu 0002, Zhi Liu 0011, Guangnan Ye
IEEE Trans. Ind. Informatics3
2024 An All-digital Compute-in-memory FPGA Architecture for Deep Learning Acceleration
abstract
Field Programmable Gate Array (FPGA) is a versatile and programmable hardware platform, which makes it a promising candidate for accelerating Deep Neural Networks (DNNs). However, FPGA’s computing energy efficiency is low due to the domination of energy consumption by interconnect data movement. In this article, we propose an all-digital Compute-in-memory FPGA architecture for deep learning acceleration. Furthermore, we present a bit-serial computing circuit of the Digital CIM core for accelerating vector-matrix multiplication (VMM) operations. A Network-CIM-deployer ( NCIMD ) is also developed to support automatic deployment and mapping of DNN networks. NCIMD provides a user-friendly API of DNN models in Caffe format. Meanwhile, we introduce a Weight-stationary dataflow and describe the method of mapping a single layer of the network to the CIM array in the architecture. We conduct experimental tests on the proposed FPGA architecture in the field of Deep Learning (DL), as well as in non-DL fields, using different architectural layouts and mapping strategies. We also compare the results with the conventional FPGA architecture. The experimental results show that compared to the conventional FPGA architecture, the energy efficiency can achieve a maximum speedup of 16.1×, while the latency can decrease up to 40% in our proposed CIM FPGA architecture.
Yonggen Li, Xin Li 0177, Haibin Shen, Jicong Fan 0002, Yanfeng Xu, Kejie Huang
ACM Trans. Reconfigurable Technol. Syst.5
2023 A Low-Power In-Memory Multiplication and Accumulation Array With Modified Radix-4 Input and Canonical Signed Digit Weights
abstract
Data transfer between the processing and storage units has become a significant bottleneck in modern von Neumann computing systems for artificial intelligence (AI) tasks. Computing in memory (CIM) has emerged as a promising candidate for lowering latency and power consumption. However, the conventional analog CIM schemes are suffering from reliability issues, which may significantly degenerate the accuracy of the computation. Recently, digitized input data and weights have been utilized for high-reliable in-memory computing. However, the properties of the digital memory and input data are not fully utilized. This article presents a novel low-power CIM scheme to further reduce the power consumption by using a modified radix-4 (M-RD4) booth algorithm at the input and a modified canonical signed digit (M-CSD) for the network weights. The simulation results show that M-RD4 and M-CSD reduce the number of nonzero activation bits by 24.2% and the number of nonzero weight bits by 36.0% in AlexNet, respectively. The power consumption can be reduced by 41.6% on average. The computing-power ratio at the fixed-point 8 bit is 60.7 tera operations per second per watt (TOPS/W), and the density is 0.177 TOPS/mm2.
Rui Xiao 0003, Yewei Zhang, Bo Wang 0020, Yanfeng Xu, Jicong Fan 0002, Haibin Shen, Kejie Huang
IEEE Trans. Very Large Scale Integr. Syst.4
2022 An 8-Bit in Resistive Memory Computing Core With Regulated Passive Neuron and Bitline Weight Mapping
abstract
The rapid development of artificial intelligence (AI) and Internet of Things (IoT) increase the requirement for edge computing with low power and relatively high processing speed devices. The computing-in-memory (CIM) schemes based on emerging resistive nonvolatile memory (NVM) show great potential in reducing the power consumption for AI computing. However, the inconsistency of the NVM may significantly degenerate the performance of the neural network. In this article, we propose a low power resistive RAM (RRAM)-based CIM core to not only achieve high computing efficiency but also greatly enhance the robustness by bit line (BL) regulator and BL weight mapping algorithm. The simulation results show that the power consumption of our proposed 8-bit CIM core is only 12.6 mW ($256\times 256$at 8b). The spurious-free dynamic range (SFDR) and signal to noise and distortion ratio (SNDR) of the CIM core achieve 62.64 and 45.92 dB, respectively. The proposed BL weight mapping scheme improves the top-1 accuracy by 2.46% and 3.47% for AlexNet and VGG16 on ImageNet Large Scale Visual Recognition Competition 2012 (ILSVRC 2012) in 8-bit mode, respectively.
Yewei Zhang, Kejie Huang, Rui Xiao 0003, Bo Wang 0020, Yanfeng Xu, Jicong Fan 0002, Haibin Shen
IEEE Trans. Very Large Scale Integr. Syst.5