Yaqing Li

dblp:08/8120 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 BDARec: Balancing Diversity and Accuracy of Recommendation Model with Graph Neural Networks
abstract
Based on research in cognitive psychology, humans typically seek a balance between their preference for familiar things and the exploration of new ones during decision-making. Therefore, studying the relationship between accuracy and diversity in recommendation systems is particularly meaningful. In recent years, recommender systems based on Graph Neural Networks (GNNs) have garnered significant attention for enhancing recommendation accuracy or diversity. However, existing works often improve accuracy or diversity at the expense of the other aspect, which is inconsistent with the complex needs of users. In this paper, we propose a novel Recommendation model that Balances Diversity and Accuracy with GNNs, called BDARec. Firstly, BDARec proposes a balanced neighborhood aggregation strategy to select diverse and accurate neighbor nodes for updating node embeddings in user-item bipartite heterogeneous graph. Secondly, to accelerate the convergence of BDARec, an enhanced category-boosted negative sampling strategy is proposed to select negative samples from the same category positive samples with a certain probability. Thirdly, we put forward a dynamic feature for each item to measure the importance of items in training phase. Finally, we conduct extensive experiments on three real-world datasets. Experimental results show that our model can even improve recall by 22.04%, hit ratio by 16.46%, and coverage by 10.27% when compared to the state-of-the-art comparison algorithm, which verifies that the proposed model can achieve the best balance between diversity and accuracy.
Jinlong Tian, Yaqing Li
IJCNN5
2025 BE-NPU: A Bandwidth-Efficient Neural Processing Unit With Adaptive Processing Schemes for Reduced Off-Chip Bandwidth Demand
abstract
Existing neural processing units (NPUs) mainly focus on the optimized multiply-accumulate (MAC) arrays for efficient inference of convolutional neural networks (CNNs). However, off-chip data transmission usually keeps NPUs waiting during CNN inference, causing up to 38.4GB/s off-chip bandwidth (OCB) demand for mobile AI devices. And none of the previous benchmarks quantitatively evaluate the bandwidth efficiency of different NPU architectures. In addition, CNNs exhibit distinct characteristics of off-chip data transmission when applied to different fields, and it has become a challenging task for NPUs to support different CNNs efficiently with reasonable OCB demand. To address the aforementioned issues, this paper proposes the Bandwidth-Peak Performance Ratio for n percentages of ideal frame rate (BPPR-n%) to demonstrate the normalized OCB demand of different NPU architectures. A bandwidth-efficient NPU (BE-NPU) is introduced with adaptive processing schemes to reduce the OCB demand during inference of different CNNs. The adaptive processing schemes include both instruction-level and thread-level schemes. For the instruction-level scheme, decoupled execute/access is introduced into depth-first (DF) and layer-first (LF) schemes to improve the concurrency between NPU calculation (CAL) and direct memory access (DMA) instructions. For the thread-level scheme, DF and LF threads are hybridly processed to further improve overall NPU efficiency. Compared with state-of-the-art works, BE-NPU achieves 48.1%~80.6% reduction of BPPR-80% and 67.0%~95.1% reduction of BPPR-95%. The proposed architecture is synthesized with TSMC 28nm technology node. BE-NPU utilizes 14.3% additional logic gates compared with baseline implementation.
Yichuan Bai, Yaqing Li, Yuan Du
IEEE Trans. Computers4
2024 Optoelectronic Computing Evaluation and Deployment Platform Based on a 256-MAC Silicon Photonic Chip
abstract
The deceleration of Moore's Law has led to increasing difficulties in advancing the computational speed and power efficiency of Complementary-Metal-Oxide-Semiconductor (CMOS) chips. As a solution to this challenge, optical computing emerges as a promising technology, boasting low energy consumption, high processing speed, and extensive bandwidth. Yet, a critical obstacle remains: the absence of a co-simulation platform that incorporates both photonic chips and peripheral electrical circuits. This paper addresses this gap by introducing a hybrid optoelectronic computing evaluation and deployment platform utilizing Simulink tools. Based on the measured data from the silicon optical computing chip, we have deployed an image filtering algorithm and a convolutional neural network onto this platform. The optical computing chip achieves an accuracy of 86.4% on the ImageNet image dataset. Through evaluation, we have identified the most substantial impacts on calculation results. To achieve an image classification accuracy of 80%, the signal-to-noise ratio (SNR) of the low-speed DAC must be a minimum of 52 dB. These findings provide crucial insights into the optimization of optical computing systems.
Likai Li, Yichuan Bai, Shengping Liu, Sunan He, Yaqing Li, Yuan Du
ISCAS6
2024 A Compilation Framework for SRAM Computing-in-Memory Systems With Optimized Weight Mapping and Error Correction
abstract
Deploying convolution-based algorithms into SRAM computing-in-memory (CIM) systems faces various challenges, such as operator incompatibility and intrinsic non-ideal error. This paper proposes a compilation framework to address this issue. Efficient weight mapping strategies are introduced to improve the utilization of SRAM-CIM macro. The intrinsic non-ideal errors of SRAM-CIM macro are also taken into consideration, and two efficient error correction schemes are proposed, which include calibration of computation voltage linear error (CCVLE) and the mitigation of analog-to-digital quantization error (MAQE). In addition, bit-width flexibility and signed-unsigned reconfigurability are also supported to facilitate the deployment of various convolution-based algorithms. ResNet18, finite impulse response (FIR) filtering, and Gaussian image filtering are deployed into a multi-macro SRAM-CIM system. These algorithms serve as deployment representatives of convolutional neural network (CNN), digital signal processing (DSP), and digital image processing (DIP), respectively. The results show that the introduced weight mapping strategies improve the macro utilization by 63.29% and 21.10% for two types of frequently used convolution layers compared to the traditional strategy. Moreover, the proposed error correction schemes achieve similar algorithm accuracy to the floating-point results, and the deployment result of ResNet18 achieves 66.3%~70.1% top-1 classification accuracy evaluated on the ImageNet dataset with different throughput tradeoffs.
Yichuan Bai, Yaqing Li, Heng Zhang 0024, Aojie Jiang, Yuan Du
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2018 Design and Simulation Optimization of a Novel Oocyte Ultrasonic Micro-dissection Instrument
abstract
The zona pellucida (ZP) micro-dissection technology has played a key role in the field of artificial assisted reproduction (such as ZP thinning, preimplantation genetic diagnosis (PGD), etc.). Currently, laser ZP micro-cutting technology using far-infrared beam has the disadvantages of high price, thermal damage, and low degree of freedom. Piezoelectric ultrasonic micro-cutting technology has matured at the biological tissue level, but ultrasonic cutting at the single-cell level is still difficult to achieve. In this paper, based on piezoelectric ultrasonic tissue micro-cutting theory, the mechanism of ultrasonic cutting of the oocyte ZP was researched, and a cutting method for the ZP was proposed. Based on the design formula of flexure's stiffness, this paper analyzes the effects of dimension parameters on the vibration condition of micro-needle. An ultrasonic cell surgery instrument based on a three-dimensional stereoscopic flexure-guided structure was designed. The modal analysis and the harmonic response analysis of the structure were performed using finite element software. The experimental results show that the lateral amplitude of the newly designed device's needle tip is smaller than the traditional one in the specified range of frequency. Among them, when the operating frequency is around 22.1 kHz, the lateral amplitude of the needle tip is reduced to 0.021μm. Finally, the theoretical method of piezoelectric ultrasonic cutting on single cell layer is presented for the first time, which has a broad prospect and significance for artificial assisted reproduction.
Xiwei Gao, Liguo Chen, Mingqiang Pan, Su Yan 0005, Yaqing Li, Lining Sun
ICARCV6
2016 Application-Aware and Software-Defined SSD Scheme for Tencent Large-Scale Storage System
abstract
Tencent, one of the biggest Internet companies in China, contains billions of users and over 600-PB data, and leverages thousands of SSDs in the storage system to improve system performance and obtain energy savings. Existing commercial SSDs however fail to meet the needs of the ultra largescale applications due to not matching the service patterns. In order to address this problem and deliver high performance, we propose an application-aware and software-defined SSD scheme for Tencent applications, called TSSD. TSSD explores and exploits the business characteristics of Tencent, which facilitates the efficient use of SSDs. TSSD is software-defined by packaging each flash chip as a fully independent and concurrent storage unit. Each concurrent unit can be mounted as a character device, which allows the application layer to manage the flash chips in a more efficient manner, while optimizing the data layout. TSSD further employs a host-target FTL (TFTL) that uses a dedicated interface in the application layer, which efficiently connects the application layer with flash chips. Application layer hence becomes more accurately by using the flash memory chip-level information from TFTL, including the storage utilization, the degree of wear, etc. Moreover, TFTL is a programmable FTL and provides a programmable interface to the application layer. According to the running states of SSDs and workload information, TSSD makes use of the programmable interface to efficiently improve the performance of the FTL, wear leveling, and garbage collection for the specified applications. Extensive experiments use the real-world datasets from the commercial storage systems of Tencent. The results demonstrate that TSSD significantly improves the storage system performance and meets the needs of the Tencent's large-scale business applications.
Jianquan Zhang, Dan Feng 0001, Jianlin Gao, Wei Tong 0001, Jingning Liu, Yu Hua 0001, Caihua Fang, Wen Xia, Feiling Fu, Yaqing Li
ICPADS11