Yanfeng Jiang

dblp:25/8407 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 SQ-Delta: Ultra-High Delta Compression for LLMs via Joint Sparsification-Quantization
abstract
Current delta compression methods struggle to achieve ultra-high compression, failing to minimize the deployment costs of multiple full-parameter fine-tuned large language models. To address this issue, we propose a novel distribution-driven delta compression framework SQ-Delta, which utilizes joint sparsification-quantization to maximize the compression ratio. We have observed that the matrix-computed intermediate results for the delta weight exhibit small variance and min-max range characteristics. Exploiting this distribution phenomenon, we introduce Group-wise Dropout to perform dropout on the delta weight using an optimal group size. Furthermore, using Separate Quantization, the sparse delta weight is quantized and decomposed to achieve a lower bit. Experimental results show that SQ-Delta achieves 16× compression with improved accuracy compared to baselines for WizardMath and WizardCoder models across different parameter scales. Moreover, SQ-Delta demonstrates the ability for ultra-high compression, achieving 128× compression for the WizardMath-7B model and 512× compression for the WizardMath-70B model. Our code is available at https://github.com/llwx593/SQ-Delta.
Yanfeng Jiang, Zelan Yang, Bohua Chen
ICME1
2025 ADFQ-ViT: Activation-Distribution-Friendly post-training Quantization for Vision Transformers
Yanfeng Jiang, Xueshuo Xie, Fei Yang 0007, Tao Li 0022
Neural Networks1
2024 MEFold: Memory-Efficient Optimization for Protein Language Models via Chunk and Quantization
abstract
Protein language models are currently experiencing a surge in demand owing to their remarkable accuracy in protein structure prediction. Nevertheless, their applications are hindered by the significant computation and memory requirements. The existing optimization strategies primarily focus on computational efficiency while often neglecting memory optimization, thereby restricting their suitability for devices with limited resources. In this paper, we propose MEFold, a novel memory-efficient optimization framework for protein language models that enables efficient inference on resource-constrained devices. MEFold consists of Look-up Table Chunk and Fine-grained Quantization. Look-up Table Chunk reduces the memory of intermediate activations by chunk and avoids the overhead of obtaining the optimal chunk size configuration through pre-computing. For the memory of model parameters, Fine-grained Quantization, delicately controls the scope of quantization to ensure that memory reduction is achieved while preventing declines in accuracy and computational speed. Experimental results show that, compared to the original model, for protein sequences ranging from 74 to 1024 in length, our method significantly reduces the peak memory during inference from 14.7-54.2GB to 6.0-14.4GB, while minimizing the impact on inference latency. On CASP14 and CAMEO datasets, the accuracy loss compared to the original model is below 1%. Moreover, our optimization provides various memory-saving alternatives. Our code is available at https://github.com/llwx593/MEFold.
Yanfeng Jiang, Zhengxian Lu, Fei Yang 0007, Tao Li 0022
IJCNN1
2024 Quantitative evaluation of deep learning frameworks in heterogeneous computing environment
Zhengxian Lu, Chengkun Du, Yanfeng Jiang, Xueshuo Xie, Tao Li 0022, Fei Yang 0007
CCF Trans. High Perform. Comput.3
2023 Exploring Post-Training Quantization of Protein Language Models
abstract
Recent advancements in unsupervised protein language models (ProteinLMs), like ESM-1b [27] and ESM-2 [21], have shown promise in different protein prediction tasks. However, these models face challenges due to their high computational demands, significant memory needs, and latency, restricting their usage on devices with limited resources. To tackle this, we explore post-training quantization (PTQ) for ProteinLMs, focusing on ESMFold [21], a simplified version of AlphaFold [16] based on ESM-2 ProteinLM. Our study is the first attempt to quantize all weights and activations of ProteinLMs. We observed that the typical uniform quantization method performs poorly on ESMFold, causing a significant drop in TM-Score when using 8-bit quantization. We conducted extensive quantization experiments and discovered unique challenges associated with ESMFold. Specifically, we found that the activation ranges before Layer Normalization are highly asymmetric. This asymmetry makes it difficult to represent the data effectively when using low-bit fixed-point formats. To address these challenges, we propose a new PTQ method for ProteinLMs, utilizing piecewise linear quantization for asymmetric activation values to ensure accurate approximation. We demonstrated the effectiveness of our method in protein structure prediction tasks, showing that ESMFold can be quantized to low-bit widths without compromising accuracy. Additionally, we applied our method to the contact prediction task, showcasing its versatility. In summary, our study introduces an innovative PTQ method for ProteinLMs, addressing specific quantization challenges and potentially leading to the development of more efficient ProteinLMs with significant implications for various protein-related applications.
Fei Yang 0007, Yanfeng Jiang, Aimin Pan
BIBM5
2023 Double-Ended Superposition Anti-Noise Resistance Monitoring Write Termination Scheme for Reliable Write Operation in STT-MRAM
abstract
Although resistance monitoring write termination (RM-WT) scheme for STT-MRAM can reduce the write energy, the degradation of read margin due to low tunnel magnetoresistance ratio (TMR) and intrusion of noise with process variation still seriously deteriorates the stability of the WT operation. In this paper, a double-ended superposition anti-noise write termination (DSA-WT) scheme is proposed and implemented, in which the voltage changes on both BL and SL can be superimposed to boost sensing margin (SM). Schmitt trigger (ST) is adopted to take the place of the inverter (INV) in the traditional WT scheme, which is demonstrated to be helpful for stability improvement. Based on 65-nm CMOS technology, the proposed DSA-WT scheme shows 15% ~33% sensing margin boosting under various PVT conditions and 1000 times lower read bit-error-rate (BER) compared with the other WT schemes. The write done (WD) delay and the energy-delay-product (EDP) achieve 49.6% and 47.2% improvements compared to the state-of-art self-referenced single-ended RM-WT scheme (SS-RM-WT), respectively.
An Yang, Zhilin Jiang, Yanfeng Jiang
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 RISC-V-Based Evaluation and Strategy Exploration of MRAM Triple-Level Hybrid Cache Systems
abstract
Magnetic random access memory (MRAM) is considered one of the most promising memories among nonvolatile memories (NVMs), to replace static random access memory (SRAM) in the computing system to improve its cache performance. In this article, performance evaluation and strategy exploration of three-level hybrid cache systems are conducted under the framework of Reduced InstrucTIon Set Computer-Five (RISC-V) instruction set. First, system-level joint calls of Nvsim, Gem5, McPAT, and other cache emulation architectures are implemented based on the system-level simulator called MAGPIE tool. Based on the performance comparison of spin transfer torque (STT)-MRAM, spin-orbit torque (SOT)-MRAM, and SRAM with varied capacities, a variety of combined CPU three-level hybrid cache architectures are simulated. The performances of the proposed hybrid cache systems and the bus impact of interconnecting are evaluated, with better analysis of the cost of CPU system implementation and interactions among levels of cache. The power consumption of different processors with various instruction sets and cores is also evaluated. It is demonstrated that SOT-MRAM and STT-MRAM show great potential applications as L2 and L3 caches in the RISC-V system. RISC-V shows the potential benefits for future CPU cache. Meanwhile, an adaptive stride prefetching strategy is proposed to address the problem of delay hysteresis and high miss rate of STT-MRAM in the L3 cache. The applicability of this strategy to different storage technologies and the comparison with cutting-edge technologies are also illustrated. Simulation results show that the prefetching strategy can achieve at most 64.12% and 94.55% optimization effects on the miss latency and the miss rate of L3 cache, respectively.
Shaopu Han, Yanfeng Jiang
IEEE Trans. Very Large Scale Integr. Syst.2
2022 Long-Term Person Re-identification with Dramatic Appearance Change: Algorithm and Benchmark
abstract
For person re-identification (Re-ID) task, most of previous studies assumed that the pedestrians do not change their appearances. The works on cross-appearance Re-ID, including datasets and algorithms, are still few. Therefore, this paper contributes a cross-season appearance change Re-ID dataset, namely NKUP+, including more than 300 IDs from surveillance videos over 10 months, to support the studies of the cross-appearance Re-ID. In addition, we propose a network named M2Net, which integrates multi-modality features from the RGB images, contour images and human parsing images. By ignoring irrelevant misleading information for cross-appearance retrieval in RGB images, M2Net can learn features that are robust to appearance changes. Meanwhile, we propose a sampling strategy called RAS to contain a variety of appearances in one batch. And appearance loss and multi-appearance loss are designed to guide the network to learn both same-appearance and cross-appearance features. Finally, we evaluated our method on NKUP+/PRCC/DeepChange datasets, and the results showed that, compared with the baseline, our method renders significant improvement, leading to the state-of-the-art performance over other methods. Our dataset is available at https://github.com/nkicsl/NKUP-dataset.
Tao Li 0022, Yanfeng Jiang, Kai Wang 0001
ACM Multimedia4
2022 Prediction model of the impact of innovation and entrepreneurship on China's digital economy based on neural network integration systems
Yanfeng Jiang
Neural Comput. Appl.1