Fei Yang 0007

dblp:19/2504-7 · DBLP profile ↗
← Back
18ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0003-4802-3191ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Theory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 OptiCo: Adaptive Distributed Training Optimization via Collaborative Agent Reasoning
abstract
Sheng Chen, Tang Zhe, Weixing Zhang, Fei Yang, Yuanyuan. Wang, Tianlin Li, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tang Zhe, Fei Yang 0007, Tianlin Li, Yang Liu 0003
ACL (1)4
2026 HVM: A HV-Attention Based Multi-modal Framework for Mathematical Expression Recognition
Linnan Jiang, Fei Yang 0007, Jiaojiao Ye, WeiXing Zhang
ICIC (22)2
2026 LaRA: Layer-wise rank allocation for efficient fine-tuning of pruned large language models
Yuhua Zhou, Changhai Zhou, Shiyang Zhang, Fei Yang 0007, Aimin Pan
Inf. Process. Manag.4
2025 ADFQ-ViT: Activation-Distribution-Friendly post-training Quantization for Vision Transformers
Yanfeng Jiang, Xueshuo Xie, Fei Yang 0007, Tao Li 0022
Neural Networks4
2025 On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
abstract
ABSTRACT Background: Transformer models have emerged as potent solutions to a wide array of multidisciplinary challenges. The deployment of transformer architectures is significantly hindered by their extensive computational and memory requirements, necessitating reliance on advanced efficient distributed training methodologies. Motivation: Prior research has delved into the performance bottlenecks associated with distributed training, aiming to unravel these bottlenecks and suggest optimization directions. However, such analyses often overlook three aspects unique to transformer models: the specialized architecture, the dependency on various distributed strategies, and the requirement to balance computational and memory overhead. Method: This paper aims to bridge this gap by offering a comprehensive examination of the performance bottlenecks inherent in the distributed training of transformer models, leveraging both theoretical analysis and empirical investigation. We propose an analytical framework tailored to these unique aspects of transformers, facilitating a holistic evaluation of model architectures, distributed strategies, and resource consumption. Based on this analytical framework, we conduct a comparative analysis of theoretical performances and further systematically explore how various distributed training strategies fare in real‐world scenarios. Results: Most of the experimental results can be well explained by the analytical outcomes derived from the analytical framework. Notably, our findings suggest an advantage of pipeline parallelism over data parallelism for transformer models. Moreover, we shed light on some unexpected outcomes, such as the potential for increased total memory overhead due to suboptimal model partitioning within pipeline parallelism. Additionally, we underscore the significance of communication block size and waiting time to further enhance performance.
Zhengxian Lu, Fangyu Wang, Fei Yang 0007, Tao Li 0022
Softw. Pract. Exp.4
2025 Deep Face Leakage: Inverting High-Quality Faces From Gradients Using Residual Optimization
abstract
Collaborative learning has gained significant traction for training deep learning models without sharing the original data of participants, particularly when dealing with sensitive data such as facial images. However, current gradient inversion attacks are employed to progressively reconstruct private data from gradients, and they have shown successful in extracting private training data. Nonetheless, our observations reveal that these methods exhibit suboptimal performance in face reconstruction and result in the loss of numerous facial details. In this paper, we propose DFLeak, an effective approach to boost face leakage from gradients using residual optimization and thwart the privacy of facial applications in collaborative learning. In particular, we first introduce a superior initialization method to stabilize the inversion process. Second, we propose to integrate prior-free face restoration (PFFR) results into the gradient inversion optimization process in a residual manner, which enriches facial details. We further design a pixel update schedule to mitigate the adverse effects of image regularization terms and preserve fine facial details. Comprehensive experimentation demonstrates the effectiveness of our approach in achieving more realistic and higher-quality facial image reconstructions, surpassing the performance of state-of-the-art gradient inversion attacks.
Tao Xiang 0001, Shangwei Guo, Fei Yang 0007, Tianwei Zhang 0004
IEEE Trans. Image Process.4
2024 Wi-Fi based Gait Recognition using Spectrogram and Phase
abstract
Conventional approaches of human gait recognition are plagued by issues such as invasion on privacy, constraints in terms of space, limits in lighting conditions, and inconveniences associated with wearable gadgets. Herein, we explore the potential advantages of Wi-Fi signals and fuse two modalities, namely phase and spectrogram, of Wi-Fi Channel State Information (CSI) as gait features. This integration leads to the development of a robust human gait recognition system called MultiGaFi. Specifically, we extract time-frequency features of spectrograms through a well-designed network, while temporal features related to phase changes are captured through gated recurrent unit (GRU). Subsequently, the time-frequency features and temporal features are linearly fused to facilitate human gait recognition. We implement MultiGaFi on commodity Wi-Fi devices across four distinct indoor environments, and the empirical findings indicate that MultiGaFi achieves an impressive average accuracy of 98.11% within these specific indoor environments.
Fei Yang 0007, Aimin Pan, Zhewei Mei
ICME2
2024 Holmes: Towards Distributed Training Across Clusters with Heterogeneous NIC Environment
abstract
Large language models (LLMs) such as GPT-3, OPT, and LLaMA have demonstrated remarkable accuracy in a wide range of tasks. However, training these models can incur significant expenses, often requiring tens of thousands of GPUs for months of continuous operation. Typically, this training is carried out in specialized GPU clusters equipped with homogeneous high-speed Remote Direct Memory Access (RDMA) network interface cards (NICs). The acquisition and maintenance of such dedicated clusters is challenging. Current LLM training frameworks, like Megatron-LM and Megatron-DeepSpeed, focus primarily on optimizing training within homogeneous cluster settings. In this paper, we introduce Holmes, a training framework for LLMs that employs thoughtfully crafted data and model parallelism strategies over the heterogeneous NIC environment. Our primary technical contribution lies in a novel scheduling method that intelligently allocates distinct computational tasklets in LLM training to specific groups of GPU devices based on the characteristics of their connected NICs. Furthermore, our proposed framework, utilizing pipeline parallel techniques, demonstrates scalability to multiple GPU clusters, even in scenarios without high-speed interconnects between nodes in distinct clusters. We conducted comprehensive experiments that involved various scenarios in the heterogeneous NIC environment. In most cases, our framework achieves performance levels close to those achievable with homogeneous RDMA-capable networks (InfiniBand or RoCE), significantly exceeding training efficiency within the pure Ethernet environment. Additionally, we verified that our framework outperforms other mainstream LLM frameworks under heterogeneous NIC environment in terms of training efficiency and can be seamlessly integrated with them.
Fei Yang 0007, Fangyu Wang, Fu Wu, Jiezhong Qiu, Aimin Pan
ICPP1
2024 MEFold: Memory-Efficient Optimization for Protein Language Models via Chunk and Quantization
abstract
Protein language models are currently experiencing a surge in demand owing to their remarkable accuracy in protein structure prediction. Nevertheless, their applications are hindered by the significant computation and memory requirements. The existing optimization strategies primarily focus on computational efficiency while often neglecting memory optimization, thereby restricting their suitability for devices with limited resources. In this paper, we propose MEFold, a novel memory-efficient optimization framework for protein language models that enables efficient inference on resource-constrained devices. MEFold consists of Look-up Table Chunk and Fine-grained Quantization. Look-up Table Chunk reduces the memory of intermediate activations by chunk and avoids the overhead of obtaining the optimal chunk size configuration through pre-computing. For the memory of model parameters, Fine-grained Quantization, delicately controls the scope of quantization to ensure that memory reduction is achieved while preventing declines in accuracy and computational speed. Experimental results show that, compared to the original model, for protein sequences ranging from 74 to 1024 in length, our method significantly reduces the peak memory during inference from 14.7-54.2GB to 6.0-14.4GB, while minimizing the impact on inference latency. On CASP14 and CAMEO datasets, the accuracy loss compared to the original model is below 1%. Moreover, our optimization provides various memory-saving alternatives. Our code is available at https://github.com/llwx593/MEFold.
Yanfeng Jiang, Zhengxian Lu, Fei Yang 0007, Tao Li 0022
IJCNN6
2024 Quantitative evaluation of deep learning frameworks in heterogeneous computing environment
Zhengxian Lu, Chengkun Du, Yanfeng Jiang, Xueshuo Xie, Tao Li 0022, Fei Yang 0007
CCF Trans. High Perform. Comput.6
2023 Exploring Post-Training Quantization of Protein Language Models
abstract
Recent advancements in unsupervised protein language models (ProteinLMs), like ESM-1b [27] and ESM-2 [21], have shown promise in different protein prediction tasks. However, these models face challenges due to their high computational demands, significant memory needs, and latency, restricting their usage on devices with limited resources. To tackle this, we explore post-training quantization (PTQ) for ProteinLMs, focusing on ESMFold [21], a simplified version of AlphaFold [16] based on ESM-2 ProteinLM. Our study is the first attempt to quantize all weights and activations of ProteinLMs. We observed that the typical uniform quantization method performs poorly on ESMFold, causing a significant drop in TM-Score when using 8-bit quantization. We conducted extensive quantization experiments and discovered unique challenges associated with ESMFold. Specifically, we found that the activation ranges before Layer Normalization are highly asymmetric. This asymmetry makes it difficult to represent the data effectively when using low-bit fixed-point formats. To address these challenges, we propose a new PTQ method for ProteinLMs, utilizing piecewise linear quantization for asymmetric activation values to ensure accurate approximation. We demonstrated the effectiveness of our method in protein structure prediction tasks, showing that ESMFold can be quantized to low-bit widths without compromising accuracy. Additionally, we applied our method to the contact prediction task, showcasing its versatility. In summary, our study introduces an innovative PTQ method for ProteinLMs, addressing specific quantization challenges and potentially leading to the development of more efficient ProteinLMs with significant implications for various protein-related applications.
Fei Yang 0007, Yanfeng Jiang, Aimin Pan
BIBM2
2023 Where to Go Next for Recommender Systems? ID- vs. Modality-based Recommender Models Revisited
abstract
Recommendation models that utilize unique identities (IDs for short) to represent distinct users and items have been state-of-the-art (SOTA) and dominated the recommender systems (RS) literature for over a decade. Meanwhile, the pre-trained modality encoders, such as BERT [9] and Vision Transformer [11], have become increasingly powerful in modeling the raw modality features of an item, such as text and images. Given this, a natural question arises: can a purely modality-based recommendation model (MoRec) outperforms or matches a pure ID-based model (IDRec) by replacing the itemID embedding with a SOTA modality encoder? In fact, this question was answered ten years ago when IDRec beats MoRec by a strong margin in both recommendation accuracy and efficiency.
Zheng Yuan 0013, Fajie Yuan, Yu Song 0007, Youhua Li, Junchen Fu, Fei Yang 0007, Yunzhu Pan, Yongxin Ni
SIGIR6
2023 A novel seminar learning framework for weakly supervised salient object detection
Yan Liu 0080, Yunzhou Zhang, Zhenyu Wang 0010, Fei Yang 0007, Sonya A. Coleman, Dermot Kerr
Eng. Appl. Artif. Intell.4
2022 Complementary characteristics fusion network for weakly supervised salient object detection
Yan Liu 0080, Yunzhou Zhang, Zhenyu Wang 0010, Fei Yang 0007, Cao Qin, Sonya A. Coleman, Dermot Kerr
Image Vis. Comput.4
2022 ARoBERT: An ASR Robust Pre-Trained Language Model for Spoken Language Understanding
abstract
Spoken Language Understanding (SLU) aims to interpret the meanings of human speeches in order to support various human-machine interaction systems. A key technique for SLU is Automatic Speech Recognition (ASR), which transcribes speech signals into text contents. As the output texts of modern ASR systems unavoidably contain errors, mainstream SLU models either trained or tested on texts transcribed by ASR systems would not be sufficiently error robust. We present ARoBERT, an ASR Robust BERT model, which can be fine-tuned to solve a variety of SLU tasks with noisy inputs. To guarantee the robustness of ARoBERT, during pretraining, we decrease the fluctuations of language representations when some parts of the input texts are replaced by homophones or synophones. Specifically, we propose two novel self-supervised pre-training tasks for ARoBERT, namely Phonetically-aware Masked Language Modeling (PMLM) and ASR Model-adaptive Masked Language Modeling (AMMLM). The PMLM task explicitly fuses the knowledge of word phonetic similarities into the pre-training process, which forces homophones and synophones to share similar representations. In AMMLM, a data-driven algorithm is further introduced to mine typical ASR errors such that ARoBERT can tolerate ASR model errors. In the experiments, we evaluate ARoBERT over multiple datasets. The results show the superiority of ARoBERT, which consistently outperforms strong baselines. We have also shown that ARoBERT outperforms state-of-the-arts on a public benchmark. Currently, ARoBERT has been deployed in an online production system with significant improvements.
Chengyu Wang 0001, Suyang Dai, Yipeng Wang 0013, Fei Yang 0007, Minghui Qiu, Jun Huang 0007
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Meta Distant Transfer Learning for Pre-trained Language Models
abstract
With the wide availability of Pre-trained Language Models (PLMs), multi-task fine-tuning across domains has been extensively applied.For tasks related to distant domains with different class label sets, PLMs may memorize nontransferable knowledge for the target domain and suffer from negative transfer.Inspired by meta-learning, we propose the Meta Distant Transfer Learning (Meta-DTL) framework to learn the cross-task knowledge for PLM-based methods.Meta-DTL first employs task representation learning to mine implicit relations among multiple tasks and classes.Based on the results, it trains a PLM-based meta-learner to capture the transferable knowledge across tasks.The weighted maximum entropy regularizers are proposed to make meta-learner more task-agnostic and unbiased.Finally, the meta-learner can be fine-tuned to fit each task with better parameter initialization.We evaluate Meta-DTL using both BERT and ALBERT on seven public datasets.Experiment results confirm the superiority of Meta-DTL as it consistently outperforms strong baselines.We find that Meta-DTL is highly effective when very few data is available for the target task.
Chengyu Wang 0001, Haojie Pan, Minghui Qiu, Jun Huang 0007, Fei Yang 0007, Yin Zhang 0006
EMNLP (1)5
2021 The π-Calculus is Behaviourally Complete and Orbit-Finitely Executable
Bas Luttik, Fei Yang 0007
Log. Methods Comput. Sci.2
2016 On the Executability of Interactive Computation
Bas Luttik, Fei Yang 0007
CiE2