Yefan Zhou

dblp:237/4333 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Efficient and distributed learning · 38% Deep learning architectures and training · 22% Trustworthy machine learning · 18%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 23 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › training optimization
layer-wise learning rate scheduling
1.422024
Model Balancing Helps Low-data Training and Fine-tuning · EMNLP 2024
Temperature Balancing, Layer-wise Weight Analysis, and Neural Network Training · NeurIPS 2023
Machine learning › Deep learning architectures and training
loss landscape
1.422024
MD tree: a model-diagnostic tree grown on loss landscape · ICML 2024
A Three-regime Model of Network Pruning · ICML 2023
Machine learning › Efficient and distributed learning
model compression
1.422024
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models · NeurIPS 2024
A Three-regime Model of Network Pruning · ICML 2023
Machine learning › Efficient and distributed learning › model compression
pruning
1.422024
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models · NeurIPS 2024
A Three-regime Model of Network Pruning · ICML 2023
Machine learning › Kernel, tree and ensemble methods › ensemble learning
deep ensembles
0.812024
Sharpness-diversity tradeoff: improving flat ensembles with SharpBalance · NeurIPS 2024
Machine learning › Optimization for machine learning › implicit regularization
heavy-tailed self-regularization
0.812024
Model Balancing Helps Low-data Training and Fine-tuning · EMNLP 2024
Machine learning › Trustworthy machine learning
interpretability
0.812024
MD tree: a model-diagnostic tree grown on loss landscape · ICML 2024
Machine learning › Efficient and distributed learning › model compression
large language model compression
0.812024
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › pruning › structured pruning
layer pruning
0.812024
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models · NeurIPS 2024
Machine learning › Efficient and distributed learning › parameter-efficient fine-tuning
low-rank adaptation
0.812024
AlphaLoRA: Assigning LoRA Experts Based on Layer Training Quality · EMNLP 2024
Machine learning › Deep learning architectures and training
mixture of experts
0.812024
AlphaLoRA: Assigning LoRA Experts Based on Layer Training Quality · EMNLP 2024
Machine learning › Trustworthy machine learning › interpretability › model debugging
model diagnosis
0.812024
MD tree: a model-diagnostic tree grown on loss landscape · ICML 2024
Machine learning › Trustworthy machine learning
out-of-distribution generalization
0.812024
Sharpness-diversity tradeoff: improving flat ensembles with SharpBalance · NeurIPS 2024
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.812024
AlphaLoRA: Assigning LoRA Experts Based on Layer Training Quality · EMNLP 2024
Machine learning › Trustworthy machine learning
robustness
0.812024
Sharpness-diversity tradeoff: improving flat ensembles with SharpBalance · NeurIPS 2024
Machine learning › Efficient and distributed learning › model compression › sparsity
sparsity allocation
0.812024
AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models · NeurIPS 2024
Machine learning › Optimization for machine learning
learning rate schedule
0.712023
Temperature Balancing, Layer-wise Weight Analysis, and Neural Network Training · NeurIPS 2023
Robotics › Robot manipulation › grasping
grasp detection
0.612022
Learn to Grasp with Less Supervision: A Data-Efficient Maximum Likelihood Grasp Sampling Loss · ICRA 2022
Robotics › Robot manipulation
grasping
0.612022
Learn to Grasp with Less Supervision: A Data-Efficient Maximum Likelihood Grasp Sampling Loss · ICRA 2022
Machine learning › Transfer learning and domain adaptation › few-shot learning
few-shot transfer
0.212024
MD tree: a model-diagnostic tree grown on loss landscape · ICML 2024
Machine learning › Learning theory
generalization bounds
0.212024
Sharpness-diversity tradeoff: improving flat ensembles with SharpBalance · NeurIPS 2024
Computational science and engineering
scientific machine learning
0.212024
Model Balancing Helps Low-data Training and Fine-tuning · EMNLP 2024
Machine learning › Deep learning architectures and training
convolutional neural network
0.212022
Learn to Grasp with Less Supervision: A Data-Efficient Maximum Likelihood Grasp Sampling Loss · ICRA 2022

Methods — techniques the papers use, named apart from their topics

heavy-tailed self-regularization theory · 3.7layer-wise learning rate scheduling · 1.5sharpness-aware minimization · 1.4training-free expert allocation · 0.8theoretical analysis · 0.8loss landscape metrics · 0.8ensemble training · 0.8empirical spectral density · 0.8classification · 0.8statistical mechanics · 0.7
YearPublicationVenuePosition
2024 Model Balancing Helps Low-data Training and Fine-tuning
abstract
Recent advances in foundation models have emphasized the need to align pre-trained models with specialized domains using small, curated datasets.Studies on these foundation models underscore the importance of low-data training and fine-tuning.This topic, well-known in natural language processing (NLP), has also gained increasing attention in the emerging field of scientific machine learning (SciML).To address the limitations of low-data training and fine-tuning, we draw inspiration from Heavy-Tailed Self-Regularization (HT-SR) theory, analyzing the shape of empirical spectral densities (ESDs) and revealing an imbalance in training quality across different model layers.To mitigate this issue, we adapt a recently proposed layer-wise learning rate scheduler, TempBalance, which effectively balances training quality across layers and enhances low-data training and fine-tuning for both NLP and SciML tasks.Notably, TempBalance demonstrates increasing performance gains as the amount of available tuning data decreases.Comparative analyses further highlight the effectiveness of TempBalance and its adaptability as an "add-on" method for improving model performance.
Tianyu Pang, Yefan Zhou, Pu Ren, Yaoqing Yang 0002
EMNLP4
2024 AlphaLoRA: Assigning LoRA Experts Based on Layer Training Quality
abstract
Parameter-efficient fine-tuning methods, such as Low-Rank Adaptation (LoRA), are known to enhance training efficiency in Large Language Models (LLMs).Due to the limited parameters of LoRA, recent studies seek to combine LoRA with Mixture-of-Experts (MoE) to boost performance across various tasks.However, inspired by the observed redundancy in traditional MoE structures, prior studies find that LoRA experts within the MoE architecture also exhibit redundancy, suggesting a need to vary the allocation of LoRA experts across different layers.In this paper, we leverage Heavy-Tailed Self-Regularization (HT-SR) Theory to design a fine-grained allocation strategy.Our analysis reveals that the number of experts per layer correlates with layer training quality, which exhibits significant variability across layers.Based on this, we introduce AlphaLoRA, a theoretically principled and training-free method for allocating LoRA experts to reduce redundancy further.Experiments on three models across ten language processing and reasoning benchmarks demonstrate that AlphaLoRA achieves comparable or superior performance over all baselines.Our code is available at https://github.com/morelife2017/alphalora.
Peijun Qing, Chongyang Gao, Yefan Zhou, Xingjian Diao, Yaoqing Yang 0002, Soroush Vosoughi
EMNLP3
2024 MD tree: a model-diagnostic tree grown on loss landscape
abstract
This paper considers ”model diagnosis”, which we formulate as a classification problem. Given a pre-trained neural network (NN), the goal is to predict the source of failure from a set of failure modes (such as a wrong hyperparameter, inadequate model size, and insufficient data) without knowing the training configuration of the pre-trained NN. The conventional diagnosis approach uses training and validation errors to determine whether the model is underfitting or overfitting. However, we show that rich information about NN performance is encoded in the optimization loss landscape, which provides more actionable insights than validation-based measurements. Therefore, we propose a diagnosis method called MD tree based on loss landscape metrics and experimentally demonstrate its advantage over classical validation-based approaches. We verify the effectiveness of MD tree in multiple practical scenarios: (1) use several models trained on one dataset to diagnose a model trained on another dataset, essentially a few-shot dataset transfer problem; (2) use small models (or models trained with small data) to diagnose big models (or models trained with big data), essentially a scale transfer problem. In a dataset transfer task, MD tree achieves an accuracy of 87.7%, outperforming validation-based approaches by 14.88%. Our code is available at https://github.com/YefanZhou/ModelDiagnosis.
Yefan Zhou, Qinxue Cao, Konstantin Schürholt, Yaoqing Yang 0002
ICML1
2024 Sharpness-diversity tradeoff: improving flat ensembles with SharpBalance
abstract
Recent studies on deep ensembles have identified the sharpness of the local minima of individual learners and the diversity of the ensemble members as key factors in improving test-time performance. Building on this, our study investigates the interplay between sharpness and diversity within deep ensembles, illustrating their crucial role in robust generalization to both in-distribution (ID) and out-of-distribution (OOD) data. We discover a trade-off between sharpness and diversity: minimizing the sharpness in the loss landscape tends to diminish the diversity of individual members within the ensemble, adversely affecting the ensemble's improvement. The trade-off is justified through our rigorous theoretical analysis and verified empirically through extensive experiments. To address the issue of reduced diversity, we introduce SharpBalance, a novel training approach that balances sharpness and diversity within ensembles. Theoretically, we show that our training strategy achieves a better sharpness-diversity trade-off. Empirically, we conducted comprehensive evaluations in various data sets (CIFAR-10, CIFAR-100, TinyImageNet) and showed that SharpBalance not only effectively improves the sharpness-diversity trade-off but also significantly improves ensemble performance in ID and OOD scenarios.
Haiquan Lu, Xiaotian Liu, Yefan Zhou, Qunli Li, Kurt Keutzer, Michael W. Mahoney, Yujun Yan, Huanrui Yang, Yaoqing Yang 0002
NeurIPS3
2024 AlphaPruning: Using Heavy-Tailed Self Regularization Theory for Improved Layer-wise Pruning of Large Language Models
abstract
Recent work on pruning large language models (LLMs) has shown that one can eliminate a large number of parameters without compromising performance, making pruning a promising strategy to reduce LLM model size. Existing LLM pruning strategies typically assign uniform pruning ratios across layers, limiting overall pruning ability; and recent work on layerwise pruning of LLMs is often based on heuristics that can easily lead to suboptimal performance. In this paper, we leverage Heavy-Tailed Self-Regularization (HT-SR) Theory, in particular the shape of empirical spectral densities (ESDs) of weight matrices, to design improved layerwise pruning ratios for LLMs. Our analysis reveals a wide variability in how well-trained, and thus relatedly how prunable, different layers of an LLM are. Based on this, we propose AlphaPruning, which uses shape metrics to allocate layerwise sparsity ratios in a more theoretically-principled manner. AlphaPruning can be used in conjunction with multiple existing LLM pruning methods. Our empirical results show that AlphaPruning prunes LLaMA-7B to 80% sparsity while maintaining reasonable perplexity, marking a first in the literature on LLMs.
Haiquan Lu, Yefan Zhou, Shiwei Liu 0003, Zhangyang Wang, Michael W. Mahoney, Yaoqing Yang 0002
NeurIPS2
2023 A Three-regime Model of Network Pruning
abstract
Recent work has highlighted the complex influence training hyperparameters, e.g., the number of training epochs, can have on the prunability of machine learning models. Perhaps surprisingly, a systematic approach to predict precisely how adjusting a specific hyperparameter will affect prunability remains elusive. To address this gap, we introduce a phenomenological model grounded in the statistical mechanics of learning. Our approach uses temperature-like and load-like parameters to model the impact of neural network (NN) training hyperparameters on pruning performance. A key empirical result we identify is a sharp transition phenomenon: depending on the value of a load-like parameter in the pruned model, increasing the value of a temperature-like parameter in the pre-pruned model may either enhance or impair subsequent pruning performance. Based on this transition, we build a three-regime model by taxonomizing the global structure of the pruned NN loss landscape. Our model reveals that the dichotomous effect of high temperature is associated with transitions between distinct types of global structures in the post-pruned model. Based on our results, we present three case-studies: 1) determining whether to increase or decrease a hyperparameter for improved pruning; 2) selecting the best model to prune from a family of models; and 3) tuning the hyperparameter of the Sharpness Aware Minimization method for better pruning performance.
Yefan Zhou, Yaoqing Yang 0002, Arin Chang, Michael W. Mahoney
ICML1
2023 Temperature Balancing, Layer-wise Weight Analysis, and Neural Network Training
abstract
Regularization in modern machine learning is crucial, and it can take various forms in algorithmic design: training set, model family, error function, regularization terms, and optimizations. In particular, the learning rate, which can be interpreted as a temperature-like parameter within the statistical mechanics of learning, plays a crucial role in neural network training. Indeed, many widely adopted training strategies basically just define the decay of the learning rate over time. This process can be interpreted as decreasing a temperature, using either a global learning rate (for the entire model) or a learning rate that varies for each parameter. This paper proposes TempBalance, a straightforward yet effective layer-wise learning rate method. TempBalance is based on Heavy-Tailed Self-Regularization (HT-SR) Theory, an approach which characterizes the implicit self-regularization of different layers in trained models. We demonstrate the efficacy of using HT-SR-motivated metrics to guide the scheduling and balancing of temperature across all network layers during model training, resulting in improved performance during testing. We implement TempBalance on CIFAR10, CIFAR100, SVHN, and TinyImageNet datasets using ResNets, VGGs and WideResNets with various depths and widths. Our results show that TempBalance significantly outperforms ordinary SGD and carefully-tuned spectral norm regularization. We also show that TempBalance outperforms a number of state-of-the-art optimizers and learning rate schedulers.
Yefan Zhou, Tianyu Pang, Keqin Liu, Charles H. Martin, Michael W. Mahoney, Yaoqing Yang 0002
NeurIPS1
2022 Learn to Grasp with Less Supervision: A Data-Efficient Maximum Likelihood Grasp Sampling Loss
abstract
Robotic grasping for a diverse set of objects is essential in many robot manipulation tasks. One promising approach is to learn deep grasping models from large training datasets of object images and grasp labels. However, empirical grasping datasets are typically sparsely labeled (i.e., a small number of successful grasp labels**Labels refer to marking the image to indicate a successful robotic grasp. in each image). The data sparsity issue can lead to insufficient supervision and false-negative labels, and thus results in poor learning results. This paper proposes a Maximum Likelihood Grasp Sampling Loss (MLGSL) to tackle the data sparsity issue. The proposed method supposes that successful grasps are stochastically sampled from the predicted grasp distribution and maximizes the observing likelihood. MLGSL is utilized for training a fully convolutional network that generates thousands of grasps simultaneously. Training results suggest that models based on MLGSL can learn to grasp with datasets composing of 2 labels per image. Compared to previous works, which require training datasets of 16 labels per image, MLGSL is 8× more data-efficient. Meanwhile, physical robot experiments demonstrate an equivalent performance at a 90.7% grasp success rate on household objects. Codes and videos are available at [1].
Xinghao Zhu, Yefan Zhou, Yongxiang Fan, Lingfeng Sun, Jianyu Chen 0002, Masayoshi Tomizuka
ICRA2
2021 A Dataset-Dispersion Perspective on Reconstruction Versus Recognition in Single-View 3D Reconstruction Networks
abstract
Neural networks (NN) for single-view 3D reconstruction (SVR) have gained in popularity. Recent work points out that for SVR, most cutting-edge NNs have limited performance on reconstructing unseen objects because they rely primarily on recognition (i.e., classification-based methods) rather than shape reconstruction. To understand this issue in depth, we provide a systematic study on when and why NNs prefer recognition to reconstruction and vice versa. Our finding shows that a leading factor in determining recognition versus reconstruction is how “dispersed” the training data is. Thus, we introduce the dispersion score, a new data-driven metric, to quantify this leading factor and study its effect on NNs. We hypothesize that NNs are biased toward recognition when training images are more dispersed and training shapes are less dispersed. Our hypothesis is supported and the dispersion score is proved effective through our experiments on synthetic and benchmark datasets. We show that the proposed metric is a principal way to analyze reconstruction quality and provides novel information in addition to the conventional reconstruction score. We have open-sourced our code.1
Yefan Zhou, Yiru Shen, Yujun Yan, Chen Feng 0002, Yaoqing Yang 0002
3DV1