Yanping Huang

dblp:00/10104 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
21since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 7 first-author · 16 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Agentix: An Efficient Serving Engine for LLM Agents as General Programs
Michael Luo, Xiaoxiang Shi, Colin Cai, Tianjun Zhang, Justin Wong, Yanping Huang, Joseph Gonzalez 0001, Ion Stoica
NSDI8
2026 Physics informed Dual-Layer Bidirectional Gated Recurrent Unit for Nuclear-Grade Electric Gate Valves Fault Prognostics
Chenwei Tang, Jiancheng Lv 0001, Yanping Huang, Yanshan Li
Eng. Appl. Artif. Intell.5
2026 Decoupled hypergraph modeling for multimodal sentiment analysis
Yanping Huang, Jiawen Deng 0006, Yan Zhuang 0002, Jiali You 0002, Fuji Ren
Neurocomputing1
2025 A graph attention network integrated with gated spiking neural P systems for session-based recommendation
Xinzhu Bai, Hong Peng 0001, Yanping Huang, Jun Wang 0013, Qian Yang 0002, Antonio Ramírez-de-Arellano
Expert Syst. Appl.3
2024 Controlled Decoding from Language Models
abstract
KL-regularized reinforcement learning (RL) is a popular alignment framework to control the language model responses towards high reward outcomes. We pose a tokenwise RL objective and propose a modular solver for it, called *controlled decoding (CD)*. CD exerts control through a separate *prefix scorer* module, which is trained to learn a value function for the reward. The prefix scorer is used at inference time to control the generation from a frozen base model, provably sampling from a solution to the RL objective. We empirically demonstrate that CD is effective as a control mechanism on popular benchmarks. We also show that prefix scorers for multiple rewards may be combined at inference time, effectively solving a multi-objective RL problem with no additional training. We show that the benefits of applying CD transfer to an unseen base model with no further tuning as well. Finally, we show that CD can be applied in a blockwise decoding fashion at inference-time, essentially bridging the gap between the popular best-of-$K$ strategy and tokenwise control through reinforcement learning. This makes CD a promising approach for alignment of language models.
Sidharth Mudgal, Jong Lee, Harish Ganapathy, YaGuang Li, Yanping Huang, Heng-Tze Cheng, Trevor Strohman, Jilin Chen, Alex Beutel, Ahmad Beirami
ICML6
2024 Stylus: Automatic Adapter Selection for Diffusion Models
abstract
Beyond scaling base models with more data or parameters, fine-tuned adapters provide an alternative way to generate high fidelity, custom images at reduced costs. As such, adapters have been widely adopted by open-source communities, accumulating a database of over 100K adapters—most of which are highly customized with insufficient descriptions. To generate high quality images, this paper explores the problem of matching the prompt to a Stylus of relevant adapters, built on recent work that highlight the performance gains of composing adapters. We introduce Stylus, which efficiently selects and automatically composes task-specific adapters based on a prompt's keywords. Stylus outlines a three-stage approach that first summarizes adapters with improved descriptions and embeddings, retrieves relevant adapters, and then further assembles adapters based on prompts' keywords by checking how well they fit the prompt. To evaluate Stylus, we developed StylusDocs, a curated dataset featuring 75K adapters with pre-computed adapter embeddings. In our evaluation on popular Stable Diffusion checkpoints, Stylus achieves greater CLIP/FID Pareto efficiency and is twice as preferred, with humans and multimodal models as evaluators, over the base model.
Michael Luo, Justin Wong, Brandon Trabucco, Yanping Huang, Joseph Gonzalez 0001, Ruslan Salakhutdinov, Ion Stoica
NeurIPS4
2024 Sequence recommendation using multi-level self-attention network with gated spiking neural P systems
Xinzhu Bai, Yanping Huang, Hong Peng 0001, Jun Wang 0013, Qian Yang 0002, David Orellana-Martín, Antonio Ramírez-de-Arellano, Mario J. Pérez-Jiménez
Inf. Sci.2
2024 Scaling Instruction-Finetuned Language Models
abstract
Finetuning language models on a collection of datasets phrased as instructions has been shown to improve model performance and generalization to unseen tasks. In this paper we explore instruction finetuning with a particular focus on (1) scaling the number of tasks, (2) scaling the model size, and (3) finetuning on chain-of-thought data. We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks (MMLU, BBH, TyDiQA, MGSM, open-ended generation, RealToxicityPrompts). For instance, Flan-PaLM 540B instruction-finetuned on 1.8K tasks outperforms PaLM 540B by a large margin (+9.4% on average). Flan-PaLM 540B achieves state-of-the-art performance on several benchmarks (at time of release), such as 75.2% on five-shot MMLU. We also publicly release Flan-T5 checkpoints,1 which achieve strong few-shot performance even compared to much larger models, such as PaLM 62B. Overall, instruction finetuning is a general method for improving the performance and usability of pretrained language models.
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Yunxuan Li, Xuezhi Wang 0002, Mostafa Dehghani 0001, Siddhartha Brahma, Albert Webson, Shixiang Gu, Zhuyun Dai, Mirac Suzgun, Aakanksha Chowdhery, Alex Castro-Ros, Marie Pellat, Kevin Robinson, Dasha Valter, Sharan Narang, Adams Yu, Vincent Y. Zhao, Yanping Huang, Andrew M. Dai, Hongkun Yu 0001, Slav Petrov, Ed H. Chi, Jeffrey Dean, Jacob Devlin, Adam Roberts, Denny Zhou, Quoc V. Le, Jason Wei
J. Mach. Learn. Res.25
2023 Massively Multilingual Shallow Fusion with Large Language Models
abstract
While large language models (LLM) have made impressive progress in natural language processing, it remains unclear how to utilize them in improving automatic speech recognition (ASR). In this work, we propose to train a single multilingual language model (LM) for shallow fusion in multiple languages. We push the limits of the multilingual LM to cover up to 84 languages by scaling up using a mixture-of-experts LLM, i.e., generalist language model (GLaM). When the number of experts increases, GLaM dynamically selects only two at each decoding step to keep the inference computation roughly constant. We then apply GLaM to a multilingual shallow fusion task based on a state-of-the-art end-to-end model. Compared to a dense LM of similar computation during inference, GLaM reduces the WER of an English long-tail test set by 4.4% relative. In a multilingual shallow fusion task, GLaM improves 41 out of 50 languages with an average relative WER reduction of 3.85%, and a maximum reduction of 10%. Compared to the baseline model, GLaM achieves an average WER reduction of 5.53% over 43 languages.
Tara N. Sainath, Bo Li 0028, Nan Du 0002, Yanping Huang, Andrew M. Dai, Yu Zhang 0033, Rodrigo Cabrera, Trevor Strohman
ICASSP5
2023 Lifelong Language Pretraining with Distribution-Specialized Experts
abstract
Pretraining on a large-scale corpus has become a standard method to build general language models (LMs). Adapting a model to new data distributions targeting different downstream tasks poses significant challenges. Naive fine-tuning may incur catastrophic forgetting when the over-parameterized LMs overfit the new data but fail to preserve the pretrained features. Lifelong learning (LLL) aims to enable information systems to learn from a continuous data stream across time. However, most prior work modifies the training recipe assuming a static fixed network architecture. We find that additional model capacity and proper regularization are key elements to achieving strong LLL performance. Thus, we propose Lifelong-MoE, an extensible MoE (Mixture-of-Experts) architecture that dynamically adds model capacity via adding experts with regularized pretaining. Our results show that by only introducing a limited number of extra experts while keeping the computation cost constant, our model can steadily adapt to data distribution shifts while preserving the previous knowledge. Compared to existing lifelong learning approaches, Lifelong-MoE achieves better few-shot performance on NLP tasks. More impressively, Lifelong-MoE surpasses multi-task learning on 19 downstream NLU tasks.
Wuyang Chen 0001, Yanqi Zhou, Nan Du 0002, Yanping Huang, James Laudon, Claire Cui
ICML4
2023 Brainformers: Trading Simplicity for Efficiency
abstract
Transformers are central to recent successes in natural language processing and computer vision. Transformers have a mostly uniform backbone where layers alternate between feed-forward and self-attention in order to build a deep network. Here we investigate this design choice and find that more complex blocks that have different permutations of layer primitives can be more efficient. Using this insight, we develop a complex block, named Brainformer, that consists of a diverse sets of layers such as sparsely gated feed-forward layers, dense feed-forward layers, attention layers, and various forms of layer normalization and activation functions. Brainformer consistently outperforms the state-of-the-art dense and sparse Transformers, in terms of both quality and efficiency. A Brainformer model with 8 billion activated parameters per token demonstrates 2x faster training convergence and 5x faster step time compared to its GLaM counterpart. In downstream task evaluation, Brainformer also demonstrates a 3% higher SuperGLUE score with fine-tuning compared to GLaM with a similar number of activated parameters. Finally, Brainformer largely outperforms a Primer dense model derived with NAS with similar computation per token on fewshot evaluations.
Yanqi Zhou, Nan Du 0002, Yanping Huang, Daiyi Peng, Chang Lan, Siamak Shakeri, David R. So, Andrew M. Dai, Yifeng Lu, Quoc V. Le, Claire Cui, James Laudon, Jeffrey Dean
ICML3
2023 SPAE: Semantic Pyramid AutoEncoder for Multimodal Generation with Frozen LLMs
abstract
In this work, we introduce Semantic Pyramid AutoEncoder (SPAE) for enabling frozen LLMs to perform both understanding and generation tasks involving non-linguistic modalities such as images or videos. SPAE converts between raw pixels and interpretable lexical tokens (or words) extracted from the LLM's vocabulary. The resulting tokens capture both the rich semantic meaning and the fine-grained details needed for visual reconstruction, effectively translating the visual content into a language comprehensible to the LLM, and empowering it to perform a wide array of multimodal tasks. Our approach is validated through in-context learning experiments with frozen PaLM 2 and GPT 3.5 on a diverse set of image understanding and generation tasks. Our method marks the first successful attempt to enable a frozen LLM to generate image content while surpassing state-of-the-art performance in image understanding tasks, under the same setting, by over 25%.
Lijun Yu, Yong Cheng 0003, Zhiruo Wang 0001, Wolfgang Macherey, Yanping Huang, David A. Ross, Irfan A. Essa, Yonatan Bisk, Ming-Hsuan Yang 0001, Kevin Murphy 0002, Alex Hauptmann 0001, Lu Jiang 0004
NeurIPS6
2023 AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
Zhuohan Li 0001, Lianmin Zheng, Yinmin Zhong, Vincent Liu 0001, Ying Sheng 0007, Xin Jin 0008, Yanping Huang, Hao Zhang 0025, Joseph Gonzalez 0001, Ion Stoica
OSDI7
2023 Sentiment classification using bidirectional LSTM-SNP model and attention mechanism
Yanping Huang, Qian Liu 0034, Hong Peng 0001, Jun Wang 0013, Qian Yang 0002, David Orellana-Martín
Expert Syst. Appl.1
2023 An Attention-Aware Long Short-Term Memory-Like Spiking Neural Model for Sentiment Analysis
abstract
LSTM-SNP model is a recently developed long short-term memory (LSTM) network, which is inspired from the mechanisms of spiking neural P (SNP) systems. In this paper, LSTM-SNP is utilized to propose a novel model for aspect-level sentiment analysis, termed as ALS model. The LSTM-SNP model has three gates: reset gate, consumption gate and generation gate. Moreover, attention mechanism is integrated with LSTM-SNP model. The ALS model can better capture the sentiment features in the text to compute the correlation between context and aspect words. To validate the effectiveness of the ALS model for aspect-level sentiment analysis, comparison experiments with 17 baseline models are conducted on three real-life data sets. The experimental results demonstrate that the ALS model has a simpler structure and can achieve better performance compared to these baseline models.
Qian Liu 0034, Yanping Huang, Qian Yang 0002, Hong Peng 0001, Jun Wang 0013
Int. J. Neural Syst.2
2023 Attention-enabled gated spiking neural P model for aspect-level sentiment classification
Yanping Huang, Hong Peng 0001, Qian Liu 0034, Qian Yang 0002, Jun Wang 0013, David Orellana-Martín, Mario J. Pérez-Jiménez
Neural Networks1
2022 GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
abstract
Scaling language models with more data, compute and parameters has driven significant progress in natural language processing. For example, thanks to scaling, GPT-3 was able to achieve strong results on in-context learning tasks. However, training these large dense models requires significant amounts of computing resources. In this paper, we propose and develop a family of language models named \glam (\textbf{G}eneralist \textbf{La}nguage \textbf{M}odel), which uses a sparsely activated mixture-of-experts architecture to scale the model capacity while also incurring substantially less training cost compared to dense variants. The largest \glam has 1.2 trillion parameters, which is approximately 7x larger than GPT-3. It consumes only 1/3 of the energy used to train GPT-3 and requires half of the computation flops for inference, while still achieving better overall fewshot performance across 29 NLP tasks.
Nan Du 0002, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, William Fedus, Maarten Bosma, Zongwei Zhou, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathy Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang 0043, Quoc V. Le, Claire Cui
ICML2
2022 Mixture-of-Experts with Expert Choice Routing
abstract
Sparsely-activated Mixture-of-experts (MoE) models allow the number of parameters to greatly increase while keeping the amount of computation for a given token or a given sample unchanged. However, a poor expert routing strategy (e.g. one resulting in load imbalance) can cause certain experts to be under-trained, leading to an expert being under or over-specialized. Prior work allocates a fixed number of experts to each token using a top-k function regardless of the relative importance of different tokens. To address this, we propose a heterogeneous mixture-of-experts employing an expert choice method. Instead of letting tokens select the top-k experts, we have experts selecting the top-k tokens. As a result, each token can be routed to a variable number of experts and each expert can have a fixed bucket size. We systematically study pre-training speedups using the same computational resources of the Switch Transformer top-1 and GShard top-2 gating of prior work and find that our method improves training convergence time by more than 2×. For the same computational cost, our method demonstrates higher performance in fine-tuning 11 selected tasks in the GLUE and SuperGLUE benchmarks. For a smaller activation cost, our method outperforms the T5 dense model in 7 out of the 11 tasks.
Yanqi Zhou, Tao Lei 0001, Hanxiao Liu, Nan Du 0002, Yanping Huang, Vincent Y. Zhao, Andrew M. Dai, Quoc V. Le, James Laudon
NeurIPS5
2022 Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep Learning
Lianmin Zheng, Zhuohan Li 0001, Hao Zhang 0025, Yonghao Zhuang 0001, Yanping Huang, Yida Wang 0003, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, Joseph Gonzalez 0001, Ion Stoica
OSDI6
2021 GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer
ICLR6
2021 Do CNNs Encode Data Augmentations?
abstract
Data augmentations are important ingredients in the recipe for training robust neural networks, especially in computer vision. A fundamental question is whether neural network features encode data augmentation transformations. To answer this question, we introduce a systematic approach to investigate which layers of neural networks are the most predictive of augmentation transformations. Our approach uses features in pre-trained vision models with minimal additional processing to predict common properties transformed by augmentation (scale, aspect ratio, hue, saturation, contrast, and brightness). Surprisingly, neural network features not only predict data augmentation transformations, but they predict many transformations with high accuracy. After validating that neural networks encode features corresponding to augmentation transformations, we show that these features are encoded in the early layers of modern CNNs, though the augmentation signal fades in deeper layers.
Eddie Yan, Yanping Huang
IJCNN2
2020 Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
abstract
The vast majority of deep models use multiple gradient signals, typically corresponding to a sum of multiple loss terms, to update a shared set of trainable weights. However, these multiple updates can impede optimal training by pulling the model in conflicting directions. We present Gradient Sign Dropout (GradDrop), a probabilistic masking procedure which samples gradients at an activation layer based on their level of consistency. GradDrop is implemented as a simple deep layer that can be used in any deep net and synergizes with other gradient balancing approaches. We show that GradDrop outperforms the state-of-the-art multiloss methods within traditional multitask and transfer learning settings, and we discuss how GradDrop reveals links between optimal multiloss training and gradient stochasticity.
Jiquan Ngiam, Yanping Huang, Thang Luong, Henrik Kretzschmar, Yuning Chai, Dragomir Anguelov
NeurIPS3
2019 Regularized Evolution for Image Classifier Architecture Search
abstract
The effort devoted to hand-crafting neural network image classifiers has motivated the use of architecture search to discover them automatically. Although evolutionary algorithms have been repeatedly applied to neural network topologies, the image classifiers thus discovered have remained inferior to human-crafted ones. Here, we evolve an image classifier— AmoebaNet-A—that surpasses hand-designs for the first time. To do this, we modify the tournament selection evolutionary algorithm by introducing an age property to favor the younger genotypes. Matching size, AmoebaNet-A has comparable accuracy to current state-of-the-art ImageNet models discovered with more complex architecture-search methods. Scaled to larger size, AmoebaNet-A sets a new state-of-theart 83.9% top-1 / 96.6% top-5 ImageNet accuracy. In a controlled comparison against a well known reinforcement learning algorithm, we give evidence that evolution can obtain results faster with the same hardware, especially at the earlier stages of the search. This is relevant when fewer compute resources are available. Evolution is, thus, a simple method to effectively discover high-quality architectures.
Esteban Real, Alok Aggarwal, Yanping Huang, Quoc V. Le
AAAI3
2019 GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
abstract
Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or infrastructure. These solutions are often architecture-specific and do not transfer to other machine learning tasks. To address the need for efficient and task-independent model parallelism, we introduce TensorPipe, a pipeline parallelism library that allows scaling any network that can be expressed as a sequence of layers. By pipelining different sub-sequences of layers on separate accelerators, TensorPipe provides the flexibility of scaling a variety of different networks to gigantic sizes efficiently. Moreover, TensorPipe utilizes a novel batch-splitting pipelining algorithm, resulting in almost linear speedup when a model is partitioned across multiple accelerators. We demonstrate the advantages of TensorPipe by training large-scale neural networks on two different tasks with distinct network architectures: (i)Image Classification: We train a 557-million-parameter AmoebaNet model and attain a top-1 accuracy of 84.4% on ImageNet-2012, (ii)Multilingual Neural Machine Translation: We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models.
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le
NeurIPS1
2016 Bayesian Inference and Online Learning in Poisson Neuronal Networks
abstract
Motivated by the growing evidence for Bayesian computation in the brain, we show how a two-layer recurrent network of Poisson neurons can perform both approximate Bayesian inference and learning for any hidden Markov model. The lower-layer sensory neurons receive noisy measurements of hidden world states. The higher-layer neurons infer a posterior distribution over world states via Bayesian inference from inputs generated by sensory neurons. We demonstrate how such a neuronal network with synaptic plasticity can implement a form of Bayesian inference similar to Monte Carlo methods such as particle filtering. Each spike in a higher-layer neuron represents a sample of a particular hidden world state. The spiking activity across the neural population approximates the posterior distribution over hidden states. In this model, variability in spiking is regarded not as a nuisance but as an integral feature that provides the variability necessary for sampling during inference. We demonstrate how the network can learn the likelihood model, as well as the transition probabilities underlying the dynamics, using a Hebbian learning rule. We present results illustrating the ability of the network to perform inference and learning for arbitrary hidden Markov models.
Yanping Huang, Rajesh P. N. Rao
Neural Comput.1
2014 Neurons as Monte Carlo Samplers: Bayesian Inference and Learning in Spiking Networks
Yanping Huang, Rajesh P. N. Rao
NIPS1
2012 How Prior Probability Influences Decision Making: A Unifying Probabilistic Model
abstract
How does the brain combine prior knowledge with sensory evidence when making decisions under uncertainty? Two competing descriptive models have been proposed based on experimental data. The first posits an additive offset to a decision variable, implying a static effect of the prior. However, this model is inconsistent with recent data from a motion discrimination task involving temporal integration of uncertain sensory evidence. To explain this data, a second model has been proposed which assumes a time-varying influence of the prior. Here we present a normative model of decision making that incorporates prior knowledge in a principled way. We show that the additive offset model and the time-varying prior model emerge naturally when decision making is viewed within the framework of partially observable Markov decision processes (POMDPs). Decision making in the model reduces to (1) computing beliefs given observations and prior information in a Bayesian manner, and (2) selecting actions based on these beliefs to maximize the expected sum of future rewards. We show that the model can explain both data previously explained using the additive offset model as well as more recent data on the time-varying influence of prior knowledge on decision making.
Yanping Huang, Abram L. Friesen, Timothy D. Hanks, Michael N. Shadlen, Rajesh P. N. Rao
NIPS1