VLDB 2026 Research / reviewers in the wild / expert
Tuo Zhao
dblp:82/199
· DBLP profile ↗
124ranked-venue papers
9as first author
70since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 111 · 7 first-author · 64 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Theory of computation · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM AlignmentabstractReward modeling lies at the core of reinforcement learning from human feedback (RLHF), yet most existing reward models rely on scalar or pairwise judgments that fail to capture the multifaceted nature of human preferences.Recent studies have explored rubrics-as-rewards (RaR) that uses structured criteria to capture multiple dimensions of response quality.However, producing rubrics that are both reliable and scalable remains a key challenge.In this work, we introduce OpenRubrics, a diverse, large-scale collection of (prompt, rubric) pairs for training rubric-generation and rubric-based reward models.To elicit discriminative and comprehensive evaluation signals, we introduce Contrastive Rubric Generation (CRG), which derives both hard rules (explicit constraints) and principles (implicit qualities) by contrasting preferred and rejected responses.We further remove noisy rubrics via preserving preference-label consistency.Across multiple reward-modeling benchmarks, our rubricbased reward model, RUBRIC-RM, surpasses strong size-matched baselines by 8.4%.These gains transfer to policy models on instructionfollowing and biomedical benchmarks.The model weights and datasets are publicly available at https://huggingface.co/OpenRubrics.* These authors contributed equally to this work, order was determined randomly (by rolling a die). Tianci Liu 0003, Ran Xu 0002, Tony Yu, Ilgee Hong, Carl Yang 0001, Tuo Zhao, Haoyu Wang 0004 |
ACL (1) | 6 |
| 2026 | CoLLM: Industrial Large-Small Model Collaboration With Fuzzy Decision-Making Agent and Self-ReflectionabstractIn industrial applications, large models have exhibited superior generalization capabilities that are unattainable with smaller models. However, when faced with edge scenarios and highly diverse industrial samples, their deployment remains challenging due to high computational costs and unreliable output. To address these challenges, we propose CoLLM, a fuzzy large-small model collaborative framework, which dynamically selects between small and large models based on the characteristics exhibited by the samples. Specifically, this approach estimates uncertainty from input samples to guide model selection: low-uncertainty samples are processed by the small model for efficiency, while high-uncertainty or complex samples are routed to the large model for improved accuracy. It first constructs a fuzzy decision-making agent based on the fuzzy neural network (FNN) to assess sample complexity and determine the appropriate model for inference. Furthermore, a self-reflection mechanism is proposed to refine the large model's output, reducing the risk of unreliable output. Experimental results in industrial time series datasets demonstrate that our framework improves the computational efficiency of large models up to 14.54x while maintaining or improving prediction accuracy. Haiteng Wang, Lei Ren 0001, Tuo Zhao, Lu Jiao |
IEEE Trans. Fuzzy Syst. | 3 |
| 2026 | AMR-Net: Adaptive Temporal-Channel Multiresolution Network for Industrial Time-Series PredictionabstractAccurate and fast industrial time-series prediction is essential for safe and reliable operation of industrial equipment. Recent deep learning methods enable extracting complex temporal patterns by utilizing large-scale parameters and multiresolution feature extraction. However, they cause substantial computational complexity and limit their application at the edge. In this article, we design an adaptive temporal-channel multiresolution network (AMR-Net) that dynamically adjusts time-series resolution to avoid redundant computation. The motivation is that low-resolution feature representations are sufficient for predicting “easy” samples, whereas “hard” samples require high-resolution features to capture fine-grained information. For the AMR-Net, time series are initially input through a temporal-channel resolution decomposition (TCRD) module, which efficiently extracts low-resolution representations. Samples exhibiting high prediction confidence are expedited through early exit mechanisms, avoiding further processing. Meanwhile, high-resolution subnetworks capture the fine-grained information to discern the “hard” samples. Experiments on CMAPSS and N-CMAPSS datasets demonstrate that AMR-Net can improve computational speed by 15x while maintaining high accuracy. Haiteng Wang, Lei Ren 0001, Tuo Zhao |
IEEE Trans. Ind. Informatics | 3 |
| 2025 | MS-RainMamba: Learning Multi-Scale State Space Models for Single Image DerainingabstractDespite the significant advances of Convolutional neural networks (CNNs) and Transformers in image deraining, they either suffer from limited receptive fields or incur quadratic complexity, leading to an imbalance between performance and efficiency. Recently, state space models (SSMs) have demonstrated significant potential in modeling long-range dependencies while maintaining linear complexity. However, existing Mamba-based approaches lack the exploration of useful complementary information from multiple image scales, which could be beneficial for facilitating rain removal. In this paper, we propose an effective multi-scale state-space model-based framework (MS-RainMamba) to explore richer scale-space information for better image deraining. Specifically, we design a local-enhanced state space module to better aggregate rich local and global information. In contrast to existing methods that adopt fixed-scale scanning for feature extraction, we develop a multi-scale hierarchical 2D scanning technique to better help image restoration. Experimental results on six benchmarks show that the proposed method performs favorably against state-of-the-art models. Zhanshuo Liu, Tuo Zhao, Tingting Zhao 0001, Yarui Chen, Ning Xie 0003 |
ICASSP | 3 |
| 2025 | Discriminative Finetuning of Generative Large Language Models without Reward Models and Human Preference DataabstractSupervised fine-tuning (SFT) has become a crucial step for aligning pretrained large language models (LLMs) using supervised datasets of input-output pairs. However, despite being supervised, SFT is inherently limited by its generative training objective. To address its limitations, the existing common strategy is to follow SFT with a separate phase of preference optimization (PO), which relies on either human-labeled preference data or a strong reward model to guide the learning process. In this paper, we address the limitations of SFT by exploring one of the most successful techniques in conventional supervised learning: discriminative learning. We introduce Discriminative Fine-Tuning (DFT), an improved variant of SFT, which mitigates the burden of collecting human-labeled preference data or training strong reward models. Unlike SFT that employs a generative approach and overlooks negative data, DFT adopts a discriminative paradigm that increases the probability of positive answers while suppressing potentially negative ones, aiming for data prediction instead of token prediction. Our contributions include: (i) a discriminative probabilistic framework for fine-tuning LLMs by explicitly modeling the discriminative likelihood of an answer among all possible outputs given an input; (ii) efficient algorithms to optimize this discriminative likelihood; and (iii) extensive experiments demonstrating DFT’s effectiveness, achieving performance better than SFT and comparable to if not better than SFT$\rightarrow$PO. The code can be found at https://github.com/Optimization-AI/DFT. Siqi Guo 0003, Ilgee Hong, Vicente Balmaseda, Changlong Yu, Haoming Jiang, Tuo Zhao, Tianbao Yang |
ICML | 8 |
| 2025 | Deep Reinforcement Learning from Hierarchical Preference DesignabstractReward design is a fundamental, yet challenging aspect of reinforcement learning (RL). Researchers typically utilize feedback signals from the environment to handcraft a reward function, but this process is not always effective due to the varying scale and intricate dependencies of the feedback signals. This paper shows by exploiting certain structures, one can ease the reward design process. Specifically, we propose a hierarchical reward design framework – HERON for scenarios: (I) The feedback signals naturally present hierarchy; (II) The reward is sparse, but with less important surrogate feedback to help policy learning. Both scenarios allow us to design a hierarchical decision tree induced by the importance ranking of the feedback signals to compare RL trajectories. With such preference data, we can then train a reward model for policy learning. We apply HERON to several RL applications, and we find that our framework can not only train high performing agents on a variety of difficult tasks, but also provide additional benefits such as improved sample efficiency and robustness. Alexander Bukharin, Tuo Zhao |
ICML | 4 |
| 2025 | Think-RM: Enabling Long-Horizon Reasoning in Generative Reward ModelsabstractReinforcement learning from human feedback (RLHF) has become a powerful post-training paradigm for aligning large language models with human preferences. A core challenge in RLHF is constructing accurate reward signals, where the conventional Bradley-Terry reward models (BT RMs) often suffer from sensitivity to data size and coverage, as well as vulnerability to reward hacking. Generative reward models (GenRMs) offer a more robust alternative by generating chain-of-thought (CoT) rationales followed by a final verdict. However, existing GenRMs rely on shallow, vertically scaled reasoning, limiting their capacity to handle nuanced or complex tasks. Moreover, their pairwise preference outputs are incompatible with standard RLHF algorithms that require pointwise reward signals. In this work, we introduce Think-RM, a training framework that enables long-horizon reasoning in GenRMs by modeling an internal thinking process. Rather than producing structured, externally provided rationales, Think-RM generates flexible, self-guided reasoning traces that support advanced capabilities such as self-reflection, hypothetical reasoning, and divergent reasoning. To elicit these reasoning abilities, we first warm-up the models by supervised fine-tuning (SFT) over long CoT data. We then further improve the model's long-horizon abilities by rule-based reinforcement learning (RL). In addition, we propose a novel pairwise RLHF pipeline that directly optimizes policies from pairwise comparisons, eliminating the need for pointwise reward conversion. Experiments show that Think-RM outperforms baselines on both in-distribution and out-of-distribution tasks, with particularly strong gains on reasoning-heavy benchmarks: more than 10\% and 5\% on RewardBench's Chat Hard and Reasoning, and 12\% on RM-Bench's Math domain. When combined with our pairwise RLHF pipeline, it demonstrates superior end-policy performance compared to traditional approaches. This depth-oriented approach not only broadens the GenRM design space but also establishes a new paradigm for preference-based policy optimization in RLHF. Ilgee Hong, Changlong Yu, Weixiang Yan, Zhenghao Xu, Haoming Jiang, Qingru Zhang, Xin Liu 0039, Chao Zhang 0014, Tuo Zhao |
NeurIPS | 11 |
| 2025 | AdaSPEC: Selective Knowledge Distillation for Efficient Speculative DecodersabstractSpeculative Decoding (SD) accelerates large language model inference by employing a small draft model to generate predictions, which are then verified by a larger target model. The effectiveness of SD hinges on the alignment between these models, which is typically enhanced by Knowledge Distillation (KD). However, conventional KD methods aim to minimize the KL divergence between the draft and target models across all tokens, a goal that is misaligned with the true objective of SD, which is to maximize token acceptance rate. Therefore, draft models often struggle to fully assimilate the target model's knowledge due to capacity constraints, leading to suboptimal performance. To address this challenge, we propose AdaSPEC, a novel method that incorporates selective token filtering into the KD process. AdaSPEC utilizes a reference model to identify and filter out difficult-to-fit tokens, enabling the distillation of a draft model that better aligns with the target model on simpler tokens. This approach improves the overall token acceptance rate without compromising generation quality. We evaluate AdaSPEC across diverse tasks, including arithmetic reasoning, instruction-following, coding, and summarization, using model configurations of 31M/1.4B and 350M/2.7B parameters. Our results demonstrate that AdaSPEC consistently outperforms the state-of-the-art DistillSpec method, achieving higher acceptance rates across all tasks (up to 15\%). The code is publicly available at \url{https://github.com/yuezhouhu/adaspec}. Yuezhou Hu, Tuo Zhao |
NeurIPS | 4 |
| 2025 | A Minimalist Example of Edge-of-Stability and Progressive SharpeningabstractRecent advances in deep learning optimization have unveiled two intriguing phenomena under large learning rates: Edge of Stability (EoS) and Progressive Sharpening (PS), challenging classical Gradient Descent (GD) analyses. Current research approaches, using either generalist frameworks or minimalist examples, face significant limitations in explaining these phenomena. This paper advances the minimalist approach by introducing a two-layer network with a two-dimensional input, where one dimension is relevant to the response and the other is irrelevant. Through this model, we rigorously prove the existence of progressive sharpening and self-stabilization under large learning rates, and establish non-asymptotic analysis of the training dynamics and sharpness along the entire GD trajectory. Besides, we connect our minimalist example to existing works by reconciling the existence of a well-behaved "stable set" between minimalist and generalist analyses, and extending the analysis of Gradient Flow Solution sharpness to our two-dimensional input scenario. These findings provide new insights into the EoS phenomenon from both parameter and input data distribution perspectives, potentially informing more effective optimization strategies in deep learning practice. Simon S. Du, Tuo Zhao |
NeurIPS | 4 |
| 2025 | Ask a Strong LLM Judge when Your Reward Model is UncertainabstractReward model (RM) plays a pivotal role in reinforcement learning with human feedback (RLHF) for aligning large language models (LLMs). However, classical RMs trained on human preferences are vulnerable to reward hacking and generalize poorly to out-of-distribution (OOD) inputs.
By contrast, strong LLM judges equipped with reasoning capabilities demonstrate superior generalization, even without additional training, but incur significantly higher inference costs, limiting their applicability in online RLHF.
In this work, we propose an uncertainty-based routing framework that efficiently complements a fast RM with a strong but costly LLM judge. Our approach formulates advantage estimation in policy gradient (PG) methods as pairwise preference classification, enabling principled uncertainty quantification to guide routing. Uncertain pairs are forwarded to the LLM judge, while confident ones are evaluated by the RM. Experiments on RM benchmarks demonstrate that our uncertainty-based routing strategy significantly outperforms random judge calling at the same cost, and downstream alignment results showcase its effectiveness in improving online RLHF. Zhenghao Xu, Qingru Zhang, Ilgee Hong, Changlong Yu, Wenlin Yao, Haoming Jiang, Lihong Li 0001, Hyokun Yun, Tuo Zhao |
NeurIPS | 12 |
| 2025 | The influences of four dimensions of perceived fit on individuals' utilisation of SPOCs: an extension of the task-technology fit modelabstractSmall Private Online Courses (SPOCs) platform enables individuals to carry out their online and offline learning activities. In order to understand individuals’ utilisation of SPOCs, this study develops a research model to examine the joint influences of four dimensions of perceived fit manifested in perceived technology-task fit (TTF), perceived individual-technology fit (ITF), perceived online-offline task fit (OTF), and perceived online-offline interactivity fit (OIF). A survey is conducted at a famous university in China, and 371 data are collected from students who select courses on the SPOC platform. Structural equation modelling method is used to examine the research model. The empirical results suggest that ITF is the most significant antecedent of individual performance expectancy, followed by OTF, TTF, and OIF. Furthermore, individual performance expectancy positively influences satisfaction and continuance intention in the SPOC platform. Moreover, this study incorporates self-regulation as a moderator to explore behavioural differences between individuals with high and low self-regulation. This study extends the traditional perceived fit framework by introducing OTF and OIF and uncovers the antecedents and boundary conditions of individuals’ utilisation of SPOCs in the emerging research context. Tuo Zhao |
Behav. Inf. Technol. | 3 |
| 2025 | Good regularity creates large learning rate implicit biases: edge of stability, balancing, and catapultabstractLarge learning rates, when applied to gradient descent for nonconvex optimization, yield various implicit biases including the edge of stability, balancing, and catapult. These phenomena cannot be well explained by classical optimization theory. Though significant theoretical progress has been made in understanding these implicit biases, it remains unclear for which objective functions they are more likely to occur --- more precisely, for which functions there exists a larger set of initial conditions that lead to these phenomena? This paper provides an initial step in answering this question and also shows that these implicit biases are in fact various tips of the same iceberg. To establish these results, we develop a global convergence theory under large learning rates, for a family of nonconvex functions without globally Lipschitz continuous gradient, which was typically assumed in existing convergence analysis. Specifically, these phenomena are more likely to occur when the optimization objective function has good regularity. This regularity, together with gradient descent using a large learning rate that favors flatter regions, results in these nontrivial dynamical behaviors. Another corollary is the first non-asymptotic convergence rate bound for large-learning-rate gradient descent optimization of nonconvex functions. Although our theory only applies to specific functions so far, the possibility of extrapolating it to neural networks is also experimentally validated, for which different choices of loss, activation functions, and other techniques such as batch normalization can all affect regularity significantly and lead to very different training dynamics. Yuqing Wang 0005, Zhenghao Xu, Tuo Zhao, Molei Tao |
J. Mach. Learn. Res. | 3 |
| 2025 | FDformer: A Fuzzy Dynamic Transformer-Based Network for Efficient Industrial Time Series PredictionabstractIndustrial time series prediction is highly important for the predictive maintenance of Industrial Internet of Things devices. Deep learning methods have demonstrated state-of-the-art (SOTA) performance in the field of time series prediction. However, time series data from complex industrial scenarios often contain substantial uncertainty. This makes it difficult for deterministic deep learning models to achieve accurate predictions. Moreover, existing static methods often fail to meet the real-time requirements of industrial environments. To address the challenges, this study introduces fuzzy learning into deep learning models to overcome the drawbacks of fixed model representations. Therefore, we propose a fuzzy dynamic transformer (FDformer) that can adaptively adjust network depth according to the complexity of individual samples. Subsequently, we design a fuzzy feature extraction mechanism to capture feature information within the fuzzy membership degree, enabling the feature-level fusion of the fuzzy representation with the dynamic depth representation. Finally, we propose a training method for dynamically allocating loss weights, emphasizing the contribution of various samples to different exits, thereby improving the performance of time-series dynamic networks. Experiments on multiple datasets indicate that FDformer achieves minimal computational costs and excellent prediction accuracy across multiple datasets, outperforming SOTA algorithms. Lei Ren 0001, Tuo Zhao, Haiteng Wang |
IEEE Trans. Fuzzy Syst. | 2 |
| 2024 | RoseLoRA: Row and Column-wise Sparse Low-rank Adaptation of Pre-trained Language Model for Knowledge Editing and Fine-tuningabstractPre-trained language models, trained on largescale corpora, demonstrate strong generalizability across various NLP tasks.Finetuning these models for specific tasks typically involves updating all parameters, which is resource-intensive.Parameter-efficient finetuning (PEFT) methods, such as the popular LoRA family, introduce low-rank matrices to learn only a few parameters efficiently.However, during inference, the product of these matrices updates all pre-trained parameters, complicating tasks like knowledge editing that require selective updates.We propose a novel PEFT method, which conducts row and column-wise sparse low-rank adaptation (RoseLoRA), to address this challenge.RoseLoRA identifies and updates only the most important parameters for a specific task, maintaining efficiency while preserving other model knowledge.By adding a sparsity constraint on the product of low-rank matrices and converting it to row and column-wise sparsity, we ensure efficient and precise model updates.Our theoretical analysis guarantees the lower bound of the sparsity with respective to the matrix product.Extensive experiments on five benchmarks across twenty datasets demonstrate that RoseLoRA outperforms baselines in both general fine-tuning and knowledge editing tasks. Haoyu Wang 0004, Tianci Liu 0003, Ruirui Li 0002, Monica Xiao Cheng, Tuo Zhao, Jing Gao 0004 |
EMNLP | 5 |
| 2024 | BlendFilter: Advancing Retrieval-Augmented Large Language Models via Query Generation Blending and Knowledge FilteringabstractHaoyu Wang, Ruirui Li, Haoming Jiang, Jinjin Tian, Zhengyang Wang, Chen Luo, Xianfeng Tang, Monica Xiao Cheng, Tuo Zhao, Jing Gao. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Haoyu Wang 0004, Ruirui Li 0002, Haoming Jiang, Jinjin Tian, Chen Luo 0003, Xianfeng Tang, Monica Xiao Cheng, Tuo Zhao, Jing Gao 0004 |
EMNLP | 9 |
| 2024 | LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language ModelsabstractQuantization is an indispensable technique for serving Large Language Models (LLMs) and has recently found its way into LoRA fine-tuning (Dettmers et al., 2023). In this work we focus on the scenario where quantization and LoRA fine- tuning are applied together on a pre-trained model. In such cases it is common to observe a consistent gap in the performance on downstream tasks between full fine-tuning and quantization plus LoRA fine-tuning approach. In response, we propose LoftQ (LoRA-Fine-Tuning-aware Quantization), a novel quantization framework that simultaneously quantizes an LLM and finds a proper low-rank initialization for LoRA fine-tuning. Such an initialization alleviates the discrep- ancy between the quantized and full-precision model and significantly improves the generalization in downstream tasks. We evaluate our method on natural lan- guage understanding, question answering, summarization, and natural language generation tasks. Experiments show that our method is highly effective and out- performs existing quantization methods, especially in the challenging 2-bit and 2/4-bit mixed precision regimes. We will release our code. Yifan Yu 0008, Chen Liang 0006, Nikos Karampatziakis, Weizhu Chen, Tuo Zhao |
ICLR | 7 |
| 2024 | Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMsabstractIn human-written articles, we often leverage the subtleties of text style, such as bold and italics, to guide the attention of readers. These textual emphases are vital for the readers to grasp the conveyed information. When interacting with large language models (LLMs), we have a similar need -- steering the model to pay closer attention to user-specified information, e.g., an instruction. Existing methods, however, are constrained to process plain text and do not support such a mechanism. This motivates us to introduce PASTA -- Post-hoc Attention STeering Approach, a method that allows LLMs to read text with user-specified emphasis marks. To this end, PASTA identifies a small subset of attention heads and applies precise attention reweighting on them, directing the model attention to user-specified parts. Like prompting, PASTA is applied at inference time and does not require changing any model parameters. Experiments demonstrate that PASTA can substantially enhance an LLM's ability to follow user instructions or integrate new knowledge from user inputs, leading to a significant performance improvement on a variety of tasks, e.g., an average accuracy improvement of 22\% for LLAMA-7B. Our code is publicly available at https://github.com/QingruZhang/PASTA . Qingru Zhang, Chandan Singh, Xiaodong Liu 0003, Bin Yu 0001, Jianfeng Gao 0001, Tuo Zhao |
ICLR | 7 |
| 2024 | Beyond Point Prediction: Score Matching-based Pseudolikelihood Estimation of Neural Marked Spatio-Temporal Point ProcessabstractSpatio-temporal point processes (STPPs) are potent mathematical tools for modeling and predicting events with both temporal and spatial features. Despite their versatility, most existing methods for learning STPPs either assume a restricted form of the spatio-temporal distribution, or suffer from inaccurate approximations of the intractable integral in the likelihood training objective. These issues typically arise from the normalization term of the probability density function. Moreover, existing works only provide point prediction for events without quantifying their uncertainty, such as confidence intervals for the event’s arrival time and confidence regions for the event’s location, which is crucial given the considerable randomness of the data. To tackle these challenges, we introduce SMASH: a Score MAtching-based pSeudolikeliHood estimator for learning marked STPPs. Specifically, our framework adopts a normalization-free objective by estimating the pseudolikelihood of marked STPPs through score-matching and predicts confidence intervals/regions for event time and location by generating samples through a score-based sampling algorithm. The superior performance of our proposed framework is demonstrated through extensive experiments on both point and confidence interval/region prediction of events. Zichong Li, Qunzhi Xu, Zhenghao Xu, Yajun Mei, Tuo Zhao, Hongyuan Zha |
ICML | 5 |
| 2024 | To Cool or not to Cool? Temperature Network Meets Large Foundation Models via DROabstractThe temperature parameter plays a profound role during training and/or inference with large foundation models (LFMs) such as large language models (LLMs) and CLIP models. Particularly, it adjusts the logits in the softmax function in LLMs, which is crucial for next token generation, and it scales the similarities in the contrastive loss for training CLIP models. A significant question remains: “ Is it viable to learn a neural network to predict a personalized temperature of any input data for enhancing LFMs?" In this paper, we present a principled framework for learning a small yet generalizable temperature prediction network (TempNet) to improve LFMs. Our solution is composed of a novel learning framework with robust losses underpinned by constrained distributionally robust optimization (DRO), and a properly designed TempNet with theoretical inspiration. TempNet can be trained together with a large foundation model from scratch or learned separately given a pretrained foundation model. It is not only useful for predicting personalized temperature to promote the training of LFMs but also generalizable and transferable to new tasks. Our experiments on LLMs and CLIP models demonstrate that TempNet greatly improves the performance of existing solutions or models. Zi-Hao Qiu, Siqi Guo 0003, Mao Xu, Tuo Zhao, Lijun Zhang 0005, Tianbao Yang |
ICML | 4 |
| 2024 | A Cloud-Edge Intelligent Collaborative Framework and Its Applications in AIGC and Digital TwinsabstractWith the development of modern information technology and 5G, cloud-edge intelligent collaboration can make full use of the powerful computing power of cloud computing and the real-time response capability of edge computing to improve the overall efficiency of industrial intelligent systems. Thus, it shows strong application potential in digital twins, foundation models, and meta-universes. However, most of the existing researches focus on resource collaboration, data collaboration and application collaboration in cloud-edge computing framework, and there are many shortcomings in the research of intelligent collaboration framework. Therefore, we first propose a cloud-edge intelligent collaboration framework, which consists of three parts: terminal layer, edge artificial intelligence (AI) layer and cloud AI layer, including two core processes: cloud-edge intelligent training and cloud-edge intelligent inference. Then, we systematically analyze the key technologies of cloud-edge intelligent collaboration, including model segmentation, model early exit and foundation models. Finally, we put forward the application of cloud-edge intelligent collaboration in digital twin and AI Generated Content (AIGC), and through the analysis of typical application scenarios. Haiteng Wang, Lu Jiao, Tuo Zhao, Lei Ren 0001 |
IECON | 3 |
| 2024 | Robust Reinforcement Learning from Corrupted Human FeedbackabstractReinforcement learning from human feedback (RLHF) provides a principled framework for aligning AI systems with human preference data. For various reasons, e.g., personal bias, context ambiguity, lack of training, etc, human annotators may give incorrect or inconsistent preference labels. To tackle this challenge, we propose a robust RLHF approach -- $R^3M$, which models the potentially corrupted preference label as sparse outliers. Accordingly, we formulate the robust reward learning as an $\ell_1$-regularized maximum likelihood estimation problem. Computationally, we develop an efficient alternating optimization algorithm, which only incurs negligible computational overhead compared with the standard RLHF approach. Theoretically, we prove that under proper regularity conditions, $R^3M$ can consistently learn the underlying reward and identify outliers, provided that the number of outlier labels scales sublinearly with the preference sample size. Furthermore, we remark that $R^3M$ is versatile and can be extended to various preference optimization methods, including direct preference optimization (DPO). Our experiments on robotic control and natural language generation with large language models (LLMs) show that $R^3M$ improves robustness of the reward against several types of perturbations to the preference data. Alexander Bukharin, Ilgee Hong, Haoming Jiang, Zichong Li, Qingru Zhang, Tuo Zhao |
NeurIPS | 7 |
| 2024 | Adaptive Preference Scaling for Reinforcement Learning with Human FeedbackabstractReinforcement learning from human feedback (RLHF) is a prevalent approach to align AI systems with human values by learning rewards from human preference data. Due to various reasons, however, such data typically takes the form of rankings over pairs of trajectory segments, which fails to capture the varying strengths of preferences across different pairs. In this paper, we propose a novel adaptive preference loss, underpinned by distributionally robust optimization (DRO), designed to address this uncertainty in preference strength. By incorporating an adaptive scaling parameter into the loss for each pair, our method increases the flexibility of the reward function. Specifically, it assigns small scaling parameters to pairs with ambiguous preferences, leading to more comparable rewards, and large scaling parameters to those with clear preferences for more distinct rewards. Computationally, our proposed loss function is strictly convex and univariate with respect to each scaling parameter, enabling its efficient optimization through a simple second-order algorithm. Our method is versatile and can be readily adapted to various preference optimization frameworks, including direct preference optimization (DPO). Our experiments with robotic control and natural language generation with large language models (LLMs) show that our method not only improves policy performance but also aligns reward function selection more closely with policy optimization, simplifying the hyperparameter tuning process. Ilgee Hong, Zichong Li, Alexander Bukharin, Haoming Jiang, Tianbao Yang, Tuo Zhao |
NeurIPS | 7 |
| 2024 | Provable Acceleration of Nesterov's Accelerated Gradient for Asymmetric Matrix Factorization and Linear Neural NetworksabstractWe study the convergence rate of first-order methods for rectangular matrix factorization, which is a canonical nonconvex optimization problem. Specifically, given a rank-$r$ matrix $\mathbf{A}\in\mathbb{R}^{m\times n}$, we prove that gradient descent (GD) can find a pair of $\epsilon$-optimal solutions $\mathbf{X}_T\in\mathbb{R}^{m\times d}$ and $\mathbf{Y}_T\in\mathbb{R}^{n\times d}$, where $d\geq r$, satisfying $\lVert\mathbf{X}_T\mathbf{Y}_T^\top-\mathbf{A}\rVert_F\leq\epsilon\lVert\mathbf{A}\rVert_F$ in $T=O(\kappa^2\log\frac{1}{\epsilon})$ iterations with high probability, where $\kappa$ denotes the condition number of $\mathbf{A}$. Furthermore, we prove that Nesterov's accelerated gradient (NAG) attains an iteration complexity of $O(\kappa\log\frac{1}{\epsilon})$, which is the best-known bound of first-order methods for rectangular matrix factorization. Different from small balanced random initialization in the existing literature, we adopt an unbalanced initialization, where $\mathbf{X}_0$ is large and $\mathbf{Y}_0$ is $0$. Moreover, our initialization and analysis can be further extended to linear neural networks, where we prove that NAG can also attain an accelerated linear convergence rate. In particular, we only require the width of the network to be greater than or equal to the rank of the output label matrix. In contrast, previous results achieving the same rate require excessive widths that additionally depend on the condition number and the rank of the input data matrix. Zhenghao Xu, Yuqing Wang 0005, Tuo Zhao, Rachel Ward, Molei Tao |
NeurIPS | 3 |
| 2024 | Nonparametric Classification on Low Dimensional Manifolds using Overparameterized Convolutional Residual NetworksabstractConvolutional residual neural networks (ConvResNets), though overparametersized, can achieve remarkable prediction performance in practice, which cannot be well explained by conventional wisdom. To bridge this gap, we study the performance of ConvResNeXts trained with weight decay, which cover ConvResNets as a special case, from the perspective of nonparametric classification. Our analysis allows for infinitely many building blocks in ConvResNeXts, and shows that weight decay implicitly enforces sparsity on these blocks. Specifically, we consider a smooth target function supported on a low-dimensional manifold, then prove that ConvResNeXts can adapt to the function smoothness and low-dimensional structures and efficiently learn the function without suffering from the curse of dimensionality. Our findings partially justify the advantage of overparameterized ConvResNeXts over conventional machine learning models. Kaiqi Zhang 0002, Minshuo Chen, Yuma Takeda, Mengdi Wang 0001, Tuo Zhao, Yu-Xiang Wang 0003 |
NeurIPS | 6 |
| 2024 | Deep Nonparametric Estimation of Operators between Infinite Dimensional SpacesabstractLearning operators between infinitely dimensional spaces is an important learning task arising in machine learning, imaging science, mathematical modeling and simulations, etc. This paper studies the nonparametric estimation of Lipschitz operators using deep neural networks. Non-asymptotic upper bounds are derived for the generalization error of the empirical risk minimizer over a properly chosen network class. Under the assumption that the target operator exhibits a low dimensional structure, our error bounds decay as the training sample size increases, with an attractive fast rate depending on the intrinsic dimension in our estimation. Our assumptions cover most scenarios in real applications and our results give rise to fast rates by exploiting low dimensional structures of data in operator estimation. We also investigate the influence of network structures (e.g., network width, depth, and sparsity) on the generalization error of the neural network estimator and propose a general suggestion on the choice of network structures to maximize the learning efficiency quantitatively. Hao Liu 0028, Haizhao Yang, Minshuo Chen, Tuo Zhao, Wenjing Liao |
J. Mach. Learn. Res. | 4 |
| 2024 | Sample Complexity of Neural Policy Mirror Descent for Policy Optimization on Low-Dimensional ManifoldsabstractPolicy gradient methods equipped with deep neural networks have achieved great success in solving high-dimensional reinforcement learning (RL) problems. However, current analyses cannot explain why they are resistant to the curse of dimensionality. In this work, we study the sample complexity of the neural policy mirror descent (NPMD) algorithm with deep convolutional neural networks (CNN). Motivated by the empirical observation that many high-dimensional environments have state spaces possessing low-dimensional structures, such as those taking images as states, we consider the state space to be a $d$-dimensional manifold embedded in the $D$-dimensional Euclidean space with intrinsic dimension $d\ll D$. We show that in each iteration of NPMD, both the value function and the policy can be well approximated by CNNs. The approximation errors are controlled by the size of the networks, and the smoothness of the previous networks can be inherited. As a result, by properly choosing the network size and hyperparameters, NPMD can find an $\epsilon$-optimal policy with $\tilde{O}(\epsilon^{-\frac{d}{\alpha}-2})$ samples in expectation, where $\alpha\in(0,1]$ indicates the smoothness of environment. Compared to previous work, our result exhibits that NPMD can leverage the low-dimensional structure of state space to escape from the curse of dimensionality, explaining the efficacy of deep policy gradient algorithms. Zhenghao Xu, Minshuo Chen, Mengdi Wang 0001, Tuo Zhao |
J. Mach. Learn. Res. | 5 |
| 2024 | Learning explainable task-relevant state representation for model-free deep reinforcement learning
Tingting Zhao 0001, Guixi Li, Tuo Zhao, Yarui Chen, Ning Xie 0003, Gang Niu 0001, Masashi Sugiyama |
Neural Networks | 3 |
| 2023 | Reinforcement Learning for Adaptive Mesh RefinementabstractFinite element simulations of physical systems governed by partial differential equations (PDE) crucially depend on adaptive mesh refinement (AMR) to allocate computational budget to regions where higher resolution is required. Existing scalable AMR methods make heuristic refinement decisions based on instantaneous error estimation and thus do not aim for long-term optimality over an entire simulation. We propose a novel formulation of AMR as a Markov decision process and apply deep reinforcement learning (RL) to train refinement policies directly from simulation. AMR poses a challenge for RL as both the state dimension and available action set changes at every step, which we solve by proposing new policy architectures with differing generality and inductive bias. The model sizes of these policy architectures are independent of the mesh size and hence can be deployed on larger simulations than those used at training time. We demonstrate in comprehensive experiments on static function estimation and time-dependent equations that RL policies can be trained on problems without using ground truth solutions, are competitive with a widely-used error estimator, and generalize to larger and unseen test problems. Tarik Dzanic, Brenden K. Petersen, Jun Kudo, Ketan Mittal, Vladimir Z. Tomov, Sylvain Camier, Tuo Zhao, Hongyuan Zha, Tzanio V. Kolev, Robert W. Anderson, Daniel M. Faissol |
AISTATS | 8 |
| 2023 | Joint Estimation of DOA and Distance in Noisy Reverberant ConditionsabstractSound Source localization (SSL) using microphone arrays is an active research topic with many applications, but noise and reverberation make the direction-of-arrival (DOA) and distance estimation a challenging problem. In this work, we propose a novel method to jointly estimate the DOA and distance in noisy and reverberant environments. Our method exploits the linear phase structure across frequencies in a steering vector (SV). We convert the joint estimation issue into an optimization problem, which can be solved by Newton’s method augmented by a gradient ascent method. Our method does not depend on certain microphone array geometry, and it can also be extended to estimate the elevation angle. We conducted experimental evaluations in simulated noisy and reverberant acoustic conditions, which verified the superiority of our proposed method to several established methods in estimation accuracy and computation efficiency. Suliang Bu, Tuo Zhao, Yunxin Zhao |
ICASSP | 2 |
| 2023 | Sample Complexity of Nonparametric Off-Policy Evaluation on Low-Dimensional Manifolds using Deep Networks
Minshuo Chen, Mengdi Wang 0001, Tuo Zhao |
ICLR | 4 |
| 2023 | HomoDistil: Homotopic Task-Agnostic Distillation of Pre-trained Transformers
Chen Liang 0006, Haoming Jiang, Zheng Li 0018, Xianfeng Tang, Tuo Zhao |
ICLR | 6 |
| 2023 | Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Qingru Zhang, Minshuo Chen, Alexander Bukharin, Yu Cheng 0001, Weizhu Chen, Tuo Zhao |
ICLR | 7 |
| 2023 | Machine Learning Force Fields with Data Cost Aware TrainingabstractMachine learning force fields (MLFF) have been proposed to accelerate molecular dynamics (MD) simulation, which finds widespread applications in chemistry and biomedical research. Even for the most data-efficient MLFFs, reaching chemical accuracy can require hundreds of frames of force and energy labels generated by expensive quantum mechanical algorithms, which may scale as $O(n^3)$ to $O(n^7)$, with $n$ proportional to the number of basis functions. To address this issue, we propose a multi-stage computational framework -- ASTEROID, which lowers the data cost of MLFFs by leveraging a combination of cheap inaccurate data and expensive accurate data. The motivation behind ASTEROID is that inaccurate data, though incurring large bias, can help capture the sophisticated structures of the underlying force field. Therefore, we first train a MLFF model on a large amount of inaccurate training data, employing a bias-aware loss function to prevent the model from overfitting the potential bias of this data. We then fine-tune the obtained model using a small amount of accurate training data, which preserves the knowledge learned from the inaccurate training data while significantly improving the model's accuracy. Moreover, we propose a variant of ASTEROID based on score matching for the setting where the inaccurate training data are unlabeled. Extensive experiments on MD datasets and downstream tasks validate the efficacy of ASTEROID. Our code and data are available at https://github.com/abukharin3/asteroid. Alexander Bukharin, Shengjie Wang 0001, Simiao Zuo, Weihao Gao, Tuo Zhao |
ICML | 7 |
| 2023 | Score Approximation, Estimation and Distribution Recovery of Diffusion Models on Low-Dimensional DataabstractDiffusion models achieve state-of-the-art performance in various generation tasks. However, their theoretical foundations fall far behind. This paper studies score approximation, estimation, and distribution recovery of diffusion models, when data are supported on an unknown low-dimensional linear subspace. Our result provides sample complexity bounds for distribution estimation using diffusion models. We show that with a properly chosen neural network architecture, the score function can be both accurately approximated and efficiently estimated. Further, the generated distribution based on the estimated score function captures the data geometric structures and converges to a close vicinity of the data distribution. The convergence rate depends on subspace dimension, implying that diffusion models can circumvent the curse of data ambient dimensionality. Minshuo Chen, Kaixuan Huang, Tuo Zhao, Mengdi Wang 0001 |
ICML | 3 |
| 2023 | SMURF-THP: Score Matching-based UnceRtainty quantiFication for Transformer Hawkes ProcessabstractTransformer Hawkes process models have shown to be successful in modeling event sequence data. However, most of the existing training methods rely on maximizing the likelihood of event sequences, which involves calculating some intractable integral. Moreover, the existing methods fail to provide uncertainty quantification for model predictions, e.g., confidence interval for the predicted event’s arrival time. To address these issues, we propose SMURF-THP, a score-based method for learning Transformer Hawkes process and quantifying prediction uncertainty. Specifically, SMURF-THP learns the score function of the event’s arrival time based on a score-matching objective that avoids the intractable computation. With such a learnt score function, we can sample arrival time of events from the predictive distribution. This naturally allows for the quantification of uncertainty by computing confidence intervals over the generated samples. We conduct extensive experiments in both event type prediction and uncertainty quantification on time of arrival. In all the experiments, SMURF-THP outperforms existing likelihood-based methods in confidence calibration while exhibiting comparable prediction accuracy. Zichong Li, Yanbo Xu, Simiao Zuo, Haoming Jiang, Chao Zhang 0014, Tuo Zhao, Hongyuan Zha |
ICML | 6 |
| 2023 | LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse ApproximationabstractTransformer models have achieved remarkable results in various natural language tasks, but they are often prohibitively large, requiring massive memories and computational resources. To re- duce the size and complexity of these models, we propose LoSparse (Low-Rank and Sparse ap- proximation), a novel model compression tech- nique that approximates a weight matrix by the sum of a low-rank matrix and a sparse matrix. Our method combines the advantages of both low- rank approximations and pruning, while avoid- ing their limitations. Low-rank approximation compresses the coherent and expressive parts in neurons, while pruning removes the incoherent and non-expressive parts in neurons. Pruning enhances the diversity of low-rank approxima- tions, and low-rank approximation prevents prun- ing from losing too many expressive neurons. We evaluate our method on natural language under- standing, question answering, and natural lan- guage generation tasks. We show that it signif- icantly outperforms existing compression meth- ods. Our code is publicly available at https: //github.com/yxli2123/LoSparse Yifan Yu 0008, Qingru Zhang, Chen Liang 0006, Weizhu Chen, Tuo Zhao |
ICML | 7 |
| 2023 | Less is More: Task-aware Layer-wise Distillation for Language Model CompressionabstractLayer-wise distillation is a powerful tool to compress large models (i.e. teacher models) into small ones (i.e., student models). The student distills knowledge from the teacher by mimicking the hidden representations of the teacher at every intermediate layer. However, layer-wise distillation is difficult. Since the student has a smaller model capacity than the teacher, it is often under-fitted. Furthermore, the hidden representations of the teacher contain redundant information that the student does not necessarily need for the target task's learning. To address these challenges, we propose a novel Task-aware layEr-wise Distillation (TED). TED designs task-aware filters to align the hidden representations of the student and the teacher at each layer. The filters select the knowledge that is useful for the target task from the hidden representations. As such, TED reduces the knowledge gap between the two models and helps the student to fit better on the target task. We evaluate TED in two scenarios: continual pre-training and fine-tuning. TED demonstrates significant and consistent improvements over existing distillation methods in both scenarios. Code is available at https://github.com/cliang1453/task-aware-distillation. Chen Liang 0006, Simiao Zuo, Qingru Zhang, Weizhu Chen, Tuo Zhao |
ICML | 6 |
| 2023 | Effective Minkowski Dimension of Deep Nonparametric Regression: Function Approximation and Statistical TheoriesabstractExisting theories on deep nonparametric regression have shown that when the input data lie on a low-dimensional manifold, deep neural networks can adapt to the intrinsic data structures. In real world applications, such an assumption of data lying exactly on a low dimensional manifold is stringent. This paper introduces a relaxed assumption that the input data are concentrated around a subset of $\mathbb{R}^d$ denoted by $\mathcal{S}$, and the intrinsic dimension of $\mathcal{S}$ can be characterized by a new complexity notation – effective Minkowski dimension. We prove that, the sample complexity of deep nonparametric regression only depends on the effective Minkowski dimension of $\mathcal{S}$ denoted by $p$. We further illustrate our theoretical findings by considering nonparametric regression with an anisotropic Gaussian random design $N(0,\Sigma)$, where $\Sigma$ is full rank. When the eigenvalues of $\Sigma$ have an exponential or polynomial decay, the effective Minkowski dimension of such an Gaussian random design is $p=\mathcal{O}(\sqrt{\log n})$ or $p=\mathcal{O}(n^\gamma)$, respectively, where $n$ is the sample size and $\gamma\in(0,1)$ is a small constant depending on the polynomial decay rate. Our theory shows that, when the manifold assumption does not hold, deep neural networks can still adapt to the effective Minkowski dimension of the data, and circumvent the curse of the ambient dimensionality for moderate sample sizes. Minshuo Chen, Mengdi Wang 0001, Wenjing Liao, Tuo Zhao |
ICML | 5 |
| 2023 | LightToken: A Task and Model-agnostic Lightweight Token Embedding Framework for Pre-trained Language Models
Haoyu Wang 0004, Ruirui Li 0002, Haoming Jiang, Xianfeng Tang, Bin Bi, Monica Xiao Cheng, Yaqing Wang 0001, Tuo Zhao, Jing Gao 0004 |
KDD | 10 |
| 2023 | Module-wise Adaptive Distillation for Multimodality Foundation ModelsabstractPre-trained multimodal foundation models have demonstrated remarkable generalizability but pose challenges for deployment due to their large sizes. One effective approach to reducing their sizes is layerwise distillation, wherein small student models are trained to match the hidden representations of large teacher models at each layer. Motivated by our observation that certain architecture components, referred to as modules, contribute more significantly to the student's performance than others, we propose to track the contributions of individual modules by recording the loss decrement after distillation each module and choose the module with a greater contribution to distill more frequently. Such an approach can be naturally formulated as a multi-armed bandit (MAB) problem, where modules and loss decrements are considered as arms and rewards, respectively. We then develop a modified-Thompson sampling algorithm named OPTIMA to address the nonstationarity of module contributions resulting from model updating. Specifically, we leverage the observed contributions in recent history to estimate the changing contribution of each module and select modules based on these estimations to maximize the cumulative contribution. We evaluate the effectiveness of OPTIMA through distillation experiments on various multimodal understanding and image captioning tasks, using the CoCa-Large model \citep{yu2022coca} as the teacher model. Chen Liang 0006, Ming-Hsuan Yang 0001, Matthew Brown 0001, Yin Cui, Tuo Zhao, Boqing Gong, Tianyi Zhou 0002 |
NeurIPS | 6 |
| 2023 | Robust Multi-Agent Reinforcement Learning via Adversarial Regularization: Theoretical Foundation and Stable AlgorithmsabstractMulti-Agent Reinforcement Learning (MARL) has shown promising results across several domains. Despite this promise, MARL policies often lack robustness and are therefore sensitive to small changes in their environment. This presents a serious concern for the real world deployment of MARL algorithms, where the testing environment may slightly differ from the training environment. In this work we show that we can gain robustness by controlling a policy’s Lipschitz constant, and under mild conditions, establish the existence of a Lipschitz and close-to-optimal policy. Motivated by these insights, we propose a new robust MARL framework, ERNIE, that promotes the Lipschitz continuity of the policies with respect to the state observations and actions by adversarial regularization. The ERNIE framework provides robustness against noisy observations, changing transition dynamics, and malicious actions of agents. However, ERNIE’s adversarial regularization may introduce some training instability. To reduce this instability, we reformulate adversarial regularization as a Stackelberg game. We demonstrate the effectiveness of the proposed framework with extensive experiments in traffic light control and particle environments. In addition, we extend ERNIE to mean-field MARL with a formulation based on distributionally robust optimization that outperforms its non-robust counterpart and is of independent interest. Our code is available at https://github.com/abukharin3/ERNIE. Alexander Bukharin, Yue Yu 0001, Qingru Zhang, Zhehui Chen, Simiao Zuo, Chao Zhang 0014, Songan Zhang, Tuo Zhao |
NeurIPS | 9 |
| 2023 | Model-Based Reparameterization Policy Gradient Methods: Theory and Practical AlgorithmsabstractReParameterization (RP) Policy Gradient Methods (PGMs) have been widely adopted for continuous control tasks in robotics and computer graphics. However, recent studies have revealed that, when applied to long-term reinforcement learning problems, model-based RP PGMs may experience chaotic and non-smooth optimization landscapes with exploding gradient variance, which leads to slow convergence. This is in contrast to the conventional belief that reparameterization methods have low gradient estimation variance in problems such as training deep generative models. To comprehend this phenomenon, we conduct a theoretical examination of model-based RP PGMs and search for solutions to the optimization difficulties. Specifically, we analyze the convergence of the model-based RP PGMs and pinpoint the smoothness of function approximators as a major factor that affects the quality of gradient estimation. Based on our analysis, we propose a spectral normalization method to mitigate the exploding variance issue caused by long model unrolls. Our experimental results demonstrate that proper normalization significantly reduces the gradient variance of model-based RP PGMs. As a result, the performance of the proposed method is comparable or superior to other gradient estimators, such as the Likelihood Ratio (LR) gradient estimator. Our code is available at https://github.com/agentification/RP_PGM. Shenao Zhang, Boyi Liu 0001, Zhaoran Wang 0001, Tuo Zhao |
NeurIPS | 4 |
| 2023 | Pivotal Estimation of Linear Discriminant Analysis in High DimensionsabstractWe consider the linear discriminant analysis problem in the high-dimensional settings. In this work, we propose PANDA(PivotAl liNear Discriminant Analysis), a tuning insensitive method in the sense that it requires very little effort to tune the parameters. Moreover, we prove that PANDA achieves the optimal convergence rate in terms of both the estimation error and misclassification rate. Our theoretical results are backed up by thorough numerical studies using both simulated and real datasets. In comparison with the existing methods, we observe that our proposed PANDA yields equal or better performance, and requires substantially less effort in parameter tuning. Ethan X. Fang, Yajun Mei, Qunzhi Xu, Tuo Zhao |
J. Mach. Learn. Res. | 5 |
| 2022 | CAMERO: Consistency Regularized Ensemble of Perturbed Language Models with Weight SharingabstractModel ensemble is a popular approach to produce a low-variance and well-generalized model.However, it induces large memory and inference costs, which are often not affordable for real-world deployment.Existing work has resorted to sharing weights among models.However, when increasing the proportion of the shared weights, the resulting models tend to be similar, and the benefits of using model ensemble diminish.To retain ensemble benefits while maintaining a low memory cost, we propose a consistency-regularized ensemble learning approach based on perturbed models, named CAMERO.Specifically, we share the weights of bottom layers across all models and apply different perturbations to the hidden representations for different models, which can effectively promote the model diversity.Meanwhile, we apply a prediction consistency regularizer across the perturbed models to control the variance due to the model diversity.Our experiments using large language models demonstrate that CAMERO significantly improves the generalization performance of the ensemble model.Specifically, CAMERO outperforms the standard ensemble of 8 BERT-base models on the GLUE benchmark by 0.7 with a significantly smaller model size (114.2Mvs. 880.6M).* Work was done during an internship at Microsoft Azure AI. Chen Liang 0006, Yelong Shen, Weizhu Chen, Tuo Zhao |
ACL (1) | 5 |
| 2022 | Noise Regularizes Over-parameterized Rank One Matrix Recovery, ProvablyabstractWe investigate the role of noise in optimization algorithms for learning over-parameterized models. Specifically, we consider the recovery of a rank one matrix $Y^*\in R^{d\times d}$ from a noisy observation $Y$ using an over-parameterization model. Specifically, we parameterize the rank one matrix $Y^*$ by $XX^\top$, where $X\in R^{d\times d}$. We then show that under mild conditions, the estimator, obtained by the randomly perturbed gradient descent algorithm using the square loss function, attains a mean square error of $O(\sigma^2/d)$, where $\sigma^2$ is the variance of the observational noise. In contrast, the estimator obtained by gradient descent without random perturbation only attains a mean square error of $O(\sigma^2)$. Our result partially justifies the implicit regularization effect of noise when learning over-parameterized models, and provides new understanding of training over-parameterized neural networks. Yan Li 0074, Enlu Zhou, Tuo Zhao |
AISTATS | 4 |
| 2022 | Frequency-aware SGD for Efficient Embedding Learning with Provable Benefits
Yan Li 0074, Dhruv Choudhary, Xiaohan Wei, Baichuan Yuan, Bhargav Bhushanam, Tuo Zhao, Guanghui Lan |
ICLR | 6 |
| 2022 | No Parameters Left Behind: Sensitivity Guided Adaptive Learning Rate for Training Large Transformer Models
Chen Liang 0006, Haoming Jiang, Simiao Zuo, Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen, Tuo Zhao |
ICLR | 8 |
| 2022 | Large Learning Rate Tames Homogeneity: Convergence and Balancing Effect
Yuqing Wang 0005, Minshuo Chen, Tuo Zhao, Molei Tao |
ICLR | 3 |
| 2022 | Taming Sparsely Activated Transformer with Stochastic Experts
Simiao Zuo, Xiaodong Liu 0003, Jian Jiao 0007, Young Jin Kim 0006, Hany Hassan, Ruofei Zhang, Jianfeng Gao 0001, Tuo Zhao |
ICLR | 8 |
| 2022 | Benefits of Overparameterized Convolutional Residual Networks: Function Approximation under Smoothness ConstraintabstractOverparameterized neural networks enjoy great representation power on complex data, and more importantly yield sufficiently smooth output, which is crucial to their generalization and robustness. Most existing function approximation theories suggest that with sufficiently many parameters, neural networks can well approximate certain classes of functions in terms of the function value. The neural network themselves, however, can be highly nonsmooth. To bridge this gap, we take convolutional residual networks (ConvResNets) as an example, and prove that large ConvResNets can not only approximate a target function in terms of function value, but also exhibit sufficient first-order smoothness. Moreover, we extend our theory to approximating functions supported on a low-dimensional manifold. Our theory partially justifies the benefits of using deep and wide networks in practice. Numerical experiments on adversarial robust image classification are provided to support our theory. Hao Liu 0028, Minshuo Chen, Siawpeng Er, Wenjing Liao, Tong Zhang 0001, Tuo Zhao |
ICML | 6 |
| 2022 | PLATON: Pruning Large Transformer Models with Upper Confidence Bound of Weight ImportanceabstractLarge Transformer-based models have exhibited superior performance in various natural language processing and computer vision tasks. However, these models contain enormous amounts of parameters, which restrict their deployment to real-world applications. To reduce the model size, researchers prune these models based on the weights’ importance scores. However, such scores are usually estimated on mini-batches during training, which incurs large variability/uncertainty due to mini-batch sampling and complicated training dynamics. As a result, some crucial weights could be pruned by commonly used pruning methods because of such uncertainty, which makes training unstable and hurts generalization. To resolve this issue, we propose PLATON, which captures the uncertainty of importance scores by upper confidence bound of importance estimation. In particular, for the weights with low importance scores but high uncertainty, PLATON tends to retain them and explores their capacity. We conduct extensive experiments with several Transformer-based models on natural language understanding, question answering and image classification to validate the effectiveness of PLATON. Results demonstrate that PLATON manifests notable improvement under different sparsity levels. Our code is publicly available at https://github.com/QingruZhang/PLATON. Qingru Zhang, Simiao Zuo, Chen Liang 0006, Alexander Bukharin, Weizhu Chen, Tuo Zhao |
ICML | 7 |
| 2022 | Steering vector correction in MVDR beamformer for speech enhancement
Suliang Bu, Yunxin Zhao, Tuo Zhao |
INTERSPEECH | 3 |
| 2022 | CERES: Pretraining of Graph-Conditioned Transformer for Semi-Structured Session DataabstractRui Feng, Chen Luo, Qingyu Yin, Bing Yin, Tuo Zhao, Chao Zhang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Chen Luo 0003, Qingyu Yin, Tuo Zhao, Chao Zhang 0014 |
NAACL-HLT | 5 |
| 2022 | MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided AdaptationabstractSimiao Zuo, Qingru Zhang, Chen Liang, Pengcheng He, Tuo Zhao, Weizhu Chen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Simiao Zuo, Qingru Zhang, Chen Liang 0006, Tuo Zhao, Weizhu Chen |
NAACL-HLT | 5 |
| 2022 | On Deep Generative Models for Approximation and Estimation of Distributions on ManifoldsabstractDeep generative models have experienced great empirical successes in distribution learning. Many existing experiments have demonstrated that deep generative networks can efficiently generate high-dimensional complex data from a low-dimensional easy-to-sample distribution. However, this phenomenon can not be justified by existing theories. The widely held manifold hypothesis speculates that real-world data sets, such as natural images and signals, exhibit low-dimensional geometric structures. In this paper, we take such low-dimensional data structures into consideration by assuming that data distributions are supported on a low-dimensional manifold. We prove approximation and estimation theories of deep generative networks for estimating distributions on a low-dimensional manifold under the Wasserstein-1 loss. We show that the Wasserstein-1 loss converges to zero at a fast rate depending on the intrinsic dimension instead of the ambient data dimension. Our theory leverages the low-dimensional geometric structures in data sets and justifies the practical power of deep generative models. We require no smoothness assumptions on the data distribution which is desirable in practice. Biraj Dahal, Alexander Havrilla, Minshuo Chen, Tuo Zhao, Wenjing Liao |
NeurIPS | 4 |
| 2022 | TDOA Estimation of Speech Source in Noisy Reverberant EnvironmentsabstractSound source localization is important in many applications, but noise and reverberation make the time difference of arrival (TDOA) estimation a challenging problem. In this work, we propose two novel methods to effectively estimate TDOA in noisy and reverberant environments. Our methods exploit the linear phase structure across frequencies in a steering vector (SV). To reduce potential noise and mathematical issues, we utilize absolute phases of SVs. We convert TDOA estimation into an optimization problem that is solvable by Newton's method. We conducted experimental evaluations in simulated acoustic conditions. In conditions with moderate-to-high input SNR and low reverberation, our fast-search method is superior in TDOA accuracy and is very efficient in computation. In conditions with low input SNR and high reverberation, our detailed-search method shows strong robustness and maintains advantageous performance in TDOA estimation. Suliang Bu, Tuo Zhao, Yunxin Zhao |
SLT | 2 |
| 2022 | Understanding users' trust transfer mechanism in a blockchain-enabled platform: A mixed methods study
Lin Zhang 0044, Susan A. Brown, Tuo Zhao |
Decis. Support Syst. | 4 |
| 2022 | Modeling Speech Structure to Improve T-F Masks for Speech Enhancement and RecognitionabstractTime-frequency (TF) masks are widely used in speech enhancement (SE). However, accurately estimating TF masks from noisy speech remains a challenge to both statistical or neural network (NN) approaches. Statistical model based mask estimation usually depends on a good parameter initialization, while NN-based method relies on setting proper and stable learning targets. To address these issues, we propose to extract TF speech structure from clean speech and partition noisy speech spectrogram into mutually exclusive regions. We investigate modeling clean speech by utterance-specific narrowband complex Gaussian mixture models to derive the regions, and using the region targets to supervise the training of UNet++, a high-performance NN, for predicting regions from noisy speech. For multichannel SE, we consider two scenarios of using speech regions: 1) integrating the regions with TF masks by constraining the mask values or the model parameter updates, and 2) using the predicted regions in place of TF masks. For single-channel SE, we consider using the region targets to improve TF mask targets. Furthermore, we propose to use UNet++ for TF mask estimation. Our experiment results on speech recognition (CHiME-3) and SE (CHiME-3 and LibriSpeech) have demonstrated the effectiveness of our proposed approach of modeling speech region structure to improve TF masks for speech recognition and enhancement. Suliang Bu, Yunxin Zhao, Tuo Zhao |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Named Entity Recognition with Small Strongly Labeled and Large Weakly Labeled DataabstractHaoming Jiang, Danqing Zhang, Tianyu Cao, Bing Yin, Tuo Zhao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Haoming Jiang, Danqing Zhang, Tianyu Cao 0001, Tuo Zhao |
ACL/IJCNLP (1) | 5 |
| 2021 | Super Tickets in Pre-Trained Language Models: From Model Compression to Improving GeneralizationabstractChen Liang, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu, Pengcheng He, Tuo Zhao, Weizhu Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Liang 0006, Simiao Zuo, Minshuo Chen, Haoming Jiang, Xiaodong Liu 0003, Tuo Zhao, Weizhu Chen |
ACL/IJCNLP (1) | 7 |
| 2021 | Learning to Defend by Learning to AttackabstractAdversarial training provides a principled approach for training robust neural networks. From an optimization perspective, adversarial training is essentially solving a bilevel optimization problem. The leader problem is trying to learn a robust classifier, while the follower maximization is trying to generate adversarial samples. Unfortunately, such a bilevel problem is difficult to solve due to its highly complicated structure. This work proposes a new adversarial training method based on a generic learning-to-learn (L2L) framework. Specifically, instead of applying existing hand-designed algorithms for the inner problem, we learn an optimizer, which is parametrized as a convolutional neural network. At the same time, a robust classifier is learned to defense the adversarial attack generated by the learned optimizer. Experiments over CIFAR-10 and CIFAR-100 datasets demonstrate that L2L outperforms existing adversarial training methods in both classification accuracy and computational efficiency. Moreover, our L2L framework can be extended to generative adversarial imitation learning and stabilize the training. Haoming Jiang, Zhehui Chen, Yuyang Shi 0001, Bo Dai 0001, Tuo Zhao |
AISTATS | 5 |
| 2021 | Noisy Gradient Descent Converges to Flat Minima for Nonconvex Matrix FactorizationabstractNumerous empirical evidences have corroborated the importance of noise in nonconvex optimization problems. The theory behind such empirical observations, however, is still largely unknown. This paper studies this fundamental problem through investigating the nonconvex rectangular matrix factorization problem, which has infinitely many global minima due to rotation and scaling invariance. Hence, gradient descent (GD) can converge to any optimum, depending on the initialization. In contrast, we show that a perturbed form of GD with an arbitrary initialization converges to a global optimum that is uniquely determined by the injected noise. Our result implies that the noise imposes implicit bias towards certain optima. Numerical experiments are provided to support our theory. Yan Li 0074, Song Wei, Enlu Zhou, Tuo Zhao |
AISTATS | 5 |
| 2021 | QUEACO: Borrowing Treasures from Weakly-labeled Behavior Data for Query Attribute Value ExtractionabstractWe study the problem of query attribute value extraction, which aims to identify named entities from user queries as diverse surface form attribute values and afterward transform them into formally canonical forms. Such a problem consists of two phases: named entity recognition (NER) and attribute value normalization (AVN). However, existing works only focus on the NER phase but neglect equally important AVN. To bridge this gap, this paper proposes a unified query attribute value extraction system in e-commerce search named QUEACO, which involves both two phases. Moreover, by leveraging large-scale weakly-labeled behavior data, we further improve the extraction performance with less supervision cost. Specifically, for the NER phase, QUEACO adopts a novel teacher-student network, where a teacher network that is trained on the strongly-labeled data generates pseudo-labels to refine the weakly-labeled data for training a student network. Meanwhile, the teacher network can be dynamically adapted by the feedback of the student's performance on strongly-labeled data to maximally denoise the noisy supervisions from the weak labels. For the AVN phase, we also leverage the weakly-labeled query-to-attribute behavior data to normalize surface form attribute values from queries into canonical forms from products. Extensive experiments on a real-world large-scale E-commerce dataset demonstrate the effectiveness of QUEACO. Danqing Zhang, Zheng Li 0018, Tianyu Cao 0001, Chen Luo 0003, Hanqing Lu, Yiwei Song, Tuo Zhao, Qiang Yang 0001 |
CIKM | 9 |
| 2021 | Towards Automatic Evaluation of Dialog Systems: A Model-Free Off-Policy Evaluation ApproachabstractReliable automatic evaluation of dialogue systems under an interactive environment has long been overdue.An ideal environment for evaluating dialog systems, also known as the Turing test, needs to involve human interaction, which is usually not affordable for large scale experiments.Though researchers have attempted to use metrics for language generation tasks (e.g., perplexity, BLEU) or some model-based reinforcement learning methods (e.g., self-play evaluation) for automatic evaluation, these methods only show very weak correlation with the actual human evaluation in practice.To bridge such a gap, we propose a new framework named ENIGMA for estimating human evaluation scores based on recent advances of off-policy evaluation in reinforcement learning.ENIGMA only requires a handful of pre-collected experience data, and therefore does not involve human interaction with the target policy during the evaluation, making automatic evaluations feasible.More importantly, ENIGMA is model-free and agnostic to the behavior policies for collecting the experience data (see details in Section 2), which significantly alleviates the technical difficulties of modeling complex dialogue environments and human behaviors.Our experiments show that ENIGMA significantly outperforms existing methods in terms of correlation with human evaluation scores. Haoming Jiang, Bo Dai 0001, Sherry Yang 0001, Tuo Zhao, Wei Wei 0019 |
EMNLP (1) | 4 |
| 2021 | Adversarial Regularization as Stackelberg Game: An Unrolled Optimization ApproachabstractAdversarial regularization has been shown to improve the generalization performance of deep learning models in various natural language processing tasks.Existing works usually formulate the method as a zero-sum game, which is solved by alternating gradient descent/ascent algorithms.Such a formulation treats the adversarial and the defending players equally, which is undesirable because only the defending player contributes to the generalization performance.To address this issue, we propose Stackelberg Adversarial Regularization (SALT), which formulates adversarial regularization as a Stackelberg game.This formulation induces a competition between a leader and a follower, where the follower generates perturbations, and the leader trains the model subject to the perturbations.Different from conventional approaches, in SALT, the leader is in an advantageous position.When the leader moves, it recognizes the strategy of the follower and takes the anticipated follower's outcomes into consideration.Such a leader's advantage enables us to improve the model fitting to the unperturbed data.The leader's strategic information is captured by the Stackelberg gradient, which is obtained using an unrolling algorithm.Our experimental results on a set of machine translation and natural language understanding tasks show that SALT outperforms existing adversarial regularization baselines across all tasks.Our code is publicly available. Simiao Zuo, Chen Liang 0006, Haoming Jiang, Xiaodong Liu 0003, Jianfeng Gao 0001, Weizhu Chen, Tuo Zhao |
EMNLP (1) | 8 |
| 2021 | A Hypergradient Approach to Robust Regression without Correspondence
Yujia Xie, Yixiu Mao, Simiao Zuo, Hongteng Xu, Xiaojing Ye, Tuo Zhao, Hongyuan Zha |
ICLR | 6 |
| 2021 | How Important is the Train-Validation Split in Meta-Learning?abstractMeta-learning aims to perform fast adaptation on a new task through learning a “prior” from multiple existing tasks. A common practice in meta-learning is to perform a train-validation split (\emph{train-val method}) where the prior adapts to the task on one split of the data, and the resulting predictor is evaluated on another split. Despite its prevalence, the importance of the train-validation split is not well understood either in theory or in practice, particularly in comparison to the more direct \emph{train-train method}, which uses all the per-task data for both training and evaluation. We provide a detailed theoretical study on whether and when the train-validation split is helpful in the linear centroid meta-learning problem. In the agnostic case, we show that the expected loss of the train-val method is minimized at the optimal prior for meta testing, and this is not the case for the train-train method in general without structural assumptions on the data. In contrast, in the realizable case where the data are generated from linear models, we show that both the train-val and train-train losses are minimized at the optimal prior in expectation. Further, perhaps surprisingly, our main result shows that the train-train method achieves a \emph{strictly better} excess loss in this realizable case, even when the regularization parameter and split ratio are optimally tuned for both methods. Our results highlight that sample splitting may not always be preferable, especially when the data is realizable by the model. We validate our theories by experimentally showing that the train-train method can indeed outperform the train-val method, on both simulations and real meta-learning tasks. Yu Bai 0017, Minshuo Chen, Pan Zhou 0002, Tuo Zhao, Jason D. Lee, Sham M. Kakade, Huan Wang 0016, Caiming Xiong |
ICML | 4 |
| 2021 | Besov Function Approximation and Binary Classification on Low-Dimensional Manifolds Using Convolutional Residual NetworksabstractMost of existing statistical theories on deep neural networks have sample complexities cursed by the data dimension and therefore cannot well explain the empirical success of deep learning on high-dimensional data. To bridge this gap, we propose to exploit the low-dimensional structures of the real world datasets and establish theoretical guarantees of convolutional residual networks (ConvResNet) in terms of function approximation and statistical recovery for binary classification problem. Specifically, given the data lying on a $d$-dimensional manifold isometrically embedded in $\mathbb{R}^D$, we prove that if the network architecture is properly chosen, ConvResNets can (1) approximate {\it Besov functions} on manifolds with arbitrary accuracy, and (2) learn a classifier by minimizing the empirical logistic risk, which gives an {\it excess risk} in the order of $n^{-\frac{s}{2s+2(s\vee d)}}$, where $s$ is a smoothness parameter. This implies that the sample complexity depends on the intrinsic dimension $d$, instead of the data dimension $D$. Our results demonstrate that ConvResNets are adaptive to low-dimensional structures of data sets. Hao Liu 0028, Minshuo Chen, Tuo Zhao, Wenjing Liao |
ICML | 3 |
| 2021 | Fine-Tuning Pre-trained Language Model with Weak Supervision: A Contrastive-Regularized Self-Training ApproachabstractYue Yu, Simiao Zuo, Haoming Jiang, Wendi Ren, Tuo Zhao, Chao Zhang. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Yue Yu 0001, Simiao Zuo, Haoming Jiang, Wendi Ren, Tuo Zhao, Chao Zhang 0014 |
NAACL-HLT | 5 |
| 2021 | Pessimism Meets Invariance: Provably Efficient Offline Mean-Field Multi-Agent RLabstractMean-Field Multi-Agent Reinforcement Learning (MF-MARL) is attractive in the applications involving a large population of homogeneous agents, as it exploits the permutation invariance of agents and avoids the curse of many agents. Most existing results only focus on online settings, in which agents can interact with the environment during training. In some applications such as social welfare optimization, however, the interaction during training can be prohibitive or even unethical in the societal systems. To bridge such a gap, we propose a SAFARI (peSsimistic meAn-Field vAlue iteRatIon) algorithm for off-line MF-MARL, which only requires a handful of pre-collected experience data. Theoretically, under a weak coverage assumption that the experience dataset contains enough information about the optimal policy, we prove that for an episodic mean-field MDP with a horizon $H$ and $N$ training trajectories, SAFARI attains a sub-optimality gap of $\mathcal{O}(H^2d_{\rm eff} /\sqrt{N})$, where $d_{\rm eff}$ is the effective dimension of the function class for parameterizing the value function, but independent on the number of agents. Numerical experiments are provided. Minshuo Chen, Yan Li 0074, Ethan Wang, Zhuoran Yang, Zhaoran Wang 0001, Tuo Zhao |
NeurIPS | 6 |
| 2020 | SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationabstractTransfer learning has fundamentally changed the landscape of natural language processing (NLP).Many state-of-the-art models are first pre-trained on a large text corpus and then fine-tuned on downstream tasks.However, due to limited data resources from downstream tasks and the extremely high complexity of pre-trained models, aggressive fine-tuning often causes the fine-tuned model to overfit the training data of downstream tasks and fail to generalize to unseen data.To address such an issue in a principled manner, we propose a new learning framework for robust and efficient fine-tuning for pre-trained models to attain better generalization performance.The proposed framework contains two important ingredients: 1. Smoothness-inducing regularization, which effectively manages the complexity of the model; 2. Bregman proximal point optimization, which is an instance of trustregion methods and can prevent aggressive updating.Our experiments show that the proposed framework achieves new state-of-the-art performance on a number of NLP tasks including GLUE, SNLI, SciTail and ANLI.Moreover, it also outperforms the state-of-the-art T5 model, which is the largest pre-trained model containing 11 billion parameters, on GLUE. 1 Haoming Jiang, Weizhu Chen, Xiaodong Liu 0003, Jianfeng Gao 0001, Tuo Zhao |
ACL | 6 |
| 2020 | Multi-Domain Neural Machine Translation with Word-Level Adaptive Layer-wise Domain MixingabstractMany multi-domain neural machine translation (NMT) models achieve knowledge transfer by enforcing one encoder to learn shared embedding across domains.However, this design lacks adaptation to individual domains.To overcome this limitation, we propose a novel multi-domain NMT model using individual modules for each domain, on which we apply word-level, adaptive and layer-wise domain mixing.We first observe that words in a sentence are often related to multiple domains.Hence, we assume each word has a domain proportion, which indicates its domain preference.Then word representations are obtained by mixing their embedding in individual domains based on their domain proportions.We show this can be achieved by carefully designing multi-head dot-product attention modules for different domains, and eventually taking weighted averages of their parameters by word-level layer-wise domain proportions.Through this, we can achieve effective domain knowledge sharing, and capture fine-grained domain-specific knowledge as well.Our experiments show that our proposed model outperforms existing ones in several NMT tasks. Haoming Jiang, Chen Liang 0006, Tuo Zhao |
ACL | 4 |
| 2020 | On Generalization Bounds of a Family of Recurrent Neural NetworksabstractRecurrent Neural Networks (RNNs) have been widely applied to sequential data analysis. Due to their complicated modeling structures, however, the theory behind is still largely missing. To connect theory and practice, we study the generalization properties of vanilla RNNs as well as their variants, including Minimal Gated Unit (MGU), Long Short Term Memory (LSTM), and Convolutional (Conv) RNNs. Specifically, our theory is established under the PAC-Learning framework. The generalization bound is presented in terms of the spectral norms of the weight matrices and the total number of parameters. We also establish refined generalization bounds with additional norm assumptions, and draw a comparison among these bounds. We remark: (1) Our generalization bound for vanilla RNNs is significantly tighter than the best of existing results; (2) We are not aware of any other generalization bounds for MGU and LSTM RNNs in the exiting literature; (3) We demonstrate the advantages of these variants in generalization. Minshuo Chen, Xingguo Li, Tuo Zhao |
AISTATS | 3 |
| 2020 | Calibrated Language Model Fine-Tuning for In- and Out-of-Distribution DataabstractFine-tuned pre-trained language models can suffer from severe miscalibration for both in-distribution and out-of-distribution (OOD) data due to over-parameterization.To mitigate this issue, we propose a regularized fine-tuning method.Our method introduces two types of regularization for better calibration: (1) On-manifold regularization, which generates pseudo on-manifold samples through interpolation within the data manifold.Augmented training with these pseudo samples imposes a smoothness regularization to improve in-distribution calibration.(2) Off-manifold regularization, which encourages the model to output uniform distributions for pseudo off-manifold samples to address the over-confidence issue for OOD data.Our experiments demonstrate that the proposed method outperforms existing calibration methods for text classification in terms of expectation calibration error, misclassification detection, and OOD detection on six datasets.Our code can be found at https://github.com/Lingkai-Kong/ Calibrated-BERT-Fine-Tuning. Haoming Jiang, Yuchen Zhuang, Jie Lyu 0005, Tuo Zhao, Chao Zhang 0014 |
EMNLP (1) | 5 |
| 2020 | On Computation and Generalization of Generative Adversarial Imitation Learning
Minshuo Chen, Yizhou Wang 0006, Zhuoran Yang, Xingguo Li, Zhaoran Wang 0001, Tuo Zhao |
ICLR | 7 |
| 2020 | Implicit Bias of Gradient Descent based Adversarial Training on Separable Data
Yan Li 0074, Ethan X. Fang, Tuo Zhao |
ICLR | 4 |
| 2020 | Deep Reinforcement Learning with Robust and Smooth PolicyabstractDeep reinforcement learning (RL) has achieved great empirical successes in various domains. However, the large search space of neural networks requires a large amount of data, which makes the current RL algorithms not sample efficient. Motivated by the fact that many environments with continuous state space have smooth transitions, we propose to learn a smooth policy that behaves smoothly with respect to states. We develop a new framework — \textbf{S}mooth \textbf{R}egularized \textbf{R}einforcement \textbf{L}earning ($\textbf{SR}^2\textbf{L}$), where the policy is trained with smoothness-inducing regularization. Such regularization effectively constrains the search space, and enforces smoothness in the learned policy. Moreover, our proposed framework can also improve the robustness of policy against measurement error in the state space, and can be naturally extended to distribubutionally robust setting. We apply the proposed framework to both on-policy (TRPO) and off-policy algorithm (DDPG). Through extensive experiments, we demonstrate that our method achieves improved sample efficiency and robustness. Qianli Shen, Yan Li 0074, Haoming Jiang, Zhaoran Wang 0001, Tuo Zhao |
ICML | 5 |
| 2020 | Transformer Hawkes ProcessabstractModern data acquisition routinely produce massive amounts of event sequence data in various domains, such as social media, healthcare, and financial markets. These data often exhibit complicated short-term and long-term temporal dependencies. However, most of the existing recurrent neural network based point process models fail to capture such dependencies, and yield unreliable prediction performance. To address this issue, we propose a Transformer Hawkes Process (THP) model, which leverages the self-attention mechanism to capture long-term dependencies and meanwhile enjoys computational efficiency. Numerical experiments on various datasets show that THP outperforms existing models in terms of both likelihood and event prediction accuracy by a notable margin. Moreover, THP is quite general and can incorporate additional structural knowledge. We provide a concrete example, where THP achieves improved prediction performance for learning multiple point processes when incorporating their relational information. Simiao Zuo, Haoming Jiang, Zichong Li, Tuo Zhao, Hongyuan Zha |
ICML | 4 |
| 2020 | BOND: BERT-Assisted Open-Domain Named Entity Recognition with Distant SupervisionabstractWe study the open-domain named entity recognition (NER) problem under distant supervision. The distant supervision, though does not require large amounts of manual annotations, yields highly incomplete and noisy distant labels via external knowledge bases. To address this challenge, we propose a new computational framework -- BOND, which leverages the power of pre-trained language models (e.g., BERT and RoBERTa) to improve the prediction performance of NER models. Specifically, we propose a two-stage training algorithm: In the first stage, we adapt the pre-trained language model to the NER tasks using the distant labels, which can significantly improve the recall and precision; In the second stage, we drop the distant labels, and propose a self-training approach to further improve the model performance. Thorough experiments on 5 benchmark datasets demonstrate the superiority of BOND over existing distantly supervised NER methods. The code and distantly labeled data have been released in https://github.com/cliang1453/BOND. Chen Liang 0006, Yue Yu 0001, Haoming Jiang, Siawpeng Er, Tuo Zhao, Chao Zhang 0014 |
KDD | 6 |
| 2020 | Towards Understanding Hierarchical Learning: Benefits of Neural RepresentationsabstractDeep neural networks can empirically perform efficient hierarchical learning, in which the layers learn useful representations of the data. However, how they make use of the intermediate representations are not explained by recent theories that relate them to ``shallow learners'' such as kernels. In this work, we demonstrate that intermediate \emph{neural representations} add more flexibility to neural networks and can be advantageous over raw inputs. We consider a fixed, randomly initialized neural network as a representation function fed into another trainable network. When the trainable network is the quadratic Taylor model of a wide two-layer network, we show that neural representation can achieve improved sample complexities compared with the raw input: For learning a low-rank degree-$p$ polynomial ($p \geq 4$) in $d$ dimension, neural representation requires only $\widetilde{O}(d^{\ceil{p/2}})$ samples, while the best-known sample complexity upper bound for the raw input is $\widetilde{O}(d^{p-1})$. We contrast our result with a lower bound showing that neural representations do not improve over the raw input (in the infinite width limit), when the trainable network is instead a neural tangent kernel. Our results characterize when neural representations are beneficial, and may provide a new perspective on why depth is important in deep learning. Minshuo Chen, Yu Bai 0017, Jason D. Lee, Tuo Zhao, Huan Wang 0016, Caiming Xiong, Richard Socher |
NeurIPS | 4 |
| 2020 | Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveabstractDeep residual networks (ResNets) have demonstrated better generalization performance than deep feedforward networks (FFNets). However, the theory behind such a phenomenon is still largely unknown. This paper studies this fundamental problem in deep learning from a so-called ``neural tangent kernel'' perspective. Specifically, we first show that under proper conditions, as the width goes to infinity, training deep ResNets can be viewed as learning reproducing kernel functions with some kernel function. We then compare the kernel of deep ResNets with that of deep FFNets and discover that the class of functions induced by the kernel of FFNets is asymptotically not learnable, as the depth goes to infinity. In contrast, the class of functions induced by the kernel of ResNets does not exhibit such degeneracy. Our discovery partially justifies the advantages of deep ResNets over deep FFNets in generalization abilities. Numerical results are provided to support our claim. Kaixuan Huang, Yuqing Wang 0005, Molei Tao, Tuo Zhao |
NeurIPS | 4 |
| 2020 | Differentiable Top-k with Optimal TransportabstractFinding the k largest or smallest elements from a collection of scores, i.e., top-k operation, is an important model component widely used in information retrieval, machine learning, and data mining. However, if the top-k operation is implemented in an algorithmic way, e.g., using bubble algorithm, the resulted model cannot be trained in an end-to-end way using prevalent gradient descent algorithms. This is because these implementations typically involve swapping indices, whose gradient cannot be computed. Moreover, the corresponding mapping from the input scores to the indicator vector of whether this element belongs to the top-k set is essentially discontinuous. To address the issue, we propose a smoothed approximation, namely SOFT (Scalable Optimal transport-based diFferenTiable) top-k operator. Specifically, our SOFT top-k operator approximates the output of top-k operation as the solution of an Entropic Optimal Transport (EOT) problem. The gradient of the SOFT operator can then be efficiently approximated based on the optimality conditions of EOT problem. We then apply the proposed operator to k-nearest neighbors algorithm and beam search algorithm. The numerical experiment demonstrates their achieve improved performance. Yujia Xie, Hanjun Dai, Minshuo Chen, Bo Dai 0001, Tuo Zhao, Hongyuan Zha, Wei Wei 0019, Tomas Pfister |
NeurIPS | 5 |
| 2019 | On Constrained Nonconvex Stochastic Optimization: A Case Study for Generalized Eigenvalue DecompositionabstractWe study constrained nonconvex optimization problems in machine learning and signal processing. It is well-known that these problems can be rewritten to a min-max problem in a Lagrangian form. However, due to the lack of convexity, their landscape is not well understood and how to find the stable equilibria of the Lagrangian function is still unknown. To bridge the gap, we study the landscape of the Lagrangian function. Further, we define a special class of Lagrangian functions. They enjoy the following two properties: 1. Equilibria are either stable or unstable (Formal definition in Section 2); 2.Stable equilibria correspond to the global optima of the original problem. We show that a generalized eigenvalue (GEV) problem, including canonical correlation analysis and other problems as special examples, belongs to the class. Specifically, we characterize its stable and unstable equilibria by leveraging an invariant group and symmetric property (more details in Section 3). Motivated by these neat geometric structures, we propose a simple, efficient, and stochastic primal-dual algorithm solving the online GEV problem. Theoretically, under sufficient conditions, we establish an asymptotic rate of convergence and obtain the first sample complexity result for the online GEV problem by diffusion approximations, which are widely used in applied probability. Numerical results are also provided to support our theory. Zhehui Chen, Xingguo Li, Lin Yang 0011, Jarvis D. Haupt, Tuo Zhao |
AISTATS | 5 |
| 2019 | On Computation and Generalization of Generative Adversarial Networks under Spectrum Control
Haoming Jiang, Zhehui Chen, Minshuo Chen, Dingding Wang 0001, Tuo Zhao |
ICLR (Poster) | 6 |
| 2019 | On Scalable and Efficient Computation of Large Scale Optimal TransportabstractOptimal Transport (OT) naturally arises in many machine learning applications, yet the heavy computational burden limits its wide-spread uses. To address the scalability issue, we propose an implicit generative learning-based framework called SPOT (Scalable Push-forward of Optimal Transport). Specifically, we approximate the optimal transport plan by a pushforward of a reference distribution, and cast the optimal transport problem into a minimax problem. We then can solve OT problems efficiently using primal dual stochastic gradient-type algorithms. We also show that we can recover the density of the optimal transport plan using neural ordinary differential equations. Numerical experiments on both synthetic and real datasets illustrate that SPOT is robust and has favorable convergence behavior. SPOT also allows us to efficiently sample from the optimal transport plan, which benefits downstream applications such as domain adaptation. Yujia Xie, Minshuo Chen, Haoming Jiang, Tuo Zhao, Hongyuan Zha |
ICML | 4 |
| 2019 | Toward Understanding the Importance of Noise in Training Neural NetworksabstractNumerous empirical evidence has corroborated that the noise plays a crucial rule in effective and efficient training of deep neural networks. The theory behind, however, is still largely unknown. This paper studies this fundamental problem through training a simple two-layer convolutional neural network model. Although training such a network requires to solve a non-convex optimization problem with a spurious local optimum and a global optimum, we prove that a perturbed gradient descent algorithm in conjunction with noise annealing is guaranteed to converge to a global optimum in polynomial time with arbitrary initialization. This implies that the noise enables the algorithm to efficiently escape from the spurious local optimum. Numerical experiments are provided to support our theory. Yan Li 0074, Dachao Lin, Enlu Zhou, Tuo Zhao |
ICML | 6 |
| 2019 | Efficient Approximation of Deep ReLU Networks for Functions on Low Dimensional ManifoldsabstractDeep neural networks have revolutionized many real world applications, due to their flexibility in data fitting and accurate predictions for unseen data. A line of research reveals that neural networks can approximate certain classes of functions with an arbitrary accuracy, while the size of the network scales exponentially with respect to the data dimension. Empirical results, however, suggest that networks of moderate size already yield appealing performance. To explain such a gap, a common belief is that many data sets exhibit low dimensional structures, and can be modeled as samples near a low dimensional manifold. In this paper, we prove that neural networks can efficiently approximate functions supported on low dimensional manifolds. The network size scales exponentially in the approximation error, with an exponent depending on the intrinsic dimension of the data and the smoothness of the function. Our result shows that exploiting low dimensional data structures can greatly enhance the efficiency in function approximation by neural networks. We also implement a sub-network that assigns input data to their corresponding local neighborhoods, which may be of independent interest. Minshuo Chen, Haoming Jiang, Wenjing Liao, Tuo Zhao |
NeurIPS | 4 |
| 2019 | Towards Understanding the Importance of Shortcut Connections in Residual NetworksabstractResidual Network (ResNet) is undoubtedly a milestone in deep learning. ResNet is equipped with shortcut connections between layers, and exhibits efficient training using simple first order algorithms. Despite of the great empirical success, the reason behind is far from being well understood. In this paper, we study a two-layer non-overlapping convolutional ResNet. Training such a network requires solving a non-convex optimization problem with a spurious local optimum. We show, however, that gradient descent combined with proper normalization, avoids being trapped by the spurious local optimum, and converges to a global optimum in polynomial time, when the weight of the first layer is initialized at 0, and that of the second layer is initialized arbitrarily in a ball. Numerical experiments are provided to support our theory. Minshuo Chen, Simon S. Du, Enlu Zhou, Tuo Zhao |
NeurIPS | 6 |
| 2019 | Meta Learning with Relational Information for Short SequencesabstractThis paper proposes a new meta-learning method -- named HARMLESS (HAwkes Relational Meta Learning method for Short Sequences) for learning heterogeneous point process models from a collection of short event sequence data along with a relational network. Specifically, we propose a hierarchical Bayesian mixture Hawkes process model, which naturally incorporates the relational information among sequences into point process modeling. Compared with existing methods, our model can capture the underlying mixed-community patterns of the relational network, which simultaneously encourages knowledge sharing among sequences and facilitates adaptively learning for each individual sequence. We further propose an efficient stochastic variational meta-EM algorithm, which can scale to large problems. Numerical experiments on both synthetic and real data show that HARMLESS outperforms existing methods in terms of predicting the future events. Yujia Xie, Haoming Jiang, Tuo Zhao, Hongyuan Zha |
NeurIPS | 4 |
| 2019 | On Fast Convergence of Proximal Algorithms for SQRT-Lasso Optimization: Don't Worry About its Nonsmooth Loss Function
Xingguo Li, Haoming Jiang, Jarvis D. Haupt, Raman Arora, Han Liu 0001, Mingyi Hong 0001, Tuo Zhao |
UAI | 7 |
| 2019 | Online Factorization and Partition of Complex Networks by Random Walk
Lin Yang 0011, Vladimir Braverman, Tuo Zhao, Mengdi Wang 0001 |
UAI | 4 |
| 2019 | Picasso: A Sparse Learning Library for High Dimensional Data Analysis in R and PythonabstractWe describe a new library named picasso, which implements a unified framework of pathwise coordinate optimization for a variety of sparse learning problems (e.g., sparse linear regression, sparse logistic regression, sparse Poisson regression and scaled sparse linear regression) combined with efficient active set selection strategies. Besides, the library allows users to choose different sparsity-inducing regularizers, including the convex $\ell_1$, nonvoncex MCP and SCAD regularizers. The library is coded in \texttt{C++} and has user-friendly R and Python wrappers. Numerical experiments demonstrate that picasso can scale up to large problems efficiently. Jason Ge, Xingguo Li, Haoming Jiang, Han Liu 0001, Tong Zhang 0001, Mengdi Wang 0001, Tuo Zhao |
J. Mach. Learn. Res. | 7 |
| 2019 | Symmetry, Saddle Points, and Global Optimization Landscape of Nonconvex Matrix FactorizationabstractWe propose a general theory for studying the landscape of nonconvex optimization with underlying symmetric structures for a class of machine learning problems (e.g., low-rank matrix factorization, phase retrieval, and deep linear neural networks). In particular, we characterize the locations of stationary points and the null space of Hessian matrices of the objective function via the lens of invariant groups. As a major motivating example, we apply the proposed general theory to characterize the global landscape of the nonconvex optimization in low-rank matrix factorization problem. We illustrate how the rotational symmetry group gives rise to infinitely many nonisolated strict saddle points and equivalent global minima of the objective function. By explicitly identifying all stationary points, we divide the entire parameter space into three regions: (R1) the region containing the neighborhoods of all strict saddle points where the objective has negative curvature; (R2) the region containing neighborhoods of all global minima, where the objective enjoys strong convexity along certain directions; and (R3) the complement of the above regions, where the gradient has sufficiently large magnitude. We further extend our result to the matrix sensing problem. Such global landscape implies that strong global convergence guarantees for popular iterative algorithms with arbitrary initial solutions. Xingguo Li, Raman Arora, Jarvis D. Haupt, Han Liu 0001, Zhaoran Wang 0001, Tuo Zhao |
IEEE Trans. Inf. Theory | 7 |
| 2018 | Dimensionality Reduction for Stationary Time Series via Stochastic Nonconvex OptimizationabstractStochastic optimization naturally arises in machine learning. Efficient algorithms with provable guarantees, however, are still largely missing, when the objective function is nonconvex and the data points are dependent. This paper studies this fundamental challenge through a streaming PCA problem for stationary time series data. Specifically, our goal is to estimate the principle component of time series data with respect to the covariance matrix of the stationary distribution. Computationally, we propose a variant of Oja's algorithm combined with downsampling to control the bias of the stochastic gradient caused by the data dependency. Theoretically, we quantify the uncertainty of our proposed stochastic algorithm based on diffusion approximations. This allows us to prove the asymptotic rate of convergence and further implies near optimal asymptotic sample complexity. Numerical experiments are provided to support our analysis. Minshuo Chen, Lin Yang 0011, Mengdi Wang 0001, Tuo Zhao |
NeurIPS | 4 |
| 2018 | Towards Understanding Acceleration Tradeoff between Momentum and Asynchrony in Nonconvex Stochastic OptimizationabstractAsynchronous momentum stochastic gradient descent algorithms (Async-MSGD) have been widely used in distributed machine learning, e.g., training large collaborative filtering systems and deep neural networks. Due to current technical limit, however, establishing convergence properties of Async-MSGD for these highly complicated nonoconvex problems is generally infeasible. Therefore, we propose to analyze the algorithm through a simpler but nontrivial nonconvex problems --- streaming PCA. This allows us to make progress toward understanding Aync-MSGD and gaining new insights for more general problems. Specifically, by exploiting the diffusion approximation of stochastic optimization, we establish the asymptotic rate of convergence of Async-MSGD for streaming PCA. Our results indicate a fundamental tradeoff between asynchrony and momentum: To ensure convergence and acceleration through asynchrony, we have to reduce the momentum (compared with Sync-MSGD). To the best of our knowledge, this is the first theoretical attempt on understanding Async-MSGD for distributed nonconvex stochastic optimization. Numerical experiments on both streaming PCA and training deep neural networks are provided to support our findings for Async-MSGD. Jianping Shi, Enlu Zhou, Tuo Zhao |
NeurIPS | 5 |
| 2018 | The Physical Systems Behind Optimization AlgorithmsabstractWe use differential equations based approaches to provide some {\it \textbf{physics}} insights into analyzing the dynamics of popular optimization algorithms in machine learning. In particular, we study gradient descent, proximal gradient descent, coordinate gradient descent, proximal coordinate gradient, and Newton's methods as well as their Nesterov's accelerated variants in a unified framework motivated by a natural connection of optimization algorithms to physical systems. Our analysis is applicable to more general algorithms and optimization problems {\it \textbf{beyond}} convexity and strong convexity, e.g. Polyak-\L ojasiewicz and error bound conditions (possibly nonconvex). Lin Yang 0011, Raman Arora, Vladimir Braverman, Tuo Zhao |
NeurIPS | 4 |
| 2018 | Provable Gaussian Embedding with One ObservationabstractThe success of machine learning methods heavily relies on having an appropriate representation for data at hand. Traditionally, machine learning approaches relied on user-defined heuristics to extract features encoding structural information about data. However, recently there has been a surge in approaches that learn how to encode the data automatically in a low dimensional space. Exponential family embedding provides a probabilistic framework for learning low-dimensional representation for various types of high-dimensional data. Though successful in practice, theoretical underpinnings for exponential family embeddings have not been established. In this paper, we study the Gaussian embedding model and develop the first theoretical results for exponential family embedding models. First, we show that, under a mild condition, the embedding structure can be learned from one observation by leveraging the parameter sharing between different contexts even though the data are dependent with each other. Second, we study properties of two algorithms used for learning the embedding structure and establish convergence results for each of them. The first algorithm is based on a convex relaxation, while the other solved the non-convex formulation of the problem directly. Experiments demonstrate the effectiveness of our approach. Ming Yu 0005, Zhuoran Yang, Tuo Zhao, Mladen Kolar, Zhaoran Wang 0001 |
NeurIPS | 3 |
| 2017 | Online Partial Least Square Optimization: Dropping Convexity for Better Efficiency and ScalabilityabstractMultiview representation learning is popular for latent factor analysis. Many existing approaches formulate the multiview representation learning as convex optimization problems, where global optima can be obtained by certain algorithms in polynomial time. However, many evidences have corroborated that heuristic nonconvex approaches also have good empirical computational performance and convergence to the global optima, although there is a lack of theoretical justification. Such a gap between theory and practice motivates us to study a nonconvex formulation for multiview representation learning, which can be efficiently solved by a simple stochastic gradient descent method. By analyzing the dynamics of the algorithm based on diffusion processes, we establish a global rate of convergence to the global optima. Numerical experiments are provided to support our theory. Zhehui Chen, Lin Yang 0011, Chris Junchi Li, Tuo Zhao |
ICML | 4 |
| 2017 | The Opensesame NIST 2016 Speaker Recognition Evaluation System
Qi Qian 0001, Zhibin Wang 0004, Qingen Zhao, Tianzhou Wang, Hao Li 0030, Shenghuo Zhu, Rong Jin 0001, Tuo Zhao |
INTERSPEECH | 10 |
| 2017 | On Quadratic Convergence of DC Proximal Newton Algorithm in Nonconvex Sparse LearningabstractWe propose a DC proximal Newton algorithm for solving nonconvex regularized sparse learning problems in high dimensions. Our proposed algorithm integrates the proximal newton algorithm with multi-stage convex relaxation based on the difference of convex (DC) programming, and enjoys both strong computational and statistical guarantees. Specifically, by leveraging a sophisticated characterization of sparse modeling structures (i.e., local restricted strong convexity and Hessian smoothness), we prove that within each stage of convex relaxation, our proposed algorithm achieves (local) quadratic convergence, and eventually obtains a sparse approximate local optimum with optimal statistical properties after only a few convex relaxations. Numerical experiments are provided to support our theory. Xingguo Li, Lin Yang 0011, Jason Ge, Jarvis D. Haupt, Tong Zhang 0001, Tuo Zhao |
NIPS | 6 |
| 2017 | Deep Hyperspherical LearningabstractConvolution as inner product has been the founding basis of convolutional neural networks (CNNs) and the key to end-to-end visual representation learning. Benefiting from deeper architectures, recent CNNs have demonstrated increasingly strong representation abilities. Despite such improvement, the increased depth and larger parameter space have also led to challenges in properly training a network. In light of such challenges, we propose hyperspherical convolution (SphereConv), a novel learning framework that gives angular representations on hyperspheres. We introduce SphereNet, deep hyperspherical convolution networks that are distinct from conventional inner product based convolutional networks. In particular, SphereNet adopts SphereConv as its basic convolution operator and is supervised by generalized angular softmax loss - a natural loss formulation under SphereConv. We show that SphereNet can effectively encode discriminative representation and alleviate training difficulty, leading to easier optimization, faster convergence and comparable (even better) classification accuracy over convolutional counterparts. We also provide some theoretical insights for the advantages of learning on hyperspheres. In addition, we introduce the learnable SphereConv, i.e., a natural improvement over prefixed SphereConv, and SphereNorm, i.e., hyperspherical learning as a normalization method. Experiments have verified our conclusions. Weiyang Liu, Xingguo Li, Zhen Liu 0019, Bo Dai 0001, Tuo Zhao |
NIPS | 6 |
| 2017 | Parametric Simplex Method for Sparse LearningabstractHigh dimensional sparse learning has imposed a great computational challenge to large scale data analysis. In this paper, we investiage a broad class of sparse learning approaches formulated as linear programs parametrized by a {\em regularization factor}, and solve them by the parametric simplex method (PSM). PSM offers significant advantages over other competing methods: (1) PSM naturally obtains the complete solution path for all values of the regularization parameter; (2) PSM provides a high precision dual certificate stopping criterion; (3) PSM yields sparse solutions through very few iterations, and the solution sparsity significantly reduces the computational cost per iteration. Particularly, we demonstrate the superiority of PSM over various sparse learning approaches, including Dantzig selector for sparse linear regression, sparse support vector machine for sparse linear classification, and sparse differential network estimation. We then provide sufficient conditions under which PSM always outputs sparse solutions such that its computational performance can be significantly boosted. Thorough numerical experiments are provided to demonstrate the outstanding performance of the PSM method. Haotian Pang, Han Liu 0001, Robert J. Vanderbei, Tuo Zhao |
NIPS | 4 |
| 2017 | Excavation equipment classification based on improved MFCC features and ELM
Jiuwen Cao, Tuo Zhao, Jianzhong Wang 0003, Ruirong Wang, Yun Chen 0008 |
Neurocomputing | 2 |
| 2017 | On Faster Convergence of Cyclic Block Coordinate Descent-type Methods for Strongly Convex Minimization
Xingguo Li, Tuo Zhao, Raman Arora, Han Liu 0001, Mingyi Hong 0001 |
J. Mach. Learn. Res. | 2 |
| 2016 | An Improved Convergence Analysis of Cyclic Block Coordinate Descent-type Methods for Strongly Convex MinimizationabstractThe cyclic block coordinate descent-type (CBCD-type) methods have shown remarkable computational performance for solving strongly convex minimization problems. Typical applications include many popular statistical machine learning methods such as elastic-net regression, ridge penalized logistic regression, and sparse additive regression. Existing optimization literature has shown that the CBCD-type methods attain iteration complexity of O(p⋅\log(1/ε)), where εis a pre-specified accuracy of the objective value, and p is the number of blocks. However, such iteration complexity explicitly depends on p, and therefore is at least p times worse than those of gradient descent methods. To bridge this theoretical gap, we propose an improved convergence analysis for the CBCD-type methods. In particular, we first show that for a family of quadratic minimization problems, the iteration complexity of the CBCD-type methods matches that of the GD methods in term of dependency on p (up to a \log^2 p factor). Thus our complexity bounds are sharper than the existing bounds by at least a factor of p/\log^2p. We also provide a lower bound to confirm that our improved complexity bounds are tight (up to a \log^2 p factor) if the largest and smallest eigenvalues of the Hessian matrix do not scale with p. Finally, we generalize our analysis to other strongly convex minimization problems beyond quadratic ones. Xingguo Li, Tuo Zhao, Raman Arora, Han Liu 0001, Mingyi Hong 0001 |
AISTATS | 2 |
| 2016 | Stochastic Variance Reduced Optimization for Nonconvex Sparse LearningabstractWe propose a stochastic variance reduced optimization algorithm for solving a class of large-scale nonconvex optimization problems with cardinality constraints, and provide sufficient conditions under which the proposed algorithm enjoys strong linear convergence guarantees and optimal estimation accuracy in high dimensions. Numerical experiments demonstrate the efficiency of our method in terms of both parameter estimation and computational performance. Xingguo Li, Tuo Zhao, Raman Arora, Han Liu 0001, Jarvis D. Haupt |
ICML | 2 |
| 2016 | Subpixel mapping of hyperspectral images based on collaborative representationabstractSubpixel mapping with a low resolution hyperspectral image as the only input is widely applicable due to the fact that auxiliary image with high spatial resolution is not always available in practice. In this paper, to extract spatial information without auxiliary image, the upscaled low resolution hyperspectral image is classified using collaborative representation-based classifier. Another subpixel scale classification map is available by the combination of collaborative representation-based classification, spectral unmixing and subpixel spatial attraction model. To achieve better classification performance, decision fusion is employed to elect approximate class label from these two initial classification maps for each subpixel by the voting of the neighboring subpixels. Experimental results illustrate that the proposed approach is more promising in extracting and utilizing spatial information compared with some state-of-the-art subpixel mapping approaches. Xiaoqin Xue, Yifan Zhang 0006, Tuo Zhao, Mingyi He |
IGARSS | 3 |
| 2016 | Hyperspectral and multispectral image fusion using collaborative representation with local adaptive dictionary pairabstractIn this paper, the spatial resolution of hyperspectral image (HSI) is enhanced by fusing it with multispectral image (MSI) of the same scene with a higher spatial resolution. The w-hole spectrum covered by HSI channels is divided into several regions according to MSI spectral channels. The HSI-MSI fusion problem is then simplified by fusing images of each spectral region one after another. Specifically, a fusion algorithm based on collaborative representation (CR) with local adaptive dictionary pair is proposed. Compared to the classic global dictionary, the scale of the local adaptive one is much smaller such that the related computational cost is also reduced. The employment of CR is capable of reducing the reconstruction error to guarantee an improved fusion performance. Simulative experiments are deployed for illustration and comparison. Tuo Zhao, Yifan Zhang 0006, Xiaoqin Xue, Mingyi He |
IGARSS | 1 |
| 2016 | NESTT: A Nonconvex Primal-Dual Splitting Method for Distributed and Stochastic OptimizationabstractWe study a stochastic and distributed algorithm for nonconvex problems whose objective consists a sum $N$ nonconvex $L_i/N$-smooth functions, plus a nonsmooth regularizer. The proposed NonconvEx primal-dual SpliTTing (NESTT) algorithm splits the problem into $N$ subproblems, and utilizes an augmented Lagrangian based primal-dual scheme to solve it in a distributed and stochastic manner. With a special non-uniform sampling, a version of NESTT achieves $\epsilon$-stationary solution using $\mathcal{O}((\sum_{i=1}^N\sqrt{L_i/N})^2/\epsilon)$ gradient evaluations, which can be up to $\mathcal{O}(N)$ times better than the (proximal) gradient descent methods. It also achieves Q-linear convergence rate for nonconvex $\ell_1$ penalized quadratic problems with polyhedral constraints. Further, we reveal a fundamental connection between {\it primal-dual} based methods and a few {\it primal only} methods such as IAG/SAG/SAGA. Davood Hajinezhad, Mingyi Hong 0001, Tuo Zhao, Zhaoran Wang 0001 |
NIPS | 3 |
| 2015 | Time-frequency kernel-based CNN for speech recognition
Tuo Zhao, Yunxin Zhao |
INTERSPEECH | 1 |
| 2015 | A Nonconvex Optimization Framework for Low Rank Matrix EstimationabstractWe study the estimation of low rank matrices via nonconvex optimization. Compared with convex relaxation, nonconvex optimization exhibits superior empirical performance for large scale instances of low rank matrix estimation. However, the understanding of its theoretical guarantees are limited. In this paper, we define the notion of projected oracle divergence based on which we establish sufficient conditions for the success of nonconvex optimization. We illustrate the consequences of this general framework for matrix sensing and completion. In particular, we prove that a broad class of nonconvex optimization algorithms, including alternating minimization and gradient-type methods, geometrically converge to the global optimum and exactly recover the true low rank matrices under standard conditions. Tuo Zhao, Zhaoran Wang 0001, Han Liu 0001 |
NIPS | 1 |
| 2015 | The flare package for high dimensional linear regression and precision matrix estimation in R
Xingguo Li, Tuo Zhao, Xiaoming Yuan 0001, Han Liu 0001 |
J. Mach. Learn. Res. | 2 |
| 2015 | Calibrated multivariate regression with application to neural semantic basis discovery
Han Liu 0001, Lie Wang 0002, Tuo Zhao |
J. Mach. Learn. Res. | 3 |
| 2014 | Multivariate Regression with Calibration
Han Liu 0001, Lie Wang 0002, Tuo Zhao |
NIPS | 3 |
| 2014 | Accelerated Mini-batch Randomized Block Coordinate Descent Method
Tuo Zhao, Mo Yu, Raman Arora, Han Liu 0001 |
NIPS | 1 |
| 2014 | Calibrated Precision Matrix Estimation for High-Dimensional Elliptical DistributionsabstractWe propose a semiparametric method for estimating a precision matrix of high-dimensional elliptical distributions. Unlike most existing methods, our method naturally handles heavy tailness and conducts parameter estimation under a calibration framework, thus achieves improved theoretical rates of convergence and finite sample performance on heavy-tail applications. We further demonstrate the performance of the proposed method using thorough numerical experiments. Tuo Zhao, Han Liu 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2013 | Sparse Inverse Covariance Estimation with CalibrationabstractWe propose a semiparametric procedure for estimating high dimensional sparse inverse covariance matrix. Our method, named ALICE, is applicable to the elliptical family. Computationally, we develop an efficient dual inexact iterative projection (${\rm D_2}$P) algorithm based on the alternating direction method of multipliers (ADMM). Theoretically, we prove that the ALICE estimator achieves the parametric rate of convergence in both parameter estimation and model selection. Moreover, ALICE calibrates regularizations when estimating each column of the inverse covariance matrix. So it not only is asymptotically tuning free, but also achieves an improved finite sample performance. We present numerical simulations to support our theory, and a real data example to illustrate the effectiveness of the proposed estimator. Tuo Zhao, Han Liu 0001 |
NIPS | 1 |
| 2013 | CODA: high dimensional copula discriminant analysis
Tuo Zhao, Han Liu 0001 |
J. Mach. Learn. Res. | 2 |
| 2012 | Smooth-projected Neighborhood Pursuit for High-dimensional Nonparanormal Graph EstimationabstractMany statistical methods gain robustness and exibility by sacricing convenient computational structure. In this paper, we illustrate this fundamental tradeoff by studying a semiparametric graphical model estimation problem. We explain how new computational techniques help to solve this type of problem. In particularly, we propose a smooth-projected neighborhood pursuit method for efciently estimating high dimensional nonparanormal graphs with theoretical guarantees. Besides new computational and theoretical analysis, we also provide an alternative view to analyze the tradeoff between computational efciency and statistical error under a smoothing optimization framework. We also report experimental results on text and stock datasets. Tuo Zhao, Kathryn Roeder, Han Liu 0001 |
NIPS | 1 |
| 2012 | The huge Package for High-dimensional Undirected Graph Estimation in R
Tuo Zhao, Han Liu 0001, Kathryn Roeder, John D. Lafferty, Larry A. Wasserman |
J. Mach. Learn. Res. | 1 |
| 2010 | Projected gradient method for kernel discriminant nonnegative matrix factorization and the applications
Zhizheng Liang, Youfu Li 0001, Tuo Zhao |
Signal Process. | 3 |
| 2008 | Interest filter vs. interest operator: Face recognition using Fisher linear discriminant based on interest filter representation
Tuo Zhao, Zhizheng Liang, David Zhang 0001, Quan Zou 0001 |
Pattern Recognit. Lett. | 1 |
| 2006 | Curve Mapping Based Illumination Adjustment for Face Detection
Xiaoyue Jiang, Tuo Zhao, Rongchun Zhao |
ACIVS | 2 |
| 2005 | Re-lighting and Compensation for Face Images
Xiaoyue Jiang, Tuo Zhao, Rongchun Zhao |
CAIP | 2 |