Rui Liu 0013

dblp:42/469-13 · DBLP profile ↗
← Back
8ranked-venue papers
7as first author
3since 2021 · last 2022
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Efficient and distributed learning · 33% Optimization for machine learning · 27% Deep learning architectures and training · 16%
Theoretical computer science
2 papers
Algorithms and data structures · 86% Algorithmic game theory and mechanism design · 9% Graph algorithms and graph theory · 5%

Topics — the 20 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
distributed training
1.122022
Communication-efficient Distributed Learning for Large Batch Optimization · ICML 2022
Gating Dropout: Communication-efficient Regularization for Sparsely Activated Transformers · ICML 2022
Machine learning › Efficient and distributed learning › distributed training
gradient compression
0.612022
Communication-efficient Distributed Learning for Large Batch Optimization · ICML 2022
Machine learning › Optimization for machine learning › large-scale optimization
large batch optimization
0.612022
Communication-efficient Distributed Learning for Large Batch Optimization · ICML 2022
Machine learning › Learning paradigms › continual learning
memory replay
0.612022
Transformer with Memory Replay · AAAI 2022
Machine learning › Deep learning architectures and training
mixture of experts
0.612022
Gating Dropout: Communication-efficient Regularization for Sparsely Activated Transformers · ICML 2022
Machine learning › Efficient and distributed learning › data-efficient learning
sample-efficient training
0.612022
Transformer with Memory Replay · AAAI 2022
Machine learning › Deep learning architectures and training
transformer
0.612022
Transformer with Memory Replay · AAAI 2022
Machine learning › Optimization for machine learning › adaptive optimization
adam
0.412020
Adam with Bandit Sampling for Deep Learning · NeurIPS 2020
Machine learning › Optimization for machine learning › stochastic optimization
adaptive gradient methods
0.412020
Adam with Bandit Sampling for Deep Learning · NeurIPS 2020
Machine learning › Optimization for machine learning
stochastic optimization
0.412020
Adam with Bandit Sampling for Deep Learning · NeurIPS 2020
Algorithms and data structures › learning algorithms
best arm identification
0.412019
A Bandit Approach to Maximum Inner Product Search · AAAI 2019
Algorithms and data structures › similarity search
maximum inner product search
0.412019
A Bandit Approach to Maximum Inner Product Search · AAAI 2019
Algorithms and data structures
similarity search
0.412019
A Bandit Approach to Maximum Inner Product Search · AAAI 2019
Machine learning › Kernel, tree and ensemble methods › ensemble learning › boosting
adaboost
0.312017
An Analysis of Boosted Linear Classifiers on Noisy Data with Applications to Multiple-Instance Learning · ICDM 2017
Machine learning › Kernel, tree and ensemble methods › ensemble learning
boosting
0.312017
An Analysis of Boosted Linear Classifiers on Noisy Data with Applications to Multiple-Instance Learning · ICDM 2017
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
0.312017
An Analysis of Boosted Linear Classifiers on Noisy Data with Applications to Multiple-Instance Learning · ICDM 2017
Machine learning › Learning theory › generalization bounds
margin theory
0.312017
An Analysis of Boosted Linear Classifiers on Noisy Data with Applications to Multiple-Instance Learning · ICDM 2017
Data mining
clustering
0.212015
Robust Multi-Network Clustering via Joint Cross-Domain Cluster Alignment · ICDM 2015
Algorithmic game theory and mechanism design
multi-armed bandit
0.112019
A Bandit Approach to Maximum Inner Product Search · AAAI 2019
Graph algorithms and graph theory
graph clustering
0.112015
Robust Multi-Network Clustering via Joint Cross-Domain Cluster Alignment · ICDM 2015

Methods — techniques the papers use, named apart from their topics

pre-training · 0.6memory replay · 0.6layer-wise adaptive learning rate · 0.6gradient compression · 0.6gating network · 0.6dropout regularization · 0.6convergence analysis · 0.6robust clustering · 0.4joint cluster alignment · 0.4multi-armed bandit · 0.4importance sampling · 0.4monte carlo method · 0.4bandit algorithms · 0.4analytical margin analysis · 0.3
YearPublicationVenuePosition
2022 Transformer with Memory Replay
abstract
Transformers achieve state-of-the-art performance for natural language processing tasks by pre-training on large-scale text corpora. They are extremely compute-intensive and have very high sample complexity. Memory replay is a mechanism that remembers and reuses past examples by saving to and replaying from a memory buffer. It has been successfully used in reinforcement learning and GANs due to better sample efficiency. In this paper, we propose Transformer with Memory Replay, which integrates memory replay with transformer, making transformer more sample efficient. Experiments on GLUE and SQuAD benchmark datasets showed that Transformer with Memory Replay can achieve at least 1% point increase compared to the baseline transformer model when pre-trained with the same number of examples. Further, by adopting a careful design that reduces the wall-clock time overhead of memory replay, we also empirically achieve a better runtime efficiency.
Rui Liu 0013, Barzan Mozafari
AAAI1
2022 Gating Dropout: Communication-efficient Regularization for Sparsely Activated Transformers
abstract
Sparsely activated transformers, such as Mixture of Experts (MoE), have received great interest due to their outrageous scaling capability which enables dramatical increases in model size without significant increases in computational cost. To achieve this, MoE models replace the feedforward sub-layer with Mixture-of-Experts sub-layer in transformers and use a gating network to route each token to its assigned experts. Since the common practice for efficient training of such models requires distributing experts and tokens across different machines, this routing strategy often incurs huge cross-machine communication cost because tokens and their assigned experts likely reside in different machines. In this paper, we propose Gating Dropout, which allows tokens to ignore the gating network and stay at their local machines, thus reducing the cross-machine communication. Similar to traditional dropout, we also show that Gating Dropout has a regularization effect during training, resulting in improved generalization performance. We validate the effectiveness of Gating Dropout on multilingual machine translation tasks. Our results demonstrate that Gating Dropout improves a state-of-the-art MoE model with faster wall-clock time convergence rates and better BLEU scores for a variety of model sizes and datasets.
Rui Liu 0013, Young Jin Kim 0006, Alexandre Muzio, Hany Hassan
ICML1
2022 Communication-efficient Distributed Learning for Large Batch Optimization
abstract
Many communication-efficient methods have been proposed for distributed learning, whereby gradient compression is used to reduce the communication cost. However, given recent advances in large batch optimization (e.g., large batch SGD and its variant LARS with layerwise adaptive learning rates), the compute power of each machine is being fully utilized. This means, in modern distributed learning, the per-machine computation cost is no longer negligible compared to the communication cost. In this paper, we propose new gradient compression methods for large batch optimization, JointSpar and its variant JointSpar-LARS with layerwise adaptive learning rates, that jointly reduce both the computation and the communication cost. To achieve this, we take advantage of the redundancy in the gradient computation, unlike the existing methods compute all coordinates of the gradient vector, even if some coordinates are later dropped for communication efficiency. JointSpar and its variant further reduce the training time by avoiding the wasted computation on dropped coordinates. While computationally more efficient, we prove that JointSpar and its variant also maintain the same convergence rates as their respective baseline methods. Extensive experiments show that, by reducing the time per iteration, our methods converge faster than state-of-the-art compression methods in terms of wall-clock time.
Rui Liu 0013, Barzan Mozafari
ICML1
2020 Adam with Bandit Sampling for Deep Learning
abstract
Adam is a widely used optimization method for training deep learning models. It computes individual adaptive learning rates for different parameters. In this paper, we propose a generalization of Adam, called Adambs, that allows us to also adapt to different training examples based on their importance in the model's convergence. To achieve this, we maintain a distribution over all examples, selecting a mini-batch in each iteration by sampling according to this distribution, which we update using a multi-armed bandit algorithm. This ensures that examples that are more beneficial to the model training are sampled with higher probabilities. We theoretically show that Adambs improves the convergence rate of Adam---$O(\sqrt{\frac{\log n}{T} })$ instead of $O(\sqrt{\frac{n}{T}})$ in some cases. Experiments on various models and datasets demonstrate Adambs's fast convergence in practice.
Rui Liu 0013, Barzan Mozafari
NeurIPS1
2020 Rankboost+: an improvement to Rankboost
Harold S. Connamacher, Nikil Pancha, Rui Liu 0013, Soumya Ray
Mach. Learn.3
2019 A Bandit Approach to Maximum Inner Product Search
abstract
There has been substantial research on sub-linear time approximate algorithms for Maximum Inner Product Search (MIPS). To achieve fast query time, state-of-the-art techniques require significant preprocessing, which can be a burden when the number of subsequent queries is not sufficiently large to amortize the cost. Furthermore, existing methods do not have the ability to directly control the suboptimality of their approximate results with theoretical guarantees. In this paper, we propose the first approximate algorithm for MIPS that does not require any preprocessing, and allows users to control and bound the suboptimality of the results. We cast MIPS as a Best Arm Identification problem, and introduce a new bandit setting that can fully exploit the special structure of MIPS. Our approach outperforms state-of-the-art methods on both synthetic and real-world datasets.
Rui Liu 0013, Barzan Mozafari
AAAI1
2017 An Analysis of Boosted Linear Classifiers on Noisy Data with Applications to Multiple-Instance Learning
abstract
An interesting observation about the well-known AdaBoost algorithm is that, though theory suggests it should overfit when applied to noisy data, experiments indicate it often does not do so in practice. In this paper, we study the behavior of AdaBoost on datasets with one-sided uniform class noise using linear classifiers as the base learner. We show analytically that, under some ideal conditions, this approach will not overfit, and can in fact recover a zero-error concept with respect to the true, uncorrupted instance labels. We also analytically show that AdaBoost increases the margins of predictions over boosting iterations, as has been previously suggested in the literature. We then compare the empirical behavior of AdaBoost using real world datasets with one-sided noise derived from multiple-instance data. Although our assumptions may not hold in a practical setting, our experiments show that standard AdaBoost still performs well, as suggested by our analysis, and often outperforms baseline variations in the literature that explicitly try to account for noise.
Rui Liu 0013, Soumya Ray
ICDM1
2015 Robust Multi-Network Clustering via Joint Cross-Domain Cluster Alignment
abstract
Network clustering is an important problem thathas recently drawn a lot of attentions. Most existing workfocuses on clustering nodes within a single network. In manyapplications, however, there exist multiple related networks, inwhich each network may be constructed from a different domainand instances in one domain may be related to instances in otherdomains. In this paper, we propose a robust algorithm, MCA, formulti-network clustering that takes into account cross-domain relationshipsbetween instances. MCA has several advantages overthe existing single network clustering methods. First, it is ableto detect associations between clusters from different domains, which, however, is not addressed by any existing methods. Second, it achieves more consistent clustering results on multiple networksby leveraging the duality between clustering individual networksand inferring cross-network cluster alignment. Finally, it providesa multi-network clustering solution that is more robust to noiseand errors. We perform extensive experiments on a variety ofreal and synthetic networks to demonstrate the effectiveness andefficiency of MCA.
Rui Liu 0013, Wei Cheng 0002, Hanghang Tong, Wei Wang 0010, Xiang Zhang 0001
ICDM1