EDBT 2026 Demo / reviewers in the wild / expert
Mu Li 0003
dblp:36/4526-3
· DBLP profile ↗
35ranked-venue papers
5as first author
21since 2021 · last 2025
0000-0002-4433-2301ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 1 since 2021Systems, architecture and hardware · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressiveness, and Linguistic Challenges Using Model-as-a-JudgeabstractText-to-Speech (TTS) benchmarks often fail to capture how well models handle nuanced and semantically complex text. Building on $\textit{EmergentTTS}$, we introduce $\textit{EmergentTTS-Eval}$, a comprehensive benchmark covering six challenging TTS scenarios: emotions, paralinguistics, foreign words, syntactic complexity, complex pronunciation (e.g. URLs, formulas), and questions. Crucially, our framework automates both test-case generation and evaluation, making the benchmark easily extensible. Starting from a small set of human-written seed prompts, we iteratively extend them using LLMs to target specific structural, phonetic and prosodic challenges, resulting in 1,645 diverse test samples. Moreover, we employ a model-as-a-judge approach, using a Large Audio Language Model (LALM) to assess the speech across multiple dimensions such as expressed emotion, prosodic, intonational, and pronunciation accuracy. We evaluate state-of-the-art open-source and proprietary TTS systems, such as 11Labs, Deepgram, and OpenAI's 4o-mini-TTS, on EmergentTTS-Eval, demonstrating its ability to reveal fine-grained performance differences. Results show that the model-as-a-judge approach offers robust TTS assessment and a high correlation with human preferences. We open-source the code and the dataset. Ruskin Raj Manku, Yuzhi Tang, Xingjian Shi, Mu Li 0003, Alexander J. Smola |
NeurIPS | 4 |
| 2024 | Online learning from capricious data streams via shared and new feature spaces
Peng Zhou 0008, Mu Li 0003, Yuan-Ting Yan |
Appl. Intell. | 3 |
| 2024 | Improving Semantic Segmentation via Efficient Self-TrainingabstractStarting from the seminal work of Fully Convolutional Networks (FCN), there has been significant progress on semantic segmentation. However, deep learning models often require large amounts of pixelwise annotations to train accurate and robust models. Given the prohibitively expensive annotation cost of segmentation masks, we introduce a self-training framework in this paper to leverage pseudo labels generated from unlabeled data. In order to handle the data imbalance problem of semantic segmentation, we propose a centroid sampling strategy to uniformly select training samples from every class within each epoch. We also introduce a fast training schedule to alleviate the computational burden. This enables us to explore the usage of large amounts of pseudo labels. Our Centroid Sampling based Self-Training framework (CSST) achieves state-of-the-art results on Cityscapes and CamVid datasets. On PASCAL VOC 2012 test set, our models trained with the original train set even outperform the same models trained on the much bigger augmented train set. This indicates the effectiveness of CSST when there are fewer annotations. We also demonstrate promising few-shot generalization capability from Cityscapes to BDD100K and from Cityscapes to Mapillary datasets. Yi Zhu 0001, Chongruo Wu, Zhi Zhang 0005, Tong He 0002, Hang Zhang 0005, R. Manmatha, Mu Li 0003, Alexander J. Smola |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2023 | Tailoring Instructions to Student's Learning Levels Boosts Knowledge DistillationabstractIt has been commonly observed that a teacher model with superior performance does not necessarily result in a stronger student, highlighting a discrepancy between current teacher training practices and effective knowledge transfer.In order to enhance the guidance of the teacher training process, we introduce the concept of distillation influence to determine the impact of distillation from each training sample on the student's generalization ability.In this paper, we propose Learning Good Teacher Matters (LGTM), an efficient training technique for incorporating distillation influence into the teacher's learning process.By prioritizing samples that are likely to enhance the student's generalization ability, our LGTM outperforms 10 common knowledge distillation baselines on 6 text classification tasks in the GLUE benchmark.1 Zihan Zhong, Xingjian Shi, Yi Zhu 0001, Chun Yuan 0003, Mu Li 0003 |
ACL (1) | 6 |
| 2023 | A Cheaper and Better Diffusion Language Model with Soft-Masked NoiseabstractDiffusion models that are based on iterative denoising have been recently proposed and leveraged in various generation tasks like image generation.Whereas, as a way inherently built for continuous data, existing diffusion models still have some limitations in modeling discrete data, e.g., languages.For example, the generally used Gaussian noise can not handle the discrete corruption well, and the objectives in continuous spaces fail to be stable for textual data in the diffusion process especially when the dimension is high.To alleviate these issues, we introduce a novel diffusion model for language modeling, Masked-Diffusion LM, with lower training cost and better performances, inspired by linguistic features in languages.Specifically, we design a linguistic-informed forward process which adds corruptions to the text through strategically soft-masking to better noise the textual data.Also, we directly predict the categorical distribution with cross-entropy loss function in every diffusion step to connect the continuous space and discrete space in a more efficient and straightforward way.Through experiments on 5 controlled generation tasks, we demonstrate that our Masked-Diffusion LM can achieve better generation quality than the state-of-the-art diffusion models with better efficiency.Code is available at https://github. com/SALT-NLP/Masked_Diffusioin_LM. Jiaao Chen, Aston Zhang, Mu Li 0003, Alexander J. Smola, Diyi Yang |
EMNLP | 3 |
| 2023 | Automatic Chain of Thought Prompting in Large Language Models
Zhuosheng Zhang 0001, Aston Zhang, Mu Li 0003, Alexander J. Smola |
ICLR | 3 |
| 2023 | Parameter-Efficient Fine-Tuning Design Spaces
Jiaao Chen, Aston Zhang, Xingjian Shi, Mu Li 0003, Alexander J. Smola, Diyi Yang |
ICLR | 4 |
| 2023 | Learning Multimodal Data Augmentation in Feature Space
Zichang Liu, Zhiqiang Tang 0001, Xingjian Shi, Aston Zhang, Mu Li 0003, Anshumali Shrivastava, Andrew Gordon Wilson |
ICLR | 5 |
| 2023 | AIM: Adapting Image Models for Efficient Video Action Recognition
Taojiannan Yang, Yi Zhu 0001, Yusheng Xie, Aston Zhang, Chen Chen 0001, Mu Li 0003 |
ICLR | 6 |
| 2023 | XTab: Cross-table Pretraining for Tabular TransformersabstractThe success of self-supervised learning in computer vision and natural language processing has motivated pretraining methods on tabular data. However, most existing tabular self-supervised learning models fail to leverage information across multiple data tables and cannot generalize to new tables. In this work, we introduce XTab, a framework for cross-table pretraining of tabular transformers on datasets from various domains. We address the challenge of inconsistent column types and quantities among tables by utilizing independent featurizers and using federated learning to pretrain the shared component. Tested on 84 tabular prediction tasks from the OpenML-AutoML Benchmark (AMLB), we show that (1) XTab consistently boosts the generalizability, learning speed, and performance of multiple tabular transformers, (2) by pretraining FT-Transformer via XTab, we achieve superior performance than other state-of-the-art tabular deep learning models on various tasks such as regression, binary, and multiclass classification. Bingzhao Zhu, Xingjian Shi, Nick Erickson, Mu Li 0003, George Karypis, Mahsa Shoaran |
ICML | 4 |
| 2023 | PreDiff: Precipitation Nowcasting with Latent Diffusion ModelsabstractEarth system forecasting has traditionally relied on complex physical models that are computationally expensive and require significant domain expertise.
In the past decade, the unprecedented increase in spatiotemporal Earth observation data has enabled data-driven forecasting models using deep learning techniques.
These models have shown promise for diverse Earth system forecasting tasks but either struggle with handling uncertainty or neglect domain-specific prior knowledge, resulting in averaging possible futures to blurred forecasts or generating physically implausible predictions.
To address these limitations, we propose a two-stage pipeline for probabilistic spatiotemporal forecasting: 1) We develop *PreDiff*, a conditional latent diffusion model capable of probabilistic forecasts. 2) We incorporate an explicit knowledge alignment mechanism to align forecasts with domain-specific physical constraints.
This is achieved by estimating the deviation from imposed constraints at each denoising step and adjusting the transition distribution accordingly.
We conduct empirical studies on two datasets: N-body MNIST, a synthetic dataset with chaotic behavior, and SEVIR, a real-world precipitation nowcasting dataset.
Specifically, we impose the law of conservation of energy in N-body MNIST and anticipated precipitation intensity in SEVIR.
Experiments demonstrate the effectiveness of PreDiff in handling uncertainty, incorporating domain-specific prior knowledge, and generating forecasts that exhibit high operational utility. Zhihan Gao 0001, Xingjian Shi, Boran Han, Hao Wang 0014, Xiaoyong Jin, Danielle C. Maddix, Yi Zhu 0001, Mu Li 0003, Yuyang Wang 0001 |
NeurIPS | 8 |
| 2023 | Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual RecognitionabstractThis work proposes POMP, a prompt pre-training method for vision-language models. Being memory and computation efficient, POMP enables the learned prompt to condense semantic information for a rich set of visual concepts with over twenty-thousand classes. Once pre-trained, the prompt with a strong transferable ability can be directly plugged into a variety of visual recognition tasks including image classification, semantic segmentation, and object detection, to boost recognition performances in a zero-shot manner. Empirical evaluation shows that POMP achieves state-of-the-art performances on 21 datasets, e.g., 67.0% average accuracy on 10 classification datasets (+3.1% compared to CoOp) and 84.4 hIoU on open-vocabulary Pascal VOC segmentation (+6.9 compared to ZSSeg). Shuhuai Ren, Aston Zhang, Yi Zhu 0001, Shuai Zheng 0004, Mu Li 0003, Alexander J. Smola, Xu Sun 0001 |
NeurIPS | 6 |
| 2023 | What Makes for Good Tokenizers in Vision Transformer?abstractThe architecture of transformers, which recently witness booming applications in vision tasks, has pivoted against the widespread convolutional paradigm. Relying on the tokenization process that splits inputs into multiple tokens, transformers are capable of extracting their pairwise relationships using self-attention. While being the stemming building block of transformers, what makes for a good tokenizer has not been well understood in computer vision. In this work, we investigate this uncharted problem from an information trade-off perspective. In addition to unifying and understanding existing structural modifications, our derivation leads to better design strategies for vision tokenizers. The proposed Modulation across Tokens (MoTo) incorporates inter-token modeling capability through normalization. Furthermore, a regularization objective TokenProp is embraced in the standard training regime. Through extensive experiments on various transformer architectures, we observe both improved performance and intriguing properties of these two plug-and-play designs with negligible computational overhead. These observations further indicate the importance of the commonly-omitted designs of tokenizers in vision transformer. Shengju Qian, Yi Zhu 0001, Wenbo Li 0002, Mu Li 0003, Jiaya Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Removing Batch Normalization Boosts Adversarial TrainingabstractAdversarial training (AT) defends deep neural networks against adversarial attacks. One challenge that limits its practical application is the performance degradation on clean samples. A major bottleneck identified by previous works is the widely used batch normalization (BN), which struggles to model the different statistics of clean and adversarial training samples in AT. Although the dominant approach is to extend BN to capture this mixture of distribution, we propose to completely eliminate this bottleneck by removing all BN layers in AT. Our normalizer-free robust training (NoFrost) method extends recent advances in normalizer-free networks to AT for its unexplored advantage on handling the mixture distribution challenge. We show that NoFrost achieves adversarial robustness with only a minor sacrifice on clean sample accuracy. On ImageNet with ResNet50, NoFrost achieves $74.06%$ clean accuracy, which drops merely $2.00%$ from standard training. In contrast, BN-based AT obtains $59.28%$ clean accuracy, suffering a significant $16.78%$ drop from standard training. In addition, NoFrost achieves a $23.56%$ adversarial robustness against PGD attack, which improves the $13.57%$ robustness in BN-based AT. We observe better model smoothness and larger decision margins from NoFrost, which make the models less sensitive to input perturbations and thus more robust. Moreover, when incorporating more data augmentations into NoFrost, it achieves comprehensive robustness against multiple distribution shifts. Code and pre-trained models are public at https://github.com/amazon-research/normalizer-free-robust-training. Haotao Wang, Aston Zhang, Shuai Zheng 0004, Xingjian Shi, Mu Li 0003, Zhangyang Wang |
ICML | 5 |
| 2022 | Partial and Asymmetric Contrastive Learning for Out-of-Distribution Detection in Long-Tailed RecognitionabstractExisting out-of-distribution (OOD) detection methods are typically benchmarked on training sets with balanced class distributions. However, in real-world applications, it is common for the training sets to have long-tailed distributions. In this work, we first demonstrate that existing OOD detection methods commonly suffer from significant performance degradation when the training set is long-tail distributed. Through analysis, we posit that this is because the models struggle to distinguish the minority tail-class in-distribution samples, from the true OOD samples, making the tail classes more prone to be falsely detected as OOD. To solve this problem, we propose Partial and Asymmetric Supervised Contrastive Learning (PASCL), which explicitly encourages the model to distinguish between tail-class in-distribution samples and OOD samples. To further boost in-distribution classification accuracy, we propose Auxiliary Branch Finetuning, which uses two separate branches of BN and classification layers for anomaly detection and in-distribution classification, respectively. The intuition is that in-distribution and OOD anomaly data have different underlying distributions. Our method outperforms previous state-of-the-art method by $1.29%$, $1.45%$, $0.69%$ anomaly detection false positive rate (FPR) and $3.24%$, $4.06%$, $7.89%$ in-distribution classification accuracy on CIFAR10-LT, CIFAR100-LT, and ImageNet-LT, respectively. Code and pre-trained models are available at https://github.com/amazon-research/long-tailed-ood-detection. Haotao Wang, Aston Zhang, Yi Zhu 0001, Shuai Zheng 0004, Mu Li 0003, Alexander J. Smola, Zhangyang Wang |
ICML | 5 |
| 2022 | Earthformer: Exploring Space-Time Transformers for Earth System ForecastingabstractConventionally, Earth system (e.g., weather and climate) forecasting relies on numerical simulation with complex physical models and hence is both expensive in computation and demanding on domain expertise. With the explosive growth of spatiotemporal Earth observation data in the past decade, data-driven models that apply Deep Learning (DL) are demonstrating impressive potential for various Earth system forecasting tasks. The Transformer as an emerging DL architecture, despite its broad success in other domains, has limited adoption in this area. In this paper, we propose Earthformer, a space-time Transformer for Earth system forecasting. Earthformer is based on a generic, flexible and efficient space-time attention block, named Cuboid Attention. The idea is to decompose the data into cuboids and apply cuboid-level self-attention in parallel. These cuboids are further connected with a collection of global vectors. We conduct experiments on the MovingMNIST dataset and a newly proposed chaotic $N$-body MNIST dataset to verify the effectiveness of cuboid attention and figure out the best design of Earthformer. Experiments on two real-world benchmarks about precipitation nowcasting and El Niño/Southern Oscillation (ENSO) forecasting show that Earthformer achieves state-of-the-art performance. Zhihan Gao 0001, Xingjian Shi, Hao Wang 0014, Yi Zhu 0001, Yuyang Wang 0001, Mu Li 0003, Dit-Yan Yeung |
NeurIPS | 6 |
| 2022 | MiCS: Near-linear Scaling for Training Gigantic Model on Public CloudabstractExisting general purpose frameworks for gigantic model training, i.e., dense models with billions of parameters, cannot scale efficiently on cloud environment with various networking conditions due to large communication overheads. In this paper, we propose MiCS, which Minimizes the Communication Scale to bring down communication overhead. Specifically, by decreasing the number of participants in a communication collective, MiCS can utilize heterogeneous network bandwidth, reduce network traffic over slower links, reduce the latency of communications for maintaining high network bandwidth utilization, and amortize expensive global gradient synchronization overhead. Our evaluation on AWS shows that the system throughput of MiCS is up to 2.89× that of the state-of-the-art large model training systems. MiCS achieves near-linear scaling efficiency, which is up to 1.27× that of DeepSpeed. MiCS allows us to train a proprietary model with 100 billion parameters on 512 GPUs with 99.4% weak-scaling efficiency, and it is able to saturate over 54.5% theoretical computation power of each GPU on a public cloud with less GPU memory and more restricted networks than DGX-A100 clusters. Zhen Zhang 0063, Shuai Zheng 0004, Yida Wang 0003, Justin Chiu, George Karypis, Trishul Chilimbi, Mu Li 0003, Xin Jin 0008 |
Proc. VLDB Endow. | 7 |
| 2021 | Lorien: Efficient Deep Learning Workloads DeliveryabstractModern deep learning systems embrace the compilation idea to self generate code of a deep learning model to catch up the rapidly changed deep learning operators and newly emerged hardware platforms. The performance of the self-generated code is guaranteed via auto-tuning frameworks which normally take a long time to find proper execution schedules for the given operators, which hurts both user experiences and time-to-the-market in terms of model developments and deployments. Cody Hao Yu, Xingjian Shi, Haichen Shen, Zhi Chen 0030, Mu Li 0003, Yida Wang 0003 |
SoCC | 5 |
| 2021 | CrossNorm and SelfNorm for Generalization under Distribution ShiftsabstractTraditional normalization techniques (e.g., Batch Normalization and Instance Normalization) generally and simplistically assume that training and test data follow the same distribution. As distribution shifts are inevitable in real-world applications, well-trained models with previous normalization methods can perform badly in new environments. Can we develop new normalization methods to improve generalization robustness under distribution shifts? In this paper, we answer the question by proposing Cross-Norm and SelfNorm. CrossNorm exchanges channel-wise mean and variance between feature maps to enlarge training distribution, while SelfNorm uses attention to recalibrate the statistics to bridge gaps between training and test distributions. CrossNorm and SelfNorm can complement each other, though exploring different directions in statistics usage. Extensive experiments on different fields (vision and language), tasks (classification and segmentation), settings (supervised and semi-supervised), and distribution shift types (synthetic and natural) show the effectiveness. Code is available at https://github.com/amazon-research/crossnorm-selfnorm Zhiqiang Tang 0001, Yunhe Gao, Yi Zhu 0001, Zhi Zhang 0005, Mu Li 0003, Dimitris N. Metaxas |
ICCV | 5 |
| 2021 | Blending Anti-Aliasing into Vision TransformerabstractThe transformer architectures, based on self-attention mechanism and convolution-free design, recently found superior performance and booming applications in computer vision. However, the discontinuous patch-wise tokenization process implicitly introduces jagged artifacts into attention maps, arising the traditional problem of aliasing for vision transformers. Aliasing effect occurs when discrete patterns are used to produce high frequency or continuous information, resulting in the indistinguishable distortions. Recent researches have found that modern convolution networks still suffer from this phenomenon. In this work, we analyze the uncharted problem of aliasing in vision transformer and explore to incorporate anti-aliasing properties. Specifically, we propose a plug-and-play Aliasing-Reduction Module (ARM) to alleviate the aforementioned issue. We investigate the effectiveness and generalization of the proposed method across multiple tasks and various vision transformer families. This lightweight design consistently attains a clear boost over several famous structures. Furthermore, our module also improves data efficiency and robustness of vision transformers. Shengju Qian, Hao Shao, Yi Zhu 0001, Mu Li 0003, Jiaya Jia |
NeurIPS | 4 |
| 2021 | Progressive Coordinate Transforms for Monocular 3D Object DetectionabstractRecognizing and localizing objects in the 3D space is a crucial ability for an AI agent to perceive its surrounding environment. While significant progress has been achieved with expensive LiDAR point clouds, it poses a great challenge for 3D object detection given only a monocular image. While there exist different alternatives for tackling this problem, it is found that they are either equipped with heavy networks to fuse RGB and depth information or empirically ineffective to process millions of pseudo-LiDAR points. With in-depth examination, we realize that these limitations are rooted in inaccurate object localization. In this paper, we propose a novel and lightweight approach, dubbed {\em Progressive Coordinate Transforms} (PCT) to facilitate learning coordinate representations. Specifically, a localization boosting mechanism with confidence-aware loss is introduced to progressively refine the localization prediction. In addition, semantic image representation is also exploited to compensate for the usage of patch proposals. Despite being lightweight and simple, our strategy allows us to establish a new state-of-the-art among the monocular 3D detectors on the competitive KITTI benchmark. At the same time, our proposed PCT shows great generalization to most coordinate-based 3D detection frameworks. Li Wang 0033, Li Zhang 0040, Yi Zhu 0001, Zhi Zhang 0005, Tong He 0002, Mu Li 0003, Xiangyang Xue 0001 |
NeurIPS | 6 |
| 2020 | CSER: Communication-efficient SGD with Error ResetabstractThe scalability of Distributed Stochastic Gradient Descent (SGD) is today limited by communication bottlenecks. We propose a novel SGD variant: \underline{C}ommunication-efficient \underline{S}GD with \underline{E}rror \underline{R}eset, or \underline{CSER}. The key idea in CSER is first a new technique called ``error reset'' that adapts arbitrary compressors for SGD, producing bifurcated local models with periodic reset of resulting local residual errors. Second we introduce partial synchronization for both the gradients and the models, leveraging advantages from them. We prove the convergence of CSER for smooth non-convex problems. Empirical results show that when combined with highly aggressive compressors, the CSER algorithms accelerate the distributed training by nearly $10\times$ for CIFAR-100, and by $4.5\times$ for ImageNet. Shuai Zheng 0004, Oluwasanmi Koyejo, Indranil Gupta, Mu Li 0003, Haibin Lin |
NeurIPS | 5 |
| 2020 | FeatGraph: a flexible and efficient backend for graph neural network systemsabstractGraph neural networks (GNNs) are gaining popularity as a promising approach to machine learning on graphs. Unlike traditional graph workloads where each vertex/edge is associated with a scalar, GNNs attach a feature tensor to each vertex/edge. This additional feature dimension, along with consequently more complex vertex- and edge-wise computations, has enormous implications on locality and parallelism, which existing graph processing systems fail to exploit. This paper proposes FeatGraph to accelerate GNN workloads by co-optimizing graph traversal and feature dimension computation. FeatGraph provides a flexible programming interface to express diverse GNN models by composing coarse-grained sparse templates with fine-grained user-defined functions (UDFs) on each vertex/edge. FeatGraph incorporates optimizations for graph traversal into the sparse templates and allows users to specify optimizations for UDFs with a feature dimension schedule (FDS). FeatGraph speeds up end-to-end GNN training and inference by up to 32× on CPU and 7× on GPU. Zihao Ye 0001, Da Zheng 0004, Mu Li 0003, Zheng Zhang 0001, Zhiru Zhang, Yida Wang 0003 |
SC | 6 |
| 2020 | GluonCV and GluonNLP: Deep Learning in Computer Vision and Natural Language ProcessingabstractWe present GluonCV and GluonNLP, the deep learning toolkits for computer vision and natural language processing based on Apache MXNet (incubating). These toolkits provide state-of-the-art pre-trained models, training scripts, and training logs, to facilitate rapid prototyping and promote reproducible research. We also provide modular APIs with flexible building blocks to enable efficient customization. Leveraging the MXNet ecosystem, the deep learning models in GluonCV and GluonNLP can be deployed onto a variety of platforms with different programming languages. The Apache 2.0 license has been adopted by GluonCV and GluonNLP to allow for software distribution, modification, and usage. He He 0001, Tong He 0002, Leonard Lausen, Mu Li 0003, Haibin Lin, Xingjian Shi, Chenguang Wang 0001, Junyuan Xie, Sheng Zha, Aston Zhang, Hang Zhang 0005, Zhi Zhang 0005, Shuai Zheng 0004, Yi Zhu 0001 |
J. Mach. Learn. Res. | 5 |
| 2019 | Bag of Tricks for Image Classification with Convolutional Neural NetworksabstractMuch of the recent progress made in image classification research can be credited to training procedure refinements, such as changes in data augmentations and optimization methods. In the literature, however, most refinements are either briefly mentioned as implementation details or only visible in source code. In this paper, we will examine a collection of such refinements and empirically evaluate their impact on the final model accuracy through ablation study. We will show that, by combining these refinements together, we are able to improve various CNN models significantly. For example, we raise ResNet-50's top-1 validation accuracy from 75.3% to 79.29% on ImageNet. We will also demonstrate that improvement on image classification accuracy leads to better transfer learning performance in other application domains such as object detection and semantic segmentation. Tong He 0002, Zhi Zhang 0005, Hang Zhang 0005, Junyuan Xie, Mu Li 0003 |
CVPR | 6 |
| 2019 | A Unified Optimization Approach for CNN Model Inference on Integrated GPUsabstractModern deep learning applications urge to push the model inference taking place at the edge devices for multiple reasons such as achieving shorter latency, relieving the burden of the network connecting to the cloud, and protecting user privacy. The Convolutional Neural Network (CNN) is one of the most widely used model family in the applications. Given the high computational complexity of the CNN models, it is favorable to execute them on the integrated GPUs at the edge devices, which are ubiquitous and have more power and better energy efficiency than the accompanying CPUs. However, programming on integrated GPUs efficiently is challenging due to the variety of their architectures and programming interfaces. This paper proposes an end-to-end solution to execute CNN model inference on the integrated GPUs at the edge, which uses a unified IR to represent and optimize vision-specific operators on integrated GPUs from multiple vendors, as well as leverages machine learning-based scheduling search schemes to optimize computationally-intensive operators like convolution. Our solution even provides a fallback mechanism for operators not suitable or convenient to run on GPUs. The evaluation results suggest that compared to state-of-the-art solutions backed up by the vendor-provided high-performance libraries on Intel Graphics, ARM Mali GPU, and Nvidia integrated Maxwell GPU, our solution achieves similar, or even better (up to 1.62×), performance on a number of popular image classification and object detection models. In addition, our solution has a wider model coverage and is more flexible to embrace new models. Our solution has been adopted in production services in AWS and is open-sourced. Leyuan Wang, Zhi Chen 0030, Lianmin Zheng, Mu Li 0003, Yida Wang 0003 |
ICPP | 6 |
| 2019 | Optimizing CNN Model Inference on CPUs
Ruofei Yu, Mu Li 0003, Vin Sharma, Yida Wang 0003 |
USENIX ATC | 4 |
| 2017 | Data Driven Resource Allocation for Distributed LearningabstractIn distributed machine learning, data is dispatched to multiple machines for processing. Motivated by the fact that similar data points often belong to the same or similar classes, and more generally, classification rules of high accuracy tend to be “locally simple but globally complex” (Vapnik and Bottou, 1993), we propose data dependent dispatching that takes advantage of such structure. We present an in-depth analysis of this model, providing new algorithms with provable worst-case guarantees, analysis proving existing scalable heuristics perform well in natural non worst-case conditions, and techniques for extending a dispatching rule from a small sample to the entire distribution. We overcome novel technical challenges to satisfy important conditions for accurate distributed learning, including fault tolerance and balancedness. We empirically compare our approach with baselines based on random partitioning, balanced partition trees, and locality sensitive hashing, showing that we achieve significantly higher accuracy on both synthetic and real world image and advertising datasets. We also demonstrate that our technique strongly scales with the available computing power. Travis Dick, Mu Li 0003, Venkata Pillutla, Colin White, Maria-Florina Balcan, Alexander J. Smola |
AISTATS | 2 |
| 2016 | AdaDelay: Delay Adaptive Distributed Stochastic OptimizationabstractWe develop distributed stochastic convex optimization algorithms under a delayed gradient model in which server nodes update parameters and worker nodes compute stochastic (sub)gradients. Our setup is motivated by the behavior of real-world distributed computation systems; in particular, we analyze a setting wherein worker nodes can be differently slow at different times. In contrast to existing approaches, we do not impose a worst-case bound on the delays experienced but rather allow the updates to be sensitive to the actual delays experienced. This sensitivity allows use of larger stepsizes, which can help speed up initial convergence without having to wait too long for slower machines; the global convergence rate is still preserved. We experiment with different delay patterns, and obtain noticeable improvements for large-scale real datasets with billions of examples and features. Suvrit Sra, Adams Wei Yu, Mu Li 0003, Alexander J. Smola |
AISTATS | 3 |
| 2016 | DiFacto: Distributed Factorization MachinesabstractFactorization Machines offer good performance and useful embeddings of data. However, they are costly to scale to large amounts of data and large numbers of features. In this paper we describe DiFacto, which uses a refined Factorization Machine model with sparse memory adaptive constraints and frequency adaptive regularization. We show how to distribute DiFacto over multiple machines using the Parameter Server framework by computing distributed subgradients on minibatches asynchronously. We analyze its convergence and demonstrate its efficiency in computational advertising datasets with billions examples and features. Mu Li 0003, Alexander J. Smola, Yu-Xiang Wang 0003 |
WSDM | 1 |
| 2015 | Cuckoo Linear AlgebraabstractIn this paper we present a novel data structure for sparse vectors based on Cuckoo hashing. It is highly memory efficient and allows for random access at near dense vector level rates. This allows us to solve sparse l1 programming problems exactly and without preprocessing at a cost that is identical to dense linear algebra both in terms of memory and speed. Our approach provides a feasible alternative to the hash kernel and it excels whenever exact solutions are required, such as for feature selection. Li Zhou 0006, David G. Andersen, Mu Li 0003, Alexander J. Smola |
KDD | 3 |
| 2015 | Inferring Movement Trajectories from GPS SnippetsabstractInferring movement trajectories can be a challenging task, in particular when detailed tracking information is not available due to privacy and data collection constraints. In this paper we present a complete and computationally tractable model for estimating and predicting trajectories based on sparsely sampled, anonymous GPS land-marks that we call GPS snippets. To combat data sparsity we use mapping data as side information to constrain the inference process. We show the efficacy of our approach on a set of prediction tasks over data collected from different cities in the US. Mu Li 0003, Amr Ahmed 0001, Alexander J. Smola |
WSDM | 1 |
| 2014 | Efficient mini-batch training for stochastic optimizationabstractStochastic gradient descent (SGD) is a popular technique for large-scale optimization problems in machine learning. In order to parallelize SGD, minibatch training needs to be employed to reduce the communication cost. However, an increase in minibatch size typically decreases the rate of convergence. This paper introduces a technique based on approximate optimization of a conservatively regularized objective function within each minibatch. We prove that the convergence rate does not decrease with increasing minibatch size. Experiments demonstrate that with suitable implementations of approximate optimization, the resulting algorithm can outperform standard SGD in many scenarios. Mu Li 0003, Tong Zhang 0001, Yuqiang Chen, Alexander J. Smola |
KDD | 1 |
| 2014 | Communication Efficient Distributed Machine Learning with the Parameter Server
Mu Li 0003, David G. Andersen, Alexander J. Smola |
NIPS | 1 |
| 2014 | Scaling Distributed Machine Learning with the Parameter Server
Mu Li 0003, David G. Andersen, Jun Woo Park, Alexander J. Smola, Amr Ahmed 0001, Vanja Josifovski, Eugene J. Shekita, Bor-Yiing Su |
OSDI | 1 |