VLDB 2026 Research / reviewers in the wild / expert
Zheng Xu 0002
dblp:83/2535-2
· DBLP profile ↗
40ranked-venue papers
10as first author
20since 2021 · last 2026
0009-0003-6747-3953ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 35 · 8 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 8 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How to DP-Fy Your Data: A Practical Guide to Generating Synthetic Data With Differential PrivacyabstractHigh quality data is of vital importance for unlocking the full potential of AI for end users. Villalobos et al. stated in 2024 that finding new sources of such data is getting harder as most publicly-available human generated data will soon have been used. Additionally, publicly available data often is not representative of users of a particular system — for example, a research speech dataset of contractors interacting with an AI assistant will likely be more homogeneous, well articulated and self-censored that real world commands that end users will issue. Therefore unlocking high-quality data grounded in real user interactions is of vital interest to both system creators and end users themselves. However, the direct use of user data comes with significant privacy risks, which must be addressed before the data can be used. Differential Privacy (DP) is a well established framework for reasoning about and limiting information leakage, and is a gold standard for protecting user privacy. The focus of this work, Differentially Private Synthetic data, refers to synthetic data that preserves the overall trends of source data (often user-generated), while providing strong privacy guarantees to individuals that contributed to the source dataset. DP synthetic data can unlock the value of datasets that have previously been inaccessible due to privacy concerns. Additionally, DP synthetic data can replace the use of sensitive datasets that previously have only had rudimentary protections like ad-hoc rule-based anonymization. In this survey we explore the full suite of techniques surrounding DP synthetic data, the types of privacy protections different generation approaches can offer, and the state-of-the-art for various modalities including image, tabular, text and federated (decentralized) data. We outline all the components needed in a system that generates DP synthetic data, from sensitive data handling and preparation, to tracking the use of synthetic data and empirical privacy testing. We hope that work will result in increased adoption of DP synthetic data, spur additional research in still underexplored domains, and additionally increase trust in DP synthetic data approaches. Natalia Ponomareva 0001, Zheng Xu 0002, H. Brendan McMahan, Peter Kairouz, Lucas Rosenblatt, Vincent Cohen-Addad, Cristóbal Guzmán, Ryan McKenna, Galen Andrew, Alex Bie, Alexey Kurakin, Morteza Zadimoghaddam, Sergei Vassilvitskii, Andreas Terzis |
J. Artif. Intell. Res. | 2 |
| 2025 | InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMsabstractVisual programming has the potential of providing novice programmers with a low-code experience to build customized processing pipelines.Existing systems typically require users to build pipelines from scratch, implying that novice users are expected to set up and link appropriate nodes from a blank workspace.In this paper, we introduce InstructPipe, an AI assistant for prototyping machine learning (ML) pipelines with text instructions.We contribute two large language model (LLM) modules and a code interpreter as part of our framework.The LLM modules generate pseudocode for a target pipeline, and the interpreter renders the pipeline in the node-graph editor for further human-AI collaboration.Both technical and user evaluation (N=16) shows that InstructPipe empowers users to streamline their ML pipeline workfow, reduce their learning curve, and leverage open-ended commands to spark innovative ideas. Zhongyi Zhou, Vrushank Phadnis, Xiuxiu Yuan, Xun Qian, Kristen Wright, Mark Sherwood, Jason Mayes, Yiyi Huang, Zheng Xu 0002, Yinda Zhang 0001, Johnny Lee, Alex Olwal, David Kim 0002, Ram Iyengar, Na Li 0034, Ruofei Du |
CHI | 12 |
| 2025 | Debiasing Federated Learning with Correlated Client ParticipationabstractIn cross-device federated learning (FL) with millions of mobile clients, only a small subset of clients participate in training in every communication round, and Federated Averaging (FedAvg) is the most popular algorithm in practice. Existing analyses of FedAvg usually assume the participating clients are independently sampled in each round from a uniform distribution, which does not reflect real-world scenarios. This paper introduces a theoretical framework that models client participation in FL as a Markov chain to study optimization convergence when clients have non-uniform and correlated participation across rounds.
We apply this framework to analyze a more practical pattern: every client must wait a minimum number of $R$ rounds (minimum separation) before re-participating. We theoretically prove and empirically observe that increasing minimum separation reduces the bias induced by intrinsic non-uniformity of client availability in cross-device FL systems.
Furthermore, we develop an effective debiasing algorithm for FedAvg that provably converges to the unbiased optimal solution under arbitrary minimum separation and unknown client availability distribution. Zheng Xu 0002, Gauri Joshi, Pranay Sharma, Ermin Wei |
ICLR | 3 |
| 2025 | FedKDD 2025: The 2025 International Joint Workshop on Federated Learning for Data Mining and Graph AnalyticsabstractDeep Learning has facilitated various high-stakes applications such as crime detection, urban planning, drug discovery, and healthcare. Its continuous success hinges on learning from massive data in miscellaneous sources, ranging from data with independent distributions to graph-structured data capturing intricate inter-sample relationships. Scaling up the data access requires global collaboration from distributed data owners. Yet, centralizing all data sources to an untrustworthy centralized server will put users' data at risk of privacy leakage or regulation violation. Federated Learning (FL) is a de facto decentralized learning framework that enables knowledge aggregation from distributed users without exposing private data. Though promising advances are witnessed for FL, new challenges are emerging when integrating FL with the rising needs and opportunities in data mining, graph analytics, foundation models, generative AI, and new interdisciplinary applications in science. By hosting this workshop, we aim to attract a broad range of audiences, including researchers and practitioners from academia and industry interested in the emergent challenges in FL. As an effort to advance the fundamental development of FL, this workshop will encourage ideas exchange on the trustworthiness, scalability, and robustness of distributed data mining and graph analytics and their emergent challenges. Carl Yang 0001, Guancheng Wan, Zhuangdi Zhu, Zheng Xu 0002, Junyuan Hong, Nathalie Baracaldo, Neil Shah, Amir Salman Avestimehr |
KDD (2) | 4 |
| 2025 | SPARTA: An Optimization Framework for Differentially Private Sparse Fine-TuningabstractKDD ’25, Toronto, ON, Canada Mehdi Makni, Kayhan Behdin, Gabriel Afriat, Zheng Xu 0002, Sergei Vassilvitskii, Natalia Ponomareva 0001, Rahul Mazumder, Hussein Hazimeh 0001 |
KDD (2) | 4 |
| 2024 | Heterogeneous LoRA for Federated Fine-tuning of On-Device Foundation ModelsabstractFoundation models (FMs) adapt surprisingly well to downstream tasks with fine-tuning.However, their colossal parameter space prohibits their training on resource-constrained edge-devices.For federated fine-tuning, we need to consider the smaller FMs of few billion parameters at most, namely on-device FMs (ODFMs), which can be deployed ondevice.Federated fine-tuning of ODFMs has unique challenges non-present in standard fine-tuning: i) ODFMs poorly generalize to downstream tasks due to their limited sizes making proper fine-tuning imperative to their performance, and ii) devices have limited and heterogeneous system capabilities and data that can deter the performance of fine-tuning.Tackling these challenges, we propose HET-LORA, a feasible and effective federated finetuning method for ODFMs that leverages the system and data heterogeneity at the edge.HETLORA allows heterogeneous LoRA ranks across clients for their individual system resources, and efficiently aggregates and distributes these LoRA modules in a data-aware manner by applying rank self-pruning locally and sparsity-weighted aggregation at the server.It combines the advantages of high and low-rank LoRAs, achieving improved convergence speed and final performance compared to homogeneous LoRA.Furthermore, HET-LORA has enhanced computation and communication efficiency compared to full finetuning making it more feasible for the edge. Yae Jee Cho, Zheng Xu 0002, Aldi Fahrezi, Gauri Joshi |
EMNLP | 3 |
| 2024 | User Inference Attacks on Large Language ModelsabstractText written by humans makes up the vast majority of the data used to pre-train and finetune large language models (LLMs).Many sources of this data-like code, forum posts, personal websites, and books-are easily attributed to one or a few "users".In this paper, we ask if it is possible to infer if any of a user's data was used to train an LLM.Not only would this constitute a breach of privacy, but it would also enable users to detect when their data was used for training.We develop the first effective attacks for user inferenceat times, with near-perfect success-against LLMs.Our attacks are easy to employ, requiring only black-box access to an LLM and a few samples from the user, which need not be the ones that were trained on.We find, both theoretically and empirically, that certain properties make users more susceptible to user inference: being an outlier, having highly correlated examples, and contributing a larger fraction of data.Based on these findings, we identify several methods for mitigating user inference including training with example-level differential privacy, removing within-user duplicate examples, and reducing a user's contribution to the training data.Though these provide partial mitigation, our work highlights the need to develop methods to fully protect LLMs from user inference.Pre-trained LLM Finetuned Nikhil Kandpal, Krishna Pillutla, Alina Oprea, Peter Kairouz, Christopher A. Choquette-Choo, Zheng Xu 0002 |
EMNLP | 6 |
| 2024 | Improved Communication-Privacy Trade-offs in L2 Mean Estimation under Streaming Differential PrivacyabstractWe study $L_2$ mean estimation under central differential privacy and communication constraints, and address two key challenges: firstly, existing mean estimation schemes that simultaneously handle both constraints are usually optimized for $L_\infty$ geometry and rely on random rotation or Kashin’s representation to adapt to $L_2$ geometry, resulting in suboptimal leading constants in mean square errors (MSEs); secondly, schemes achieving order-optimal communication-privacy trade-offs do not extend seamlessly to streaming differential privacy (DP) settings (e.g., tree aggregation or matrix factorization), rendering them incompatible with DP-FTRL type optimizers. In this work, we tackle these issues by introducing a novel privacy accounting method for the sparsified Gaussian mechanism that incorporates the randomness inherent in sparsification into the DP noise. Unlike previous approaches, our accounting algorithm directly operates in $L_2$ geometry, yielding MSEs that fast converge to those of the uncompressed Gaussian mechanism. Additionally, we extend the sparsification scheme to the matrix factorization framework under streaming DP and provide a precise accountant tailored for DP-FTRL type optimizers. Empirically, our method demonstrates at least a 100x improvement of compression for DP-SGD across various FL tasks. Wei-Ning Chen, Berivan Isik, Peter Kairouz, Albert No, Sewoong Oh, Zheng Xu 0002 |
ICML | 6 |
| 2024 | Privacy-Preserving Instructions for Aligning Large Language ModelsabstractService providers of large language model (LLM) applications collect user instructions in the wild and use them in further aligning LLMs with users' intentions. These instructions, which potentially contain sensitive information, are annotated by human workers in the process. This poses a new privacy risk not addressed by the typical private optimization. To this end, we propose using synthetic instructions to replace real instructions in data annotation and model fine-tuning. Formal differential privacy is guaranteed by generating those synthetic instructions using privately fine-tuned generators. Crucial in achieving the desired utility is our novel filtering algorithm that matches the distribution of the synthetic instructions to that of the real ones. In both supervised fine-tuning and reinforcement learning from human feedback, our extensive experiments demonstrate the high utility of the final set of synthetic instructions by showing comparable results to real instructions. In supervised fine-tuning, models trained with private synthetic instructions outperform leading open-source models such as Vicuna. Peter Kairouz, Sewoong Oh, Zheng Xu 0002 |
ICML | 4 |
| 2024 | FedKDD: International Joint Workshop on Federated Learning for Data Mining and Graph AnalyticsabstractDeep Learning has facilitated various high-stakes applications such as crime detection, urban planning, drug discovery, and healthcare. Its continuous success hinges on learning from massive data in miscellaneous sources, ranging from data with independent distributions to graph-structured data capturing intricate inter-sample relationships. Scaling up the data access requires global collaboration from distributed data owners. Yet, centralizing all data sources to an untrustworthy centralized server will put users' data at risk of privacy leakage or regulation violation. Federated Learning (FL) is a de facto decentralized learning framework that enables knowledge aggregation from distributed users without exposing private data. Though promising advances are witnessed for FL, new challenges are emerging when integrating FL with the rising needs and opportunities in data mining, graph analytics, foundation models, generative AI, and new interdisciplinary applications in science. By hosting this workshop, we aim to attract a broad range of audiences, including researchers and practitioners from academia and industry interested in the emergent challenges in FL. As an effort to advance the fundamental development of FL, this workshop will encourage ideas exchange on the trustworthiness, scalability, and robustness of distributed data mining and graph analytics and their emergent challenges. Junyuan Hong, Carl Yang 0001, Zhuangdi Zhu, Zheng Xu 0002, Nathalie Baracaldo, Neil Shah, Amir Salman Avestimehr |
KDD | 4 |
| 2024 | FedGKD: Toward Heterogeneous Federated Learning via Global Knowledge DistillationabstractFederated learning, as one enabling technology of edge intelligence, has gained substantial attention due to its efficacy in training deep learning models without data privacy and network bandwidth concerns. However, due to the heterogeneity of the edge computing system and data, many methods suffer from the“client-drift”issue that could considerably impede the convergence of global model training: local models on clients can drift apart, and the aggregated model can be different from the global optimum. To tackle this issue, one intuitive idea is to guide the local model training by global teachers,i.e.,past global models, where each client learns the global knowledge from past global models via adaptive knowledge distillation techniques. Inspired by these insights, we propose a novel approach for heterogeneous federated learning,FedGKD, which fuses the knowledge from historical global models and guides local training to alleviate the“client-drift”issue. In this paper, we evaluateFedGKDthrough extensive experiments across various CV and NLP datasets (i.e.,CIFAR-10/100, Tiny-ImageNet, AG News, SST5) under different heterogeneous settings. The proposed method is guaranteed to converge under common assumptions and outperforms the state-of-the-art baselines in the non-IID federated setting. Dezhong Yao 0002, Wanning Pan, Yutong Dai 0002, Yao Wan 0001, Xiaofeng Ding 0001, Chen Yu 0003, Hai Jin 0001, Zheng Xu 0002, Lichao Sun 0001 |
IEEE Trans. Computers | 8 |
| 2023 | Learning to Generate Image Embeddings with User-Level Differential PrivacyabstractSmall on-device models have been successfully trained with user-level differential privacy (DP) for next word prediction and image classification tasks in the past. However, existing methods can fail when directly applied to learn embedding models using supervised training data with a large class space. To achieve user-level DP for large image-to-embedding feature extractors, we propose DP-FedEmb, a variant of federated learning algorithms with per-user sensitivity control and noise addition, to train from user-partitioned data centralized in the datacenter. DP-FedEmb combines virtual clients, partial aggregation, private local fine-tuning, and public Pre-training to achieve strong privacy utility trade-offs. We apply DP-FedEmb to train image embedding models for faces, landmarks and natural species, and demonstrate its superior utility under same privacy budget on benchmark datasets DigiFace. EMNIST, GLD and iNaturalist. We further illustrate it is possible to achieve strong user-level DP guarantees of ∊ < 2 while controlling the utility drop within 5%, when millions of users can participate in training. Zheng Xu 0002, Maxwell D. Collins, Yuxiao Wang 0001, Liviu Panait, Sewoong Oh, Sean Augenstein, Ting Liu 0005, Florian Schroff, H. Brendan McMahan |
CVPR | 1 |
| 2023 | On the Convergence of Federated Averaging with Cyclic Client ParticipationabstractFederated Averaging (FedAvg) and its variants are the most popular optimization algorithms in federated learning (FL). Previous convergence analyses of FedAvg either assume full client participation or partial client participation where the clients can be uniformly sampled. However, in practical cross-device FL systems, only a subset of clients that satisfy local criteria such as battery status, network connectivity, and maximum participation frequency requirements (to ensure privacy) are available for training at a given time. As a result, client availability follows a *natural cyclic pattern*. We provide (to our knowledge) the first theoretical framework to analyze the convergence of FedAvg with cyclic client participation with several different client optimizers such as GD, SGD, and shuffled SGD. Our analysis discovers that cyclic client participation can achieve a faster asymptotic convergence rate than vanilla FedAvg with uniform client participation under suitable conditions, providing valuable insights into the design of client sampling protocols. Yae Jee Cho, Pranay Sharma, Gauri Joshi, Zheng Xu 0002, Satyen Kale, Tong Zhang 0001 |
ICML | 4 |
| 2023 | Beyond Uniform Lipschitz Condition in Differentially Private OptimizationabstractMost prior results on differentially private stochastic gradient descent (DP-SGD) are derived under the simplistic assumption of uniform Lipschitzness, i.e., the per-sample gradients are uniformly bounded. We generalize uniform Lipschitzness by assuming that the per-sample gradients have sample-dependent upper bounds, i.e., per-sample Lipschitz constants, which themselves may be unbounded. We provide principled guidance on choosing the clip norm in DP-SGD for convex over-parameterized settings satisfying our general version of Lipschitzness when the per-sample Lipschitz constants are bounded; specifically, we recommend tuning the clip norm only till values up to the minimum per-sample Lipschitz constant. This finds application in the private training of a softmax layer on top of a deep network pre-trained on public data. We verify the efficacy of our recommendation via experiments on 8 datasets. Furthermore, we provide new convergence results for DP-SGD on convex and nonconvex functions when the Lipschitz constants are unbounded but have bounded moments, i.e., they are heavy-tailed. Rudrajit Das, Satyen Kale, Zheng Xu 0002, Tong Zhang 0001, Sujay Sanghavi |
ICML | 3 |
| 2023 | How to DP-fy ML: A Practical Tutorial to Machine Learning with Differential PrivacyabstractMachine Learning (ML) models are ubiquitous in real world applications and are a constant focus of research. At the same time, the community has started to realize the importance of protecting the privacy of models' training data. Natalia Ponomareva 0001, Sergei Vassilvitskii, Zheng Xu 0002, H. Brendan McMahan, Alexey Kurakin, Chiyaun Zhang |
KDD | 3 |
| 2023 | (Amplified) Banded Matrix Factorization: A unified approach to private trainingabstractMatrix factorization (MF) mechanisms for differential privacy (DP) have substantially improved the state-of-the-art in privacy-utility-computation tradeoffs for ML applications in a variety of scenarios, but in both the centralized and federated settings there remain instances where either MF cannot be easily applied, or other algorithms provide better tradeoffs (typically, as $\epsilon$ becomes small).
In this work, we show how MF can subsume prior state-of-the-art algorithms in both federated and centralized training settings, across all privacy budgets. The key technique throughout is the construction of MF mechanisms with banded matrices (lower-triangular matrices with at most $\hat{b}$ nonzero bands including the main diagonal). For cross-device federated learning (FL), this enables multiple-participations with a relaxed device participation schema compatible with practical FL infrastructure (as demonstrated by a production deployment). In the centralized setting, we prove that banded matrices enjoy the same privacy amplification results as the ubiquitous DP-SGD algorithm, but can provide strictly better performance in most scenarios---this lets us always at least match DP-SGD, and often outperform it Christopher A. Choquette-Choo, Arun Ganesh, Ryan McKenna, H. Brendan McMahan, John Rush, Abhradeep Thakurta, Zheng Xu 0002 |
NeurIPS | 7 |
| 2023 | How to DP-fy ML: A Practical Guide to Machine Learning with Differential PrivacyabstractMachine Learning (ML) models are ubiquitous in real-world applications and are a constant focus of research. Modern ML models have become more complex, deeper, and harder to reason about. At the same time, the community has started to realize the importance of protecting the privacy of the training data that goes into these models. Differential Privacy (DP) has become a gold standard for making formal statements about data anonymization. However, while some adoption of DP has happened in industry, attempts to apply DP to real world complex ML models are still few and far between. The adoption of DP is hindered by limited practical guidance of what DP protection entails, what privacy guarantees to aim for, and the difficulty of achieving good privacy-utility-computation trade-offs for ML models. Tricks for tuning and maximizing performance are scattered among papers or stored in the heads of practitioners, particularly with respect to the challenging task of hyperparameter tuning. Furthermore, the literature seems to present conflicting evidence on how and whether to apply architectural adjustments and which components are “safe” to use with DP. In this survey paper, we attempt to create a self-contained guide that gives an in-depth overview of the field of DP ML. We aim to assemble information about achieving the best possible DP ML model with rigorous privacy guarantees. Our target audience is both researchers and practitioners. Researchers interested in DP for ML will benefit from a clear overview of current advances and areas for improvement. We also include theory-focused sections that highlight important topics such as privacy accounting and convergence. For a practitioner, this survey provides a background in DP theory and a clear step-by-step guide for choosing an appropriate privacy definition and approach, implementing DP training, potentially updating the model architecture, and tuning hyperparameters. For both researchers and practitioners, consistently and fully reporting privacy guarantees is critical, so we propose a set of specific best practices for stating guarantees. With sufficient computation and a sufficiently large training set or supplemental nonprivate data, both good accuracy (that is, almost as good as a non-private model) and good privacy can often be achievable. And even when computation and dataset size are limited, there are advantages to training with even a weak (but still finite) formal DP guarantee. Hence, we hope this work will facilitate more widespread deployments of DP ML models. Natalia Ponomareva 0001, Hussein Hazimeh 0001, Alexey Kurakin, Zheng Xu 0002, Carson Denison, H. Brendan McMahan, Sergei Vassilvitskii, Steve Chien, Abhradeep Thakurta |
J. Artif. Intell. Res. | 4 |
| 2022 | Diurnal or Nocturnal? Federated Learning of Multi-branch Networks from Periodically Shifting Distributions
Chen Zhu 0001, Zheng Xu 0002, Mingqing Chen, Jakub Konecný, Andrew Hard, Tom Goldstein |
ICLR | 2 |
| 2021 | Practical and Private (Deep) Learning Without Sampling or ShufflingabstractWe consider training models with differential privacy (DP) using mini-batch gradients. The existing state-of-the-art, Differentially Private Stochastic Gradient Descent (DP-SGD), requires \emph{privacy amplification by sampling or shuffling} to obtain the best privacy/accuracy/computation trade-offs. Unfortunately, the precise requirements on exact sampling and shuffling can be hard to obtain in important practical scenarios, particularly federated learning (FL). We design and analyze a DP variant of Follow-The-Regularized-Leader (DP-FTRL) that compares favorably (both theoretically and empirically) to amplified DP-SGD, while allowing for much more flexible data access patterns. DP-FTRL does not use any form of privacy amplification. Peter Kairouz, H. Brendan McMahan, Shuang Song 0001, Om Thakkar 0001, Abhradeep Thakurta, Zheng Xu 0002 |
ICML | 6 |
| 2021 | GradInit: Learning to Initialize Neural Networks for Stable and Efficient TrainingabstractInnovations in neural architectures have fostered significant breakthroughs in language modeling and computer vision. Unfortunately, novel architectures often result in challenging hyper-parameter choices and training instability if the network parameters are not properly initialized. A number of architecture-specific initialization schemes have been proposed, but these schemes are not always portable to new architectures. This paper presents GradInit, an automated and architecture agnostic method for initializing neural networks. GradInit is based on a simple heuristic; the norm of each network layer is adjusted so that a single step of SGD or Adam with prescribed hyperparameters results in the smallest possible loss value. This adjustment is done by introducing a scalar multiplier variable in front of each parameter block, and then optimizing these variables using a simple numerical scheme. GradInit accelerates the convergence and test performance of many convolutional architectures, both with or without skip connections, and even without normalization layers. It also improves the stability of the original Transformer architecture for machine translation, enabling training it without learning rate warmup using either Adam or SGD under a wide range of learning rates and momentum coefficients. Code is available at https://github.com/zhuchen03/gradinit. Chen Zhu 0001, Renkun Ni, Zheng Xu 0002, Kezhi Kong, W. Ronny Huang, Tom Goldstein |
NeurIPS | 3 |
| 2020 | Universal Adversarial TrainingabstractStandard adversarial attacks change the predicted class label of a selected image by adding specially tailored small perturbations to its pixels. In contrast, a universal perturbation is an update that can be added to any image in a broad class of images, while still changing the predicted class label. We study the efficient generation of universal adversarial perturbations, and also efficient methods for hardening networks to these attacks. We propose a simple optimization-based universal attack that reduces the top-1 accuracy of various network architectures on ImageNet to less than 20%, while learning the universal perturbation 13× faster than the standard method.To defend against these perturbations, we propose universal adversarial training, which models the problem of robust classifier generation as a two-player min-max game, and produces robust models with only 2× the cost of natural training. We also propose a simultaneous stochastic gradient method that is almost free of extra computation, which allows us to do universal adversarial training on ImageNet. Ali Shafahi, Mahyar Najibi, Zheng Xu 0002, John Dickerson 0001, Larry Davis 0001, Tom Goldstein |
AAAI | 3 |
| 2020 | The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient DescentabstractThis paper studies how neural network architecture affects the speed of training. We introduce a simple concept called gradient confusion to help formally analyze this. When gradient confusion is high, stochastic gradients produced by different data samples may be negatively correlated, slowing down convergence. But when gradient confusion is low, data samples interact harmoniously, and training proceeds quickly. Through theoretical and experimental results, we demonstrate how the neural network architecture affects gradient confusion, and thus the efficiency of training. Our results show that, for popular initialization techniques, increasing the width of neural networks leads to lower gradient confusion, and thus faster model training. On the other hand, increasing the depth of neural networks has the opposite effect. Our results indicate that alternate initialization techniques or networks using both batch normalization and skip connections help reduce the training burden of very deep networks. Karthik Abinav Sankararaman, Soham De, Zheng Xu 0002, W. Ronny Huang, Tom Goldstein |
ICML | 3 |
| 2020 | Adversarial training for fast arbitrary style transfer
Zheng Xu 0002, Kimberly Wilber, Aaron Hertzmann, Hailin Jin |
Comput. Graph. | 1 |
| 2019 | Adversarial training for free!abstractAdversarial training, in which a network is trained on adversarial examples, is one of the few defenses against adversarial attacks that withstands strong attacks. Unfortunately, the high cost of generating strong adversarial examples makes standard adversarial training impractical on large-scale problems like ImageNet. We present an algorithm that eliminates the overhead cost of generating adversarial examples by recycling the gradient information computed when updating model parameters. Our "free" adversarial training algorithm achieves comparable robustness to PGD adversarial training on the CIFAR-10 and CIFAR-100 datasets at negligible additional cost compared to natural training, and can be 7 to 30 times faster than other strong adversarial training methods. Using a single workstation with 4 P100 GPUs and 2 days of runtime, we can train a robust model for the large-scale ImageNet classification task that maintains 40% accuracy against PGD attacks. Ali Shafahi, Mahyar Najibi, Amin Ghiasi, Zheng Xu 0002, John Dickerson 0001, Christoph Studer, Larry Davis 0001, Gavin Taylor, Tom Goldstein |
NeurIPS | 4 |
| 2018 | Towards Perceptual Image Dehazing by Physics-Based Disentanglement and Adversarial TrainingabstractSingle image dehazing is a challenging under-constrained problem because of the ambiguities of unknown scene radiance and transmission. Previous methods solve this problem using various hand-designed priors or by supervised training on synthetic hazy image pairs. In practice, however, the predefined priors are easily violated and the paired image data is unavailable for supervised training. In this work, we propose Disentangled Dehazing Network, an end-to-end model that generates realistic haze-free images using only unpaired supervision. Our approach alleviates the paired training constraint by introducing a physical-model based disentanglement and reconstruction mechanism. A multi-scale adversarial training is employed to generate perceptually haze-free images. Experimental results on synthetic datasets demonstrate our superior performance compared with the state-of-the-art methods in terms of PSNR, SSIM and CIEDE2000. Through training on purely natural haze-free and hazy images from our collected HazyCity dataset, our model can generate more perceptually appealing dehazing results. Xitong Yang, Zheng Xu 0002, Jiebo Luo 0001 |
AAAI | 2 |
| 2018 | Training Student Networks for Acceleration with Conditional Adversarial Networks
Zheng Xu 0002, Yen-Chang Hsu, Jiawei Huang 0006 |
BMVC | 1 |
| 2018 | Strong Baseline for Single Image Dehazing with Deep Features and Instance Normalization
Zheng Xu 0002, Xitong Yang, Xiaoshuai Sun |
BMVC | 1 |
| 2018 | Stabilizing Adversarial Nets with Prediction Methods
Abhay Kumar Yadav, Sohil Shah, Zheng Xu 0002, David Jacobs 0001, Tom Goldstein |
ICLR (Poster) | 3 |
| 2018 | Learning to Cluster for Proposal-Free Instance SegmentationabstractThis work proposed a novel learning objective to train a deep neural network to perform end-to-end image pixel clustering. We applied the approach to instance segmentation, which is at the intersection of image semantic segmentation and object detection. We utilize the most fundamental property of instance labeling- the pairwise relationship between pixels-as the supervision to formulate the learning objective, then apply it to train a fully convolutional network (FCN) for learning to perform pixel-wise clustering. The resulting clusters can be used as the instance labeling directly. To support labeling of an unlimited number of instance, we further formulate ideas from graph coloring theory into the proposed learning objective. The evaluation on the Cityscapes dataset demonstrates strong performance and therefore proof of the concept. Moreover, our approach won the second place in the lane detection competition of 2017 CVPR Autonomous Driving Challenge, and was the top performer without using external data. Yen-Chang Hsu, Zheng Xu 0002, Zsolt Kira, Jiawei Huang 0006 |
IJCNN | 2 |
| 2018 | Visualizing the Loss Landscape of Neural NetsabstractNeural network training relies on our ability to find "good" minimizers of highly non-convex loss functions. It is well known that certain network architecture designs (e.g., skip connections) produce loss functions that train easier, and well-chosen training parameters (batch size, learning rate, optimizer) produce minimizers that generalize better. However, the reasons for these differences, and their effect on the underlying loss landscape, is not well understood. In this paper, we explore the structure of neural loss functions, and the effect of loss landscapes on generalization, using a range of visualization methods. First, we introduce a simple "filter normalization" method that helps us visualize loss function curvature, and make meaningful side-by-side comparisons between loss functions. Then, using a variety of visualizations, we explore how network architecture affects the loss landscape, and how training parameters affect the shape of minimizers. Hao Li 0022, Zheng Xu 0002, Gavin Taylor, Christoph Studer, Tom Goldstein |
NeurIPS | 2 |
| 2018 | Domain Generalization and Adaptation Using Low Rank Exemplar SVMsabstractDomain adaptation between diverse source and target domains is challenging, especially in the real-world visual recognition tasks where the images and videos consist of significant variations in viewpoints, illuminations, qualities, etc. In this paper, we propose a new approach for domain generalization and domain adaptation based on exemplar SVMs. Specifically, we decompose the source domain into many subdomains, each of which contains only one positive training sample and all negative samples. Each subdomain is relatively less diverse, and is expected to have a simpler distribution. By training one exemplar SVM for each subdomain, we obtain a set of exemplar SVMs. To further exploit the inherent structure of source domain, we introduce a nuclear-norm based regularizer into the objective function in order to enforce the exemplar SVMs to produce a low-rank output on training samples. In the prediction process, the confident exemplar SVM classifiers are selected and reweigted according to the distribution mismatch between each subdomain and the test sample in the target domain. We formulate our approach based on the logistic regression and least square SVM algorithms, which are referred to as low rank exemplar SVMs (LRE-SVMs) and low rank exemplar least square SVMs (LRE-LSSVMs), respectively. A fast algorithm is also developed for accelerating the training of LRE-LSSVMs. We further extend Domain Adaptation Machine (DAM) to learn an optimal target classifier for domain adaptation, and show that our approach can also be applied to domain adaptation with evolving target domain, where the target data distribution is gradually changing. The comprehensive experiments for object recognition and action recognition demonstrate the effectiveness of our approach for domain generalization and domain adaptation with fixed and evolving target domains. Wen Li 0001, Zheng Xu 0002, Dong Xu 0001, Dengxin Dai, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Adaptive ADMM with Spectral Penalty Parameter SelectionabstractThe alternating direction method of multipliers (ADMM) is a versatile tool for solving a wide range of constrained optimization problems. However, its performance is highly sensitive to a penalty parameter, making ADMM often unreliable and hard to automate for a non-expert user. We tackle this weakness of ADMM by proposing a method that adaptively tunes the penalty parameter to achieve fast convergence. The resulting adaptive ADMM (AADMM) algorithm, inspired by the successful Barzilai-Borwein spectral method for gradient descent, yields fast convergence and relative insensitivity to the initial stepsize and problem scaling. Zheng Xu 0002, Mário A. T. Figueiredo, Tom Goldstein |
AISTATS | 1 |
| 2017 | Adaptive Relaxed ADMM: Convergence Theory and Practical ImplementationabstractMany modern computer vision and machine learning applications rely on solving difficult optimization problems that involve non-differentiable objective functions and constraints. The alternating direction method of multipliers (ADMM) is a widely used approach to solve such problems. Relaxed ADMM is a generalization of ADMM that often achieves better performance, but its efficiency depends strongly on algorithm parameters that must be chosen by an expert user. We propose an adaptive method that automatically tunes the key algorithm parameters to achieve optimal performance without user oversight. Inspired by recent work on adaptivity, the proposed adaptive relaxed ADMM (ARADMM) is derived by assuming a Barzilai-Borwein style linear gradient. A detailed convergence analysis of ARADMM is provided, and numerical results on several applications demonstrate fast practical convergence. Zheng Xu 0002, Mário A. T. Figueiredo, Christoph Studer, Tom Goldstein |
CVPR | 1 |
| 2017 | Adaptive Consensus ADMM for Distributed OptimizationabstractThe alternating direction method of multipliers (ADMM) is commonly used for distributed model fitting problems, but its performance and reliability depend strongly on user-defined penalty parameters. We study distributed ADMM methods that boost performance by using different fine-tuned algorithm parameters on each worker node. We present a O(1/k) convergence rate for adaptive ADMM methods with node-specific parameters, and propose adaptive consensus ADMM (ACADMM), which automatically tunes parameters without user oversight. Zheng Xu 0002, Gavin Taylor, Hao Li 0022, Mário A. T. Figueiredo, Tom Goldstein |
ICML | 1 |
| 2017 | Training Quantized Nets: A Deeper UnderstandingabstractCurrently, deep neural networks are deployed on low-power portable devices by first training a full-precision model using powerful hardware, and then deriving a corresponding low-precision model for efficient inference on such systems. However, training models directly with coarsely quantized weights is a key step towards learning on embedded platforms that have limited computing resources, memory capacity, and power consumption. Numerous recent publications have studied methods for training quantized networks, but these studies have mostly been empirical. In this work, we investigate training methods for quantized neural networks from a theoretical viewpoint. We first explore accuracy guarantees for training methods under convexity assumptions. We then look at the behavior of these algorithms for non-convex problems, and show that training algorithms that exploit high-precision representations have an important greedy search phase that purely quantized training methods lack, which explains the difficulty of training using low-precision arithmetic. Hao Li 0022, Soham De, Zheng Xu 0002, Christoph Studer, Hanan Samet, Tom Goldstein |
NIPS | 3 |
| 2016 | Training Neural Networks Without Gradients: A Scalable ADMM ApproachabstractWith the growing importance of large network models and enormous training datasets, GPUs have become increasingly necessary to train neural networks. This is largely because conventional optimization algorithms rely on stochastic gradient methods that don’t scale well to large numbers of cores in a cluster setting. Furthermore, the convergence of all gradient methods, including batch methods, suffers from common problems like saturation effects, poor conditioning, and saddle points. This paper explores an unconventional training method that uses alternating direction methods and Bregman iteration to train networks without gradient descent steps. The proposed method reduces the network training problem to a sequence of minimization sub-steps that can each be solved globally in closed form. The proposed method is advantageous because it avoids many of the caveats that make gradient methods slow on highly non-convex problems. In addition, the method exhibits strong scaling in the distributed setting, yielding linear speedups even when split over thousands of cores. Gavin Taylor, Ryan Burmeister, Zheng Xu 0002, Ankit B. Patel, Tom Goldstein |
ICML | 3 |
| 2015 | Exploiting Low-rank Structure for Discriminative Sub-categorizationabstractIn visual recognition, sub-categorization has been proposed to deal with large intraclass variance of samples in a category. Instead of learning a single classifier for each category, discriminant sub-categorization approaches divide a category into several subcategories and simultaneously train classifiers for each sub-category. In this paper, we propose a novel approach for discriminative sub-categorization. Our method jointly trains the exemplar classifier for each positive sample to address the intra-variance of a category and exploits the low rank structure to preserve common information while discovering sub-categories. We formulate the problem as a convex objective function and introduce an efficient solver based on alternating direction method of multipliers. Comprehensive experiments on various datasets demonstrate the effectiveness and efficiency of the proposed method in both sub-category discovery and visual recognition. Zheng Xu 0002, Kuiyuan Yang, Tom Goldstein |
BMVC | 1 |
| 2014 | Exploiting Low-Rank Structure from Latent Domains for Domain Generalization
Zheng Xu 0002, Wen Li 0001, Li Niu 0002, Dong Xu 0001 |
ECCV (3) | 1 |
| 2013 | Mining visualnessabstractTo understand which concepts are visualizable and to what extent they can be visualized are worthwhile for multimedia and computer vision research. Unfortunately, few previous works have ever touched such topics. In this paper, we propose an unified model to automatically identify visual concepts and estimate their visual characteristics, or visualness, from a large-scale image dataset. To this end, an image heterogeneous graph is first built to integrate various visual features, and then a simultaneous ranking and clustering algorithm is introduced to generate visually and semantically compact image clusters, named visualsets. Based on the visualsets, visualizable concepts are discovered and their visualness scores are estimated. The experimental results demonstrate the effectiveness of the proposed schema. Zheng Xu 0002, Xin-Jing Wang, Chang Wen Chen |
ICME | 1 |
| 2012 | Towards indexing representative images on the webabstractEven after 20 years of research on real-world image retrieval, there is still a big gap between what search engines can provide and what users expect to see. To bridge this gap, we present an image knowledge base, ImageKB, a graph representation of structured entities, categories, and representative images, as a new basis for practical image indexing and search. ImageKB is automatically constructed via a both bottom-up and top-down, scalable approach that efficiently matches 2 billion web images onto an ontology with millions of nodes. Our approach consists of identifying duplicate image clusters from billions of images, obtaining a candidate set of entities and their images, discovering definitive texts to represent an image and identifying representative images for an entity. To date, ImageKB contains 235.3M representative images corresponding to 0.52M entities, much larger than the state-of-the-art alternative ImageNet that contains 14.2M images for 0.02M synsets. Compared to existing image databases, ImageKB reflects the distributions of both images on the web and users' interests, contains rich semantic descriptions for images and entities, and can be widely used for both text to image search and image to text understanding. Xin-Jing Wang, Zheng Xu 0002, Lei Zhang 0001, Ce Liu 0001, Yong Rui |
ACM Multimedia | 2 |