VLDB 2026 Research / reviewers in the wild / expert
Maziar Sanjabi
dblp:21/8577 · also Maziar Sanjabi Boroujeni
· DBLP profile ↗
29ranked-venue papers
2as first author
19since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT ExamplesabstractChain-of-thought (CoT) prompting combined with few-shot in-context learning (ICL) has unlocked significant reasoning capabilities in large language models (LLMs).However, ICL with CoT examples is ineffective on novel tasks when the pre-training knowledge is insufficient.We study this problem in a controlled setting using the CoT-ICL Lab (Kothapalli et al., 2025) framework, and propose meta-training techniques to learn novel abstract reasoning tasks in-context.Although CoT examples facilitate reasoning, we noticed that their excessive inclusion during metatraining degrades performance when CoT supervision is limited.To mitigate such behavior, we propose CoT-Recipe, a formal approach to modulate the mix of CoT and non-CoT examples in meta-training sequences.We demonstrate that careful modulation via CoT-Recipe can increase the accuracy of transformers on novel tasks by up to 300% even when there are no CoT examples available in-context.We confirm the broader effectiveness of these techniques by applying them to pretrained LLMs (Qwen2.5 series) for symbolic reasoning tasks and observing gains of up to 130% in accuracy. Vignesh Kothapalli, Ata Fatahi Baarzi, Hamed Firooz, Maziar Sanjabi |
ACL (1) | 4 |
| 2025 | CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context DemonstrationsabstractWe introduce CoT-ICL Lab, a framework and methodology to generate synthetic tokenized datasets and systematically study chain-of thought (CoT) in-context learning (ICL) in language models. CoT-ICL Lab allows fine grained control over the complexity of in-context examples by decoupling (1) the causal structure involved in chain token generation from (2) the underlying token processing functions. We train decoder-only transformers (up to 700M parameters) on these datasets and show that CoT accelerates the accuracy transition to higher values across model sizes. In particular, we find that model depth is crucial for leveraging CoT with limited in-context examples, while more examples help shallow models match deeper model performance. Additionally, limiting the diversity of token processing functions throughout training improves causal structure learning via ICL. We also interpret these transitions by analyzing transformer embeddings and attention maps. Overall, CoT-ICL Lab serves as a simple yet powerful testbed for theoretical and empirical insights into ICL and CoT in language models. Vignesh Kothapalli, Hamed Firooz, Maziar Sanjabi |
ACL (1) | 3 |
| 2024 | Measuring Self-Supervised Representation Quality for Downstream Classification Using Discriminative FeaturesabstractSelf-supervised learning (SSL) has shown impressive results in downstream classification tasks. However, there is limited work in understanding their failure modes and interpreting their learned representations. In this paper, we study the representation space of state-of-the-art self-supervised models including SimCLR, SwaV, MoCo, BYOL, DINO, SimSiam, VICReg and Barlow Twins. Without the use of class label information, we discover discriminative features that correspond to unique physical attributes in images, present mostly in correctly-classified representations. Using these features, we can compress the representation space by up to$40% without significantly affecting linear classification performance. We then propose Self-Supervised Representation Quality Score (or Q-Score), an unsupervised score that can reliably predict if a given sample is likely to be mis-classified during linear evaluation, achieving AUPRC of 91.45 on ImageNet-100 and 78.78 on ImageNet-1K. Q-Score can also be used as a regularization term on pre-trained encoders to remedy low-quality representations. Fine-tuning with Q-Score regularization can boost the linear probing accuracy of SSL models by up to 5.8% on ImageNet-100 and 3.7% on ImageNet-1K compared to their baselines. Finally, using gradient heatmaps and Salient ImageNet masks, we define a metric to quantify the interpretability of each representation. We show that discriminative features are strongly correlated to core attributes and, enhancing these features through Q-score regularization makes SSL representations more interpretable. Neha Mukund Kalibhat, Kanika Narang, Hamed Firooz, Maziar Sanjabi, Soheil Feizi |
AAAI | 4 |
| 2024 | Text-to-Sticker: Style Tailoring Latent Diffusion Models for Human Expression
Animesh Sinha, Anmol Kalia, Arantxa Casanova, Elliot Blanchard, David Yan, Winnie Zhang, Tony Nelli, Hardik Shah, Licheng Yu, Mitesh Kumar Singh, Ankit Ramchandani, Maziar Sanjabi, Sonal Gupta, Amy Bearman, Dhruv Mahajan 0001 |
ECCV (70) | 14 |
| 2024 | Distilling Knowledge from Text-to-Image Generative Models Improves Visio-Linguistic Reasoning in CLIPabstractImage-text contrastive models like CLIP have wide applications in zero-shot classification, image-text retrieval, and transfer learning.However, they often struggle on compositional visio-linguistic tasks (e.g., attribute-binding or object-relationships) where their performance is no better than random chance.To address this, we introduce SDS-CLIP, a lightweight and sample-efficient distillation method to enhance CLIP's compositional visio-linguistic reasoning.Our approach fine-tunes CLIP using a distillation objective borrowed from large text-to-image generative models like Stable-Diffusion, which are known for their strong visio-linguistic reasoning abilities.On the challenging Winoground benchmark, SDS-CLIP improves the visio-linguistic performance of various CLIP models by up to 7%, while on the ARO dataset, it boosts performance by up to 3%.This work underscores the potential of well-designed distillation objectives from generative models to enhance contrastive imagetext models with improved visio-linguistic reasoning capabilities. Samyadeep Basu, Shell Xu Hu, Maziar Sanjabi, Daniela Massiceti, Soheil Feizi |
EMNLP | 3 |
| 2024 | Differentially Private Representation Learning via Image CaptioningabstractDifferentially private (DP) machine learning is considered the gold-standard solution for training a model from sensitive data while still preserving privacy. However, a major barrier to achieving this ideal is its sub-optimal privacy-accuracy trade-off, which is particularly visible in DP representation learning. Specifically, it has been shown that under modest privacy budgets, most models learn representations that are not significantly better than hand-crafted features. In this work, we show that effective DP representation learning can be done via image captioning and scaling up to internet-scale multimodal datasets. Through a series of engineering tricks, we successfully train a DP image captioner (DP-Cap) on a 233M subset of LAION-2B from scratch using a reasonable amount of computation, and obtaining unprecedented high-quality image features that can be used in a variety of downstream vision and vision-language tasks. For example, under a privacy budget of $\varepsilon=8$ for the LAION dataset, a linear classifier trained on top of learned DP-Cap features attains $65.8\%$ accuracy on ImageNet-1K, considerably improving the previous SOTA of $56.5\%$. Our work challenges the prevailing sentiment that high-utility DP representation learning cannot be achieved by training from scratch. Tom Sander, Yaodong Yu, Maziar Sanjabi, Alain Durmus, Yi Ma 0001, Kamalika Chaudhuri, Chuan Guo 0001 |
ICML | 3 |
| 2024 | ViP: A Differentially Private Foundation Model for Computer VisionabstractArtificial intelligence (AI) has seen a tremendous surge in capabilities thanks to the use of foundation models trained on internet-scale data. On the flip side, the uncurated nature of internet-scale data also poses significant privacy and legal risks, as they often contain personal information or copyrighted material that should not be trained on without permission. In this work, we propose as a mitigation measure a recipe to train foundation vision models via self-supervised learning with differential privacy (DP) guarantee. We identify masked autoencoders as a suitable learning algorithm that aligns well with DP-SGD, and train *ViP*---a **Vi**sion transformer with differential **P**rivacy---under a strict privacy budget of $\epsilon=8$ on the LAION400M dataset. We evaluate the quality of representation learned by ViP using standard downstream vision tasks; in particular, ViP achieves a (non-private) linear probing accuracy of 55.7% on ImageNet, comparable to that of end-to-end trained AlexNet (trained and evaluated on ImageNet). Our result suggests that scaling to internet-scale data can be practical for private learning. Code and DP pre-trained models are available at https://github.com/facebookresearch/ViP-MAE. Yaodong Yu, Maziar Sanjabi, Yi Ma 0001, Kamalika Chaudhuri, Chuan Guo 0001 |
ICML | 2 |
| 2024 | RESPROMPT: Residual Connection Prompting Advances Multi-Step Reasoning in Large Language ModelsabstractSong Jiang, Zahra Shakeri, Aaron Chan, Maziar Sanjabi, Hamed Firooz, Yinglong Xia, Bugra Akyildiz, Yizhou Sun, Jinchao Li, Qifan Wang, Asli Celikyilmaz. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Song Jiang 0002, Zahra Shakeri, Aaron Chan, Maziar Sanjabi, Hamed Firooz, Yinglong Xia, Bugra Akyildiz, Yizhou Sun, Jinchao Li, Qifan Wang 0001, Asli Celikyilmaz |
NAACL-HLT | 4 |
| 2023 | Defending Against Patch-based Backdoor Attacks on Self-Supervised LearningabstractRecently, self-supervised learning (SSL) was shown to be vulnerable to patch-based data poisoning backdoor attacks. It was shown that an adversary can poison a small part of the unlabeled data so that when a victim trains an SSL model on it, the final model will have a back-door that the adversary can exploit. This work aims to defend self-supervised learning against such attacks. We use a three-step defense pipeline, where we first train a model on the poisoned data. In the second step, our proposed defense algorithm (PatchSearch) uses the trained model to search the training data for poisoned samples and removes them from the training set. In the third step, a final model is trained on the cleaned-up training set. Our results show that PatchSearch is an effective defense. As an example, it improves a model's accuracy on images containing the trigger from 38.2% to 63.7% which is very close to the clean model's accuracy, 64.6%. More-over, we show that PatchSearch outperforms baselines and state-of-the-art defense approaches including those using additional clean, trusted data. Our code is available at https://github.com/UCDvision/PatchSearch Ajinkya Tejankar, Maziar Sanjabi, Qifan Wang 0001, Sinong Wang, Hamed Firooz, Hamed Pirsiavash, Liang Tan 0005 |
CVPR | 2 |
| 2023 | COFFEE: Counterfactual Fairness for Personalized Text Generation in Explainable RecommendationabstractNan Wang, Qifan Wang, Yi-Chia Wang, Maziar Sanjabi, Jingzhou Liu, Hamed Firooz, Hongning Wang, Shaoliang Nie. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Qifan Wang 0001, Yi-Chia Wang, Maziar Sanjabi, Jingzhou Liu, Hamed Firooz, Hongning Wang, Shaoliang Nie |
EMNLP | 4 |
| 2023 | Where to Begin? On the Impact of Pre-Training and Initialization in Federated Learning
John Nguyen, Kshitiz Malik, Maziar Sanjabi, Michael G. Rabbat |
ICLR | 4 |
| 2023 | Analyzing Privacy Leakage in Machine Learning via Multiple Hypothesis Testing: A Lesson From FanoabstractDifferential privacy (DP) is by far the most widely accepted framework for mitigating privacy risks in machine learning. However, exactly how small the privacy parameter $\epsilon$ needs to be to protect against certain privacy risks in practice is still not well-understood. In this work, we study data reconstruction attacks for discrete data and analyze it under the framework of multiple hypothesis testing. For a learning algorithm satisfying $(\alpha, \epsilon)$-Renyi DP, we utilize different variants of the celebrated Fano’s inequality to upper bound the attack advantage of a data reconstruction adversary. Our bound can be numerically computed to relate the parameter $\epsilon$ to the desired level of privacy protection in practice, and complements the empirical evidence for the effectiveness of DP against data reconstruction attacks even at relatively large values of $\epsilon$. Chuan Guo 0001, Alexandre Sablayrolles, Maziar Sanjabi |
ICML | 3 |
| 2023 | Identifying Interpretable Subspaces in Image RepresentationsabstractWe propose Automatic Feature Explanation using Contrasting Concepts (FALCON), an interpretability framework to explain features of image representations. For a target feature, FALCON captions its highly activating cropped images using a large captioning dataset (like LAION-400m) and a pre-trained vision-language model like CLIP. Each word among the captions is scored and ranked leading to a small number of shared, human-understandable concepts that closely describe the target feature. FALCON also applies contrastive interpretation using lowly activating (counterfactual) images, to eliminate spurious concepts. Although many existing approaches interpret features independently, we observe in state-of-the-art self-supervised and supervised models, that less than 20% of the representation space can be explained by individual features. We show that features in larger spaces become more interpretable when studied in groups and can be explained with high-order scoring concepts through FALCON. We discuss how extracted concepts can be used to explain and debug failures in downstream tasks. Finally, we present a technique to transfer concepts from one (explainable) representation space to another unseen representation space by learning a simple linear transformation. Neha Mukund Kalibhat, Shweta Bhardwaj, C. Bayan Bruss, Hamed Firooz, Maziar Sanjabi, Soheil Feizi |
ICML | 5 |
| 2023 | Text-To-Concept (and Back) via Cross-Model AlignmentabstractWe observe that the mapping between an image's representation in one model to its representation in another can be learned surprisingly well with just a linear layer, even across diverse models. Building on this observation, we propose *text-to-concept*, where features from a fixed pretrained model are aligned linearly to the CLIP space, so that text embeddings from CLIP's text encoder become directly comparable to the aligned features. With text-to-concept, we convert fixed off-the-shelf vision encoders to surprisingly strong zero-shot classifiers for free, with accuracy at times even surpassing that of CLIP, despite being much smaller models and trained on a small fraction of the data compared to CLIP. We show other immediate use-cases of text-to-concept, like building concept bottleneck models with no concept supervision, diagnosing distribution shifts in terms of human concepts, and retrieving images satisfying a set of text-based constraints. Lastly, we demonstrate the feasibility of *concept-to-text*, where vectors in a model's feature space are decoded by first aligning to the CLIP before being fed to a GPT-based generative model. Our work suggests existing deep models, with presumably diverse architectures and training, represent input samples relatively similarly, and a two-way communication across model representation spaces and to humans (through language) is viable. Mazda Moayeri, Keivan Rezaei, Maziar Sanjabi, Soheil Feizi |
ICML | 3 |
| 2023 | On Tilted Losses in Machine Learning: Theory and ApplicationsabstractExponential tilting is a technique commonly used in fields such as statistics, probability, information theory, and optimization to create parametric distribution shifts. Despite its prevalence in related fields, tilting has not seen widespread use in machine learning. In this work, we aim to bridge this gap by exploring the use of tilting in risk minimization. We study a simple extension to ERM---tilted empirical risk minimization (TERM)---which uses exponential tilting to flexibly tune the impact of individual losses. The resulting framework has several useful properties: We show that TERM can increase or decrease the influence of outliers, respectively, to enable fairness or robustness; has variance-reduction properties that can benefit generalization; and can be viewed as a smooth approximation to the tail probability of losses. Our work makes connections between TERM and related objectives, such as Value-at-Risk, Conditional Value-at-Risk, and distributionally robust optimization (DRO). We develop batch and stochastic first-order optimization methods for solving TERM, provide convergence guarantees for the solvers, and show that the framework can be efficiently solved relative to common alternatives. Finally, we demonstrate that TERM can be used for a multitude of applications in machine learning, such as enforcing fairness between subgroups, mitigating the effect of outliers, and handling class imbalance. Despite the straightforward modification TERM makes to traditional ERM objectives, we find that the framework can consistently outperform ERM and deliver competitive performance with state-of-the-art, problem-specific approaches. Tian Li 0005, Ahmad Beirami, Maziar Sanjabi, Virginia Smith |
J. Mach. Learn. Res. | 3 |
| 2022 | UNIREX: A Unified Learning Framework for Language Model Rationale ExtractionabstractAn extractive rationale explains a language model’s (LM’s) prediction on a given task instance by highlighting the text inputs that most influenced the prediction. Ideally, rationale extraction should be faithful (reflective of LM’s actual behavior) and plausible (convincing to humans), without compromising the LM’s (i.e., task model’s) task performance. Although attribution algorithms and select-predict pipelines are commonly used in rationale extraction, they both rely on certain heuristics that hinder them from satisfying all three desiderata. In light of this, we propose UNIREX, a flexible learning framework which generalizes rationale extractor optimization as follows: (1) specify architecture for a learned rationale extractor; (2) select explainability objectives (\ie faithfulness and plausibility criteria); and (3) jointly train the task model and rationale extractor on the task using selected objectives. UNIREX enables replacing prior works’ heuristic design choices with a generic learned rationale extractor in (1) and optimizing it for all three desiderata in (2)-(3). To facilitate comparison between methods w.r.t. multiple desiderata, we introduce the Normalized Relative Gain (NRG) metric. On five English text classification datasets, our best UNIREX configuration outperforms baselines by an average of 32.9% NRG. Plus, UNIREX rationale extractors’ faithfulness can even generalize to unseen datasets and tasks. Aaron Chan, Maziar Sanjabi, Lambert Mathias, Liang Tan 0005, Shaoliang Nie, Xiaochang Peng, Xiang Ren 0001, Hamed Firooz |
ICML | 2 |
| 2022 | Federated Learning with Partial Model PersonalizationabstractWe consider two federated learning algorithms for training partially personalized models, where the shared and personal parameters are updated either simultaneously or alternately on the devices. Both algorithms have been proposed in the literature, but their convergence properties are not fully understood, especially for the alternating variant. We provide convergence analyses of both algorithms in the general nonconvex setting with partial participation and delineate the regime where one dominates the other. Our experiments on real-world image, text, and speech datasets demonstrate that (a) partial personalization can obtain most of the benefits of full model personalization with a small fraction of personal parameters, and, (b) the alternating update algorithm outperforms the simultaneous update algorithm by a small but consistent margin. Krishna Pillutla, Kshitiz Malik, Abdel-rahman Mohamed, Michael G. Rabbat, Maziar Sanjabi, Lin Xiao 0003 |
ICML | 5 |
| 2021 | Alternating Direction Method of Multipliers for QuantizationabstractQuantization of the parameters of machine learning models, such as deep neural networks, requires solving constrained optimization problems, where the constraint set is formed by the Cartesian product of many simple discrete sets. For such optimization problems, we study the performance of the Alternating Direction Method of Multipliers for Quantization (ADMM-Q) algorithm, which is a variant of the widely-used ADMM method applied to our discrete optimization problem. We establish the convergence of the iterates of ADMM-Q to certain stationary points. To the best of our knowledge, this is the first analysis of an ADMM-type method for problems with discrete variables/constraints. Based on our theoretical insights, we develop a few variants of ADMM-Q that can handle inexact update rules, and have improved performance via the use of "soft projection" and "injecting randomness to the algorithm". We empirically evaluate the efficacy of our proposed approaches. Tianjian Huang, Prajwal Singhania, Maziar Sanjabi, Pabitra Mitra, Meisam Razaviyayn |
AISTATS | 3 |
| 2021 | Tilted Empirical Risk Minimization
Tian Li 0005, Ahmad Beirami, Maziar Sanjabi, Virginia Smith |
ICLR | 3 |
| 2020 | Building Placements In Urban Modeling Using Conditional Generative Latent OptimizationabstractGenerating realistic urban environments by scattering or placing buildings on maps is a challenging problem. Unlike the existing procedural methods, we employ a data-driven approach to this problem. We combine two recent advances in machine learning techniques, Generative Latent optimization (GLO) together with adversarial training, to learn a model that can easily generate and place buildings on a given map. Such a model enables its users, particularly artists, to easily generate areas with specific styles, e.g. residential or commercial, just by providing examples. In contrast, traditional procedural methods require lengthy manual tuning of hyper-parameters. Using a more flexible method like ours allows artists to iterate over their designs of urban layouts much faster. Finally, our experiments on real-world data show that our method outperforms state-of-the-art methods in visual quality and can better match the underlying distribution of the building placements. Jingwen Liang, Maziar Sanjabi, Mohsen Sardari, Harold Chaput, Navid Aghdaie, Kazi A. Zaman |
ICIP | 4 |
| 2020 | Fair Resource Allocation in Federated Learning
Tian Li 0005, Maziar Sanjabi, Ahmad Beirami, Virginia Smith |
ICLR | 2 |
| 2019 | Solving a Class of Non-Convex Min-Max Games Using Iterative First Order MethodsabstractRecent applications that arise in machine learning have surged significant interest in solving min-max saddle point games. This problem has been extensively studied in the convex-concave regime for which a global equilibrium solution can be computed efficiently. In this paper, we study the problem in the non-convex regime and show that an $\varepsilon$--first order stationary point of the game can be computed when one of the player’s objective can be optimized to global optimality efficiently. In particular, we first consider the case where the objective of one of the players satisfies the Polyak-{\L}ojasiewicz (PL) condition. For such a game, we show that a simple multi-step gradient descent-ascent algorithm finds an $\varepsilon$--first order stationary point of the problem in $\widetilde{\mathcal{O}}(\varepsilon^{-2})$ iterations. Then we show that our framework can also be applied to the case where the objective of the ``max-player" is concave. In this case, we propose a multi-step gradient descent-ascent algorithm that finds an $\varepsilon$--first order stationary point of the game in $\widetilde{\cal O}(\varepsilon^{-3.5})$ iterations, which is the best known rate in the literature. We applied our algorithm to a fair classification problem of Fashion-MNIST dataset and observed that the proposed algorithm results in smoother training and better generalization. Maher Nouiehed, Maziar Sanjabi, Tianjian Huang, Jason D. Lee, Meisam Razaviyayn |
NeurIPS | 2 |
| 2019 | Multi-theme generative adversarial terrain amplificationabstractAchieving highly detailed terrain models spanning vast areas is crucial to modern computer graphics. The pipeline for obtaining such terrains is via amplification of a low-resolution terrain to refine the details given a desired theme, which is a time-consuming and labor-intensive process. Recently, data-driven methods, such as the sparse construction tree, have provided a promising direction to equip the artist with better control over the theme. These methods learn to amplify terrain details by using an exemplar of high-resolution detailed terrains to transfer the theme. In this paper, we propose Generative Adversarial Terrain Amplification (GATA) that achieves better local/global coherence compared to the existing data-driven methods while providing even more ways to control the theme. GATA is comprised of two key ingredients. Thefi rst one is a novel embedding of themes into vectors of real numbers to achieve a single tool for multi-theme amplification. The theme component can leverage existing LIDAR data to generate similar terrain features. It can also generate newfi ctional themes by tuning the embedding vector or even encoding a new example terrain into an embedding. The second one is an adversarially trained model that, conditioned on an embedding and a low-resolution terrain, generates a high-resolution terrain adhering to the desired theme. The proposed integral approach reduces the need for unnecessary manual adjustments, can speed up the development, and brings the model quality to a new level. Our implementation of the proposed method has proved successful in large-scale terrain authoring for an open-world game. Igor Borovikov, Ahmad Beirami, Maziar Sanjabi, Kazi A. Zaman |
ACM Trans. Graph. | 5 |
| 2018 | On the Convergence and Robustness of Training GANs with Regularized Optimal TransportabstractGenerative Adversarial Networks (GANs) are one of the most practical methods for learning data distributions. A popular GAN formulation is based on the use of Wasserstein distance as a metric between probability distributions. Unfortunately, minimizing the Wasserstein distance between the data distribution and the generative model distribution is a computationally challenging problem as its objective is non-convex, non-smooth, and even hard to compute. In this work, we show that obtaining gradient information of the smoothed Wasserstein GAN formulation, which is based on regularized Optimal Transport (OT), is computationally effortless and hence one can apply first order optimization methods to minimize this objective. Consequently, we establish theoretical convergence guarantee to stationarity for a proposed class of GAN optimization algorithms. Unlike the original non-smooth formulation, our algorithm only requires solving the discriminator to approximate optimality. We apply our method to learning MNIST digits as well as CIFAR-10 images. Our experiments show that our method is computationally efficient and generates images comparable to the state of the art algorithms given the same architecture and computational power. Maziar Sanjabi, Jimmy Ba, Meisam Razaviyayn, Jason D. Lee |
NeurIPS | 1 |
| 2017 | Federated Multi-Task LearningabstractFederated learning poses new statistical and systems challenges in training machine learning models over distributed networks of devices. In this work, we show that multi-task learning is naturally suited to handle the statistical challenges of this setting, and propose a novel systems-aware optimization method, MOCHA, that is robust to practical systems issues. Our method and theory for the first time consider issues of high communication cost, stragglers, and fault tolerance for distributed multi-task learning. The resulting method achieves significant speedups compared to alternatives in the federated setting, as we demonstrate through simulations on real-world federated datasets. Virginia Smith, Chao-Kai Chiang, Maziar Sanjabi, Ameet Talwalkar |
NIPS | 3 |
| 2015 | Accelerated Alternating Direction Method of MultipliersabstractRecent years have seen a revival of interest in the Alternating Direction Method of Multipliers (ADMM), due to its simplicity, versatility, and scalability. As a first order method for general convex problems, the rate of convergence of ADMM is O(1=k) [4, 25]. Given the scale of modern data mining problems, an algorithm with similar properties as ADMM but faster convergence rate can make a big difference in real world applications. In this paper, we introduce the Accelerated Alternating Direction Method of Multipliers (A2DM2) which solves problems with the same structure as ADMM. When the objective function is strongly convex, we show that A2DM2 has a O(1=k2) convergence rate. Unlike related existing literature on trying to accelerate ADMM, our analysis does not need any additional restricting assumptions. Through experiments, we show that A2DM2 converges faster than ADMM on a variety of problems. Further, we illustrate the versatility of the general A2DM2 on the problem of learning to rank, where it is shown to be competitive with the state-of-the-art specialized algorithms for the problem on both scalability and accuracy. Mojtaba Kadkhodaie, Konstantina Christakopoulou, Maziar Sanjabi, Arindam Banerjee 0001 |
KDD | 3 |
| 2012 | Optimal joint base station assignment and downlink beamforming for heterogeneous networksabstractConsider a MIMO heterogeneous network with multiple transmitters (including macro, pico and femto base stations) and many receivers (mobile users). The users are to be assigned to the base stations which then optimize their linear transmit beamformers accordingly. In this work, we consider the problem of joint base station assignment and linear beamformer design to maximize a system wide utility. We first establish the NP-hardness of the resulting optimization problem for a large family of α-fairness utility functions. Then, we propose an efficient algorithm to approximately solve this problem for the special case of sum rate maximization. The simulation results show that the algorithm improves the sum rate. Maziar Sanjabi, Meisam Razaviyayn, Zhi-Quan Luo |
ICASSP | 1 |
| 2012 | Linear Transceiver Design for Interference Alignment: Complexity and ComputationabstractConsider a multiple input-multiple output (MIMO) interference channel where each transmitter and receiver are equipped with multiple antennas. An effective approach to practically achieving high system throughput is to deploy linear transceivers (or beamformers) that can optimally exploit the spatial characteristics of the channel. The recent work of Cadambe and Jafar (IEEE Trans. Inf. Theory, vol. 54, no. 8) suggests that optimal beamformers should maximize the total degrees of freedom and achieve interference alignment in the high signal-to-noise ratio (SNR) regime. In this paper we first consider the interference alignment problem without channel extension and prove that the problem of maximizing the total achieved degrees of freedom for a given MIMO interference channel is NP-hard. Furthermore, we show that even checking the achievability of a given tuple of degrees of freedom for all receivers is NP-hard when each receiver is equipped with at least three antennas. Interestingly, the same problem becomes polynomial time solvable when each transmit/receive node is equipped with no more than two antennas. We also propose a distributed algorithm for transmit covariance matrix design that does not require the DoF tuple preassignment, under the assumption that each receiver uses a linear minimum mean square error (MMSE) beamformer. The simulation results show that the proposed algorithm outperforms the existing interference alignment algorithms in terms of system throughput. Meisam Razaviyayn, Maziar Sanjabi, Zhi-Quan Luo |
IEEE Trans. Inf. Theory | 2 |
| 2011 | Robust SINR-constrained MISO downlink beamforming: When is semidefinite programming relaxation tight?abstractWe consider the robust beamforming problem under imperfect channel state information (CSI) subject to SINR constraints in a downlink multiuser MISO system. One popular approach to solve this nonconvex optimization problem is via semidefinite relaxation (SDR). In this paper, we prove that the SDR method is tight when the channel uncertainty bound is small or when the base station is equipped with two antennas. Enbin Song, Qingjiang Shi, Maziar Sanjabi, Ruoyu Sun 0001, Zhi-Quan Luo |
ICASSP | 3 |