VLDB 2026 Research / reviewers in the wild / expert
Shibal Ibrahim
dblp:177/1113
· DBLP profile ↗
7ranked-venue papers
6as first author
6since 2021 · last 2024
0000-0002-3300-0213ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | End-to-end Feature Selection Approach for Learning Skinny TreesabstractWe propose a new optimization-based approach for feature selection in tree ensembles, an important problem in statistics and machine learning. Popular tree ensemble toolkits e.g., Gradient Boosted Trees and Random Forests support feature selection post-training based on feature importance scores, while very popular, they are known to have drawbacks. We propose Skinny Trees: an end-to-end toolkit for feature selection in tree ensembles where we train a tree ensemble while controlling the number of selected features. Our optimization-based approach learns an ensemble of differentiable trees, and simultaneously performs feature selection using a grouped $\ell_0$-regularizer. We use first-order methods for optimization and present convergence guarantees for our approach. We use a dense-to-sparse regularization scheduling scheme that can lead to more expressive and sparser tree ensembles. On 15 synthetic and real-world datasets, Skinny Trees can achieve $1.5{\times}$–$620{\times}$ feature compression rates, leading up to $10{\times}$ faster inference over dense trees, without any loss in performance. Skinny Trees lead to superior feature selection than many existing toolkits e.g., in terms of AUC performance for 25% feature budget, Skinny Trees outperforms LightGBM by 10.2% (up to 37.7%), and Random Forests by 3% (up to 12.5%). Shibal Ibrahim, Kayhan Behdin, Rahul Mazumder |
AISTATS | 1 |
| 2024 | OSSCAR: One-Shot Structured Pruning in Vision and Language Models with Combinatorial OptimizationabstractStructured pruning is a promising approach for reducing the inference costs of large vision and language models. By removing carefully chosen structures, e.g., neurons or attention heads, the improvements from this approach can be realized on standard deep learning hardware. In this work, we focus on structured pruning in the one-shot (post-training) setting, which does not require model retraining after pruning. We propose a novel combinatorial optimization framework for this problem, based on a layer-wise reconstruction objective and a careful reformulation that allows for scalable optimization. Moreover, we design a new local combinatorial optimization algorithm, which exploits low-rank updates for efficient local search. Our framework is time and memory-efficient and considerably improves upon state-of-the-art one-shot methods on vision models (e.g., ResNet50, MobileNet) and language models (e.g., OPT-1.3B – OPT-30B). For language models, e.g., OPT-2.7B, OSSCAR can lead to $125\times$ lower test perplexity on WikiText with $2\times$ inference time speedup in comparison to the state-of-the-art ZipLM approach. Our framework is also $6\times$ – $8\times$ faster. Notably, our work considers models with tens of billions of parameters, which is up to $100\times$ larger than what has been previously considered in the structured pruning literature. Our code is available at https://github.com/mazumder-lab/OSSCAR. Xiang Meng 0004, Shibal Ibrahim, Kayhan Behdin, Hussein Hazimeh 0001, Natalia Ponomareva 0001, Rahul Mazumder |
ICML | 2 |
| 2023 | COMET: Learning Cardinality Constrained Mixture of Experts with Trees and Local SearchabstractThe sparse Mixture-of-Experts (Sparse-MoE) framework efficiently scales up model capacity in various domains, such as natural language processing and vision. Sparse-MoEs select a subset of the "experts" (thus, only a portion of the overall network) for each input sample using a sparse, trainable gate. Existing sparse gates are prone to convergence and performance issues when training with first-order optimization methods. In this paper, we introduce two improvements to current MoE approaches. First, we propose a new sparse gate: COMET, which relies on a novel tree-based mechanism. COMET is differentiable, can exploit sparsity to speed up computation, and outperforms state-of-the-art gates. Second, due to the challenging combinatorial nature of sparse expert selection, first-order methods are typically prone to low-quality solutions. To deal with this challenge, we propose a novel, permutation-based local search method that can complement first-order methods in training any sparse gate, e.g., Hash routing, Top-k, DSelect-k, and COMET. We show that local search can help networks escape bad initializations or solutions. We performed large-scale experiments on various domains, including recommender systems, vision, and natural language processing. On standard vision and recommender systems benchmarks, COMET+ (COMET with local search) achieves up to 13% improvement in ROC AUC over popular gates, e.g., Hash routing and Top-k, and up to 9% over prior differentiable gates e.g., DSelect-k. When Top-k and Hash gates are combined with local search, we see up to 100X reduction in the budget needed for hyperparameter tuning. Moreover, for language modeling, our approach improves over the state-of-the-art MoEBERT model for distilling BERT on 5/7 GLUE benchmarks as well as SQuAD dataset. Shibal Ibrahim, Wenyu Chen 0003, Hussein Hazimeh 0001, Natalia Ponomareva 0001, Zhe Zhao 0001, Rahul Mazumder |
KDD | 1 |
| 2023 | GRAND-SLAMIN' Interpretable Additive Modeling with Structural ConstraintsabstractGeneralized Additive Models (GAMs) are a family of flexible and interpretable models with old roots in statistics. GAMs are often used with pairwise interactions to improve model accuracy while still retaining flexibility and interpretability but lead to computational challenges as we are dealing with order of $p^2$ terms. It is desirable to restrict the number of components (i.e., encourage sparsity) for easier interpretability, and better computational and statistical properties. Earlier approaches, considering sparse pairwise interactions, have limited scalability, especially when imposing additional structural interpretability constraints. We propose a flexible GRAND-SLAMIN framework that can learn GAMs with interactions under sparsity and additional structural constraints in a differentiable end-to-end fashion. We customize first-order gradient-based optimization to perform sparse backpropagation to exploit sparsity in additive effects for any differentiable loss function in a GPU-compatible manner. Additionally, we establish novel non-asymptotic prediction bounds for our estimators with tree-based shape functions. Numerical experiments on real-world datasets show that our toolkit performs favorably in terms of performance, variable selection and scalability when compared with popular toolkits to fit GAMs with interactions. Our work expands the landscape of interpretable modeling while maintaining prediction accuracy competitive with non-interpretable black-box models. Our code is available at https://github.com/mazumder-lab/grandslamin. Shibal Ibrahim, Gabriel Afriat, Kayhan Behdin, Rahul Mazumder |
NeurIPS | 1 |
| 2022 | Flexible Modeling and Multitask Learning using Differentiable Tree EnsemblesabstractDecision tree ensembles are widely used and competitive learning models. Despite their success, popular toolkits for learning tree ensembles have limited modeling capabilities. For instance, these toolkits support a limited number of loss functions and are restricted to single task learning. We propose a flexible framework for learning tree ensembles, which goes beyond existing toolkits to support arbitrary loss functions, missing responses, and multi-task learning. Our framework builds on differentiable (a.k.a. soft) tree ensembles, which can be trained using first-order methods. However, unlike classical trees, differentiable trees are difficult to scale. We therefore propose a novel tensor-based formulation of differentiable trees that allows for efficient vectorization on GPUs. We introduce FASTEL: a new toolkit (based on Tensorflow 2) for learning differentiable tree ensembles. We perform experiments on a collection of 28 real open-source and proprietary datasets, which demonstrate that our framework can lead to 100x more compact and 23% more expressive tree ensembles than those obtained by popular toolkits. Shibal Ibrahim, Hussein Hazimeh 0001, Rahul Mazumder |
KDD | 1 |
| 2022 | Newer is Not Always Better: Rethinking Transferability Metrics, Their Peculiarities, Stability and Performance
Shibal Ibrahim, Natalia Ponomareva 0001, Rahul Mazumder |
ECML/PKDD (1) | 1 |
| 2015 | Networked PFC drive for SPPSCIM air conditioners to enable elastic load operation in constrained power systemsabstractLarge HVAC systems are good candidates for elastic loads in demand side management optimization of smarter grid infrastructures. There is an increaseing trend of using Split-type Air Conditioners (ACs) in residential, commercial and industrial setups for fine granularity temperature control. The traditional control in each unit is through duty-cycle based hysteretic method which is simple but inefficient and requires power spikes at every start-up of the compressor. In developing countries with weak grids, people rely on backup sources like generators and UPS. Large power surge at each turn-on inflicts huge pressure on these constrained resources. Limited work has been reported that aims at energy optimization and controlled operation of Split-type ACs constructed using single phase permanent split capacitor induction machine (SPPSCIM). Our work targets operational optimization, efficiency enhancement and increased load handling capability of contrained power systems in setups employing multiple AC units. This is done through networked Variable Frequency Drives (VFD) for conventional ACs which allow elimination of surge current at turn-on and tighter set-point control while drawing sinusoidal current near unity power factor from the constrained source. Furthermore, the networked operation allows integration of existing installed base of standalone units through Building Management System (BMS) with the central control infrastructure. The prototype networked VFD is tested on a 2.3kW Mitsubishi Electric Split-type AC through Wi-Fi and the results are presented to confirm the expected results in network mode of operation. Shibal Ibrahim, Nauman Ahmad Zaffar |
IECON | 1 |