VLDB 2026 Research / reviewers in the wild / expert
Shahana Ibrahim
dblp:243/9707
· DBLP profile ↗
13ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0003-1951-5234ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multi-label Recognition under Noisy Supervision: A Confusion Mixture Modeling ApproachabstractMulti-label recognition is a critical task in artificial intelligence, aiming to identify every object present in an image. Designing a multi-label classifier is a nontrivial task both from data collection and modeling perspectives. Collecting multiple labels for each image is extremely time-consuming and costly, which often leads to noisy annotations. Furthermore, a robust classifier that performs reliably well in the presence of such noisy labels demands meticulous modeling and learning criterion design. In this work, we propose a novel probabilistic confusion model that effectively incorporates inter-label interactions in causing label noise. The proposed model incorporates a latent variable, building a hierarchical structure to the label noise generation, and represents the label noise as a mixture of confusions caused by various classes. Under the proposed multi-label confusion mixture (MCM) model, we design an end-to-end learning criterion along with a sparsity regularization, that effectively estimates the true multi-label classifier. Experiments with various real-world datasets showcase the effectiveness of our approach. Diego Linares Gonzalez, Shahana Ibrahim |
ICASSP | 2 |
| 2025 | Under-Counted Matrix Completion Without Detection FeaturesabstractUnder-counted matrix completion (UC-MC) has many important applications, especially in epidemiology and ecology where the observed data are often smaller than the actual numbers. Existing works model the under-counting effects using entry-wise miss detection probabilities, which are usually formulated as functions of detection-related side information or features (e.g., weather conditions for observing a certain species). However, such features are not always available. This work proposes a model for UC-MC that circumvents using such side information. By assuming that the detection probabilities for a large proportion of entries are similar, the under-counting probabilities are approximated by a specially structured (i.e., rank-one plus sparse) matrix. This way, a structure-regularized UC-MC formulation is attained, and no detection-related side information is used. A re-parameterization-based implementation is proposed, allowing us to employ off-the-shelf gradient-based optimizers to tackle the loss function of interest. Simulations are used to demonstrate the effectiveness of the proposed approach. Tri Nguyen 0004, Shahana Ibrahim, Rebecca A. Hutchinson, Xiao Fu 0001 |
ICASSP | 2 |
| 2025 | Robust Multi-Label Learning with Human-Guided and Foundation Model-Aided Crowd FrameworkabstractMulti-label learning has emerged as a critical task in artificial intelligence (AI) for understanding data across diverse modalities. However, a significant challenge in this domain is the acquisition of accurate labels, which is often both time-consuming and resource-intensive. Assigning multiple labels to each data instance typically requires input from multiple annotators, each bringing their own expertise or mistakes. Recent advancements in foundation models have enabled the use of pseudo-labels to supplement human annotations, but these models are often not primarily designed for multi-label tasks, introducing additional label noise. In this work, we present a novel crowd framework for multi-label learning that integrates hybrid collaboration between human annotators and foundation models. By combining their responses in a robust manner and leveraging insights from modeling and factorization techniques, the proposed framework is accompanied by a regularized end-to-end learning criterion. Experiments using several real-world datasets showcase the promise of our framework. Faizul Rakib Sayem, Shahana Ibrahim |
ICIP | 2 |
| 2024 | Noisy Label Learning with Instance-Dependent Outliers: Identifiability via Crowd WisdomabstractThe generation of label noise is often modeled as a process involving a probability transition matrix (also interpreted as the _annotator confusion matrix_) imposed onto the label distribution. Under this model, learning the ``ground-truth classifier''---i.e., the classifier that can be learned if no noise was present---and the confusion matrix boils down to a model identification problem. Prior works along this line demonstrated appealing empirical performance, yet identifiability of the model was mostly established by assuming an instance-invariant confusion matrix. Having an (occasionally) instance-dependent confusion matrix across data samples is apparently more realistic, but inevitably introduces outliers to the model. Our interest lies in confusion matrix-based noisy label learning with such outliers taken into consideration. We begin with pointing out that under the model of interest, using labels produced by only one annotator is fundamentally insufficient to detect the outliers or identify the ground-truth classifier. Then, we prove that by employing a crowdsourcing strategy involving multiple annotators, a carefully designed loss function can establish the desired model identifiability under reasonable conditions. Our development builds upon a link between the noisy label model and a column-corrupted matrix factorization mode---based on which we show that crowdsourced annotations distinguish nominal data and instance-dependent outliers using a low-dimensional subspace. Experiments show that our learning scheme substantially improves outlier detection and the classifier's testing accuracy. Tri Nguyen 0004, Shahana Ibrahim, Xiao Fu 0001 |
NeurIPS | 2 |
| 2023 | Deep Learning From Crowdsourced Labels: Coupled Cross-Entropy Minimization, Identifiability, and Regularization
Shahana Ibrahim, Tri Nguyen 0004, Xiao Fu 0001 |
ICLR | 1 |
| 2023 | Under-Counted Tensor Completion with Neural Incorporation of AttributesabstractSystematic under-counting effects are observed in data collected across many disciplines, e.g., epidemiology and ecology. Under-counted tensor completion (UC-TC) is well-motivated for many data analytics tasks, e.g., inferring the case numbers of infectious diseases at unobserved locations from under-counted case numbers in neighboring regions. However, existing methods for similar problems often lack supports in theory, making it hard to understand the underlying principles and conditions beyond empirical successes. In this work, a low-rank Poisson tensor model with an expressive unknown nonlinear side information extractor is proposed for under-counted multi-aspect data. A joint low-rank tensor completion and neural network learning algorithm is designed to recover the model. Moreover, the UC-TC formulation is supported by theoretical analysis showing that the fully counted entries of the tensor and each entry's under-counting probability can be provably recovered from partial observations---under reasonable conditions. To our best knowledge, the result is the first to offer theoretical supports for under-counted multi-aspect data completion. Simulations and real-data experiments corroborate the theoretical claims. Shahana Ibrahim, Xiao Fu 0001, Rebecca A. Hutchinson, Eugene Seo 0001 |
ICML | 1 |
| 2023 | Deep Clustering with Incomplete Noisy Pairwise Annotations: A Geometric Regularization ApproachabstractThe recent integration of deep learning and pairwise similarity annotation-based constrained clustering—i.e., deep constrained clustering (DCC)—has proven effective for incorporating weak supervision into massive data clustering: Less than 1% of pair similarity annotations can often substantially enhance the clustering accuracy. However, beyond empirical successes, there is a lack of understanding of DCC. In addition, many DCC paradigms are sensitive to annotation noise, but performance-guaranteed noisy DCC methods have been largely elusive. This work first takes a deep look into a recently emerged logistic loss function of DCC, and characterizes its theoretical properties. Our result shows that the logistic DCC loss ensures the identifiability of data membership under reasonable conditions, which may shed light on its effectiveness in practice. Building upon this understanding, a new loss function based on geometric factor analysis is proposed to fend against noisy annotations. It is shown that even under unknown annotation confusions, the data membership can still be provably identified under our proposed learning criterion. The proposed approach is tested over multiple datasets to validate our claims. Tri Nguyen 0004, Shahana Ibrahim, Xiao Fu 0001 |
ICML | 2 |
| 2021 | Learning Mixed Membership from Adjacency Graph Via Systematic Edge Query: Identifiability and AlgorithmabstractGraph clustering is a core technique for network analysis problems, e.g., community detection. This work puts forth a node clustering approach for largely incomplete adjacency graphs. Under the considered scenario, instead of having access to the complete graph, only a small amount of queries about the graph edges can be made for node clustering. This task is well-motivated in many large-scale network analysis problems, where complete graph acquisition is prohibitively costly. Prior work tackles this problem under the setting that the nodes only admit single membership and the clusters are disjoint, yet multiple membership nodes and overlapping clusters often arise in practice. Existing approaches also rely on random edge query patterns and convex optimization-based formulations, which give rise to a number of implementation and scalability challenges. This work offers a framework that provably learns the mixed membership of nodes from overlapping clusters using limited edge information. Our method is equipped with a systematic edge query pattern, which is arguably easier to implement relative to the random counterparts in certain applications, e.g., field survey based graph analysis. A lightweight scalable algorithm is proposed, and its performance characterizations are presented. Numerical experiments are used to showcase the effectiveness of our method. Shahana Ibrahim, Xiao Fu 0001 |
ICASSP | 1 |
| 2021 | Fiber-Sampled Stochastic Mirror Descent for Tensor Decomposition with β-DivergenceabstractCanonical polyadic decomposition (CPD) has been a workhorse for multimodal data analytics. This work puts forth a stochastic algorithmic framework for CPD under β-divergence, which is well-motivated in statistical learning—where the Euclidean distance is typically not preferred. Despite the existence of a series of prior works addressing this topic, pressing computational and theoretical challenges, e.g., scalability and convergence issues, still remain. In this paper, a unified stochastic mirror descent framework is developed for large-scale β-divergence CPD. Our key contribution is the integrated design of a tensor fiber sampling strategy and a flexible stochastic Bregman divergence-based mirror descent iterative procedure, which significantly reduces the computation and memory cost per iteration for various β. Leveraging the fiber sampling scheme and the multilinear algebraic structure of low-rank tensors, the proposed lightweight algorithm also ensures global convergence to a stationary point under mild conditions. Numerical results on synthetic and real data show that our framework attains significant computational saving compared with state-of-the-art methods. Wenqiang Pu, Shahana Ibrahim, Xiao Fu 0001, Mingyi Hong 0001 |
ICASSP | 2 |
| 2021 | Crowdsourcing via Annotator Co-occurrence Imputation and Provable Symmetric Nonnegative Matrix FactorizationabstractUnsupervised learning of the Dawid-Skene (D&S) model from noisy, incomplete and crowdsourced annotations has been a long-standing challenge, and is a critical step towards reliably labeling massive data. A recent work takes a coupled nonnegative matrix factorization (CNMF) perspective, and shows appealing features: It ensures the identifiability of the D&S model and enjoys low sample complexity, as only the estimates of the co-occurrences of annotator labels are involved. However, the identifiability holds only when certain somewhat restrictive conditions are met in the context of crowdsourcing. Optimizing the CNMF criterion is also costly—and convergence assurances are elusive. This work recasts the pairwise co-occurrence based D&S model learning problem as a symmetric NMF (SymNMF) problem—which offers enhanced identifiability relative to CNMF. In practice, the SymNMF model is often (largely) incomplete, due to the lack of co-labeled items by some annotators. Two lightweight algorithms are proposed for co-occurrence imputation. Then, a low-complexity shifted rectified linear unit (ReLU)-empowered SymNMF algorithm is proposed to identify the D&S model. Various performance characterizations (e.g., missing co-occurrence recoverability, stability, and convergence) and evaluations are also presented. Shahana Ibrahim, Xiao Fu 0001 |
ICML | 1 |
| 2020 | On Recoverability of Randomly Compressed Tensors With Low CP RankabstractOur interest lies in the recoverability properties of compressed tensors under the canonical polyadic decomposition (CPD) model. The considered problem is well-motivated in many applications, e.g., hyperspectral image and video compression. Prior work studied this problem under a variety of assumptions, e.g., that the latent factors of the tensor are sparse and that the compressing matrix follows a joint absolutely continuous distribution. These results leverage analytical tools such as CPD uniqueness and algebraic geometry-which are elegant. In this work, we offer an alternative result: We show that if the tensor is compressed by a subgaussian linear mapping, then the tensor is recoverable if the number of measurements is on the same order of magnitude as that of the model parameters. Unlike existing results, our proof is based on deriving a restricted isometry property (R.I.P.) under the CPD model via set covering techniques, and thus exhibits a flavor of classic compressive sensing. The new recoverability result enriches the understanding to the compressed CP tensor recovery problem. It offers theoretical guarantees for recovering tensors whose elements are not necessarily sparse; the compressing matrix is also not restricted to continuous matrices under our framework. The newly derived covering number for tensors with low CP rank may also benefit future research, e.g., sketching based tensor compression for reducing computational burden. Shahana Ibrahim, Xiao Fu 0001, Xingguo Li |
IEEE Signal Process. Lett. | 1 |
| 2019 | Crowdsourcing via Pairwise Co-occurrences: Identifiability and AlgorithmsabstractThe data deluge comes with high demands for data labeling. Crowdsourcing (or, more generally, ensemble learning) techniques aim to produce accurate labels via integrating noisy, non-expert labeling from annotators. The classic Dawid-Skene estimator and its accompanying expectation maximization (EM) algorithm have been widely used, but the theoretical properties are not fully understood. Tensor methods were proposed to guarantee identification of the Dawid-Skene model, but the sample complexity is a hurdle for applying such approaches---since the tensor methods hinge on the availability of third-order statistics that are hard to reliably estimate given limited data. In this paper, we propose a framework using pairwise co-occurrences of the annotator responses, which naturally admits lower sample complexity. We show that the approach can identify the Dawid-Skene model under realistic conditions. We propose an algebraic algorithm reminiscent of convex geometry-based structured matrix factorization to solve the model identification problem efficiently, and an identifiability-enhanced algorithm for handling more challenging and critical scenarios. Experiments show that the proposed algorithms outperform the state-of-art algorithms under a variety of scenarios. Shahana Ibrahim, Xiao Fu 0001, Nikos Kargas, Kejun Huang |
NeurIPS | 1 |
| 2019 | Estimating Phase Duration for SPaT MessagesabstractA signal phase and timing (SPaT) message describes the current phase at a signalized intersection for each lane, together with an estimate of the residual time of that phase. Accurate SPaT messages can be used to construct a speed profile for a vehicle that reduces its fuel consumption as it approaches or leaves an intersection. This paper presents SPaT estimation algorithms at an intersection with a semi-actuated signal, using real-time signal phase measurements. The algorithms are evaluated using high-resolution data from an intersection in Montgomery County, MD, USA. The algorithms can be readily implemented at signal controllers. This paper supports three findings. First, real-time information dramatically improves the accuracy of the prediction of the residual time compared with the prediction based on historical data alone. Second, as time increases, the prediction of the residual time may increase or decrease. Third, as drivers differently weight errors in predicting “end of green” and “end of red,” drivers on two different approaches may prefer different estimates of the residual time of the same phase. Shahana Ibrahim, Dileep M. Kalathil, René Osorio Sanchez, Pravin Varaiya |
IEEE Trans. Intell. Transp. Syst. | 1 |