Abdul-Saboor Sheikh

dblp:22/8697 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Probabilistic and Bayesian machine learning · 32% Time series and sequential data · 24% Generative modeling · 24%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
sparse coding
0.742016
Select-and-Sample for Spike-and-Slab Sparse Coding · NIPS 2016
A truncated EM approach for spike-and-slab sparse coding · J. Mach. Learn. Res. 2014
Why MCA? Nonlinear sparse coding with spike-and-slab prior for neurally plausible image encoding · NIPS 2012
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference
0.532016
Select-and-Sample for Spike-and-Slab Sparse Coding · NIPS 2016
Why MCA? Nonlinear sparse coding with spike-and-slab prior for neurally plausible image encoding · NIPS 2012
Select and Sample - A Model of Efficient Neural Inference and Learning · NIPS 2011
Machine learning › Generative modeling › normalizing flow
conditional normalizing flow
0.512021
Multivariate Probabilistic Time Series Forecasting via Conditioned Normalizing Flows · ICLR 2021
Machine learning › Time series and sequential data › time series analysis › time series forecasting
multivariate time series forecasting
0.512021
Multivariate Probabilistic Time Series Forecasting via Conditioned Normalizing Flows · ICLR 2021
Machine learning › Generative modeling
normalizing flow
0.512021
Multivariate Probabilistic Time Series Forecasting via Conditioned Normalizing Flows · ICLR 2021
Machine learning › Time series and sequential data › time series modeling
probabilistic time series forecasting
0.512021
Multivariate Probabilistic Time Series Forecasting via Conditioned Normalizing Flows · ICLR 2021
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › sparse bayesian learning
spike-and-slab prior
0.322014
A truncated EM approach for spike-and-slab sparse coding · J. Mach. Learn. Res. 2014
Why MCA? Nonlinear sparse coding with spike-and-slab prior for neurally plausible image encoding · NIPS 2012
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods › markov chain monte carlo
gibbs sampling
0.212016
Select-and-Sample for Spike-and-Slab Sparse Coding · NIPS 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization
0.212014
A truncated EM approach for spike-and-slab sparse coding · J. Mach. Learn. Res. 2014

Methods — techniques the papers use, named apart from their topics

normalizing flow · 0.5gibbs sampling · 0.4select-and-sample · 0.2latent subspace selection · 0.2variational inference · 0.1maximal causes analysis · 0.1variational approximation · 0.1markov chain monte carlo · 0.1expectation-maximization · 0.1
YearPublicationVenuePosition
2021 Multivariate Probabilistic Time Series Forecasting via Conditioned Normalizing Flows
Kashif Rasul, Abdul-Saboor Sheikh, Ingmar Schuster, Urs M. Bergmann, Roland Vollgraf
ICLR2
2020 Meta-Learning for Size and Fit Recommendation in Fashion
abstract
Fashion e-commerce has enjoyed an exponential growth in the last few years. A key challenge of the market players is to offer customers a personalized experience and to suggest relevant articles. In that respect, although product recommendation is a well-studied field, size and fit recommendation is still in its infancy. The size and fit topic is a very challenging problem as data is extremely sparse and noisy. Most approaches so far have exploited traditional machine learning techniques. In this work, we bring forward a meta-learning approach using an underlying deep neural network. The advantage of such an approach lies in its ability to exploit large scale data, learn across fashion categories, and absorb new data efficiently without re-training. We benchmark our method against 3 recent methods proven successful in the domain, and demonstrate various strengths of the proposed approach. To that end, we use a large-scale anonymized dataset of about 9.4 million customer-size interactions, collected over 5 years from around 384k customers.
Julia Lasserre, Abdul-Saboor Sheikh, Evgenii Koriagin, Urs Bergmann, Roland Vollgraf, Reza Shirvany
SDM2
2019 A deep learning system for predicting size and fit in fashion e-commerce
abstract
Personalized size and fit recommendations bear crucial significance for any fashion e-commerce platform. Predicting the correct fit drives customer satisfaction and benefits the business by reducing costs incurred due to size-related returns. Traditional collaborative filtering algorithms seek to model customer preferences based on their previous orders. A typical challenge for such methods stems from extreme sparsity of customer-article orders. To alleviate this problem, we propose a deep learning based content-collaborative methodology for personalized size and fit recommendation. Our proposed method can ingest arbitrary customer and article data and can model multiple individuals or intents behind a single account. The method optimizes a global set of parameters to learn population-level abstractions of size and fit relevant information from observed customer-article interactions. It further employs customer and article specific embedding variables to learn their properties. Together with learned entity embeddings, the method maps additional customer and article attributes into a latent space to derive personalized recommendations. Application of our method to two publicly available datasets demonstrate an improvement over the state-of-the-art published results. On two proprietary datasets, one containing fit feedback from fashion experts and the other involving customer purchases, we further outperform comparable methodologies, including a recent Bayesian approach for size recommendation.
Abdul-Saboor Sheikh, Romain Guigourès, Evgenii Koriagin, Yuen King Ho, Reza Shirvany, Roland Vollgraf, Urs Bergmann
RecSys1
2019 STRFs in primary auditory cortex emerge from masking-based statistics of natural sounds
abstract
We investigate how the neural processing in auditory cortex is shaped by the statistics of natural sounds. Hypothesising that auditory cortex (A1) represents the structural primitives out of which sounds are composed, we employ a statistical model to extract such components. The input to the model are cochleagrams which approximate the non-linear transformations a sound undergoes from the outer ear, through the cochlea to the auditory nerve. Cochleagram components do not superimpose linearly, but rather according to a rule which can be approximated using the max function. This is a consequence of the compression inherent in the cochleagram and the sparsity of natural sounds. Furthermore, cochleagrams do not have negative values. Cochleagrams are therefore not matched well by the assumptions of standard linear approaches such as sparse coding or ICA. We therefore consider a new encoding approach for natural sounds, which combines a model of early auditory processing with maximal causes analysis (MCA), a sparse coding model which captures both the non-linear combination rule and non-negativity of the data. An efficient truncated EM algorithm is used to fit the MCA model to cochleagram data. We characterize the generative fields (GFs) inferred by MCA with respect to in vivo neural responses in A1 by applying reverse correlation to estimate spectro-temporal receptive fields (STRFs) implied by the learned GFs. Despite the GFs being non-negative, the STRF estimates are found to contain both positive and negative subfields, where the negative subfields can be attributed to explaining away effects as captured by the applied inference method. A direct comparison with ferret A1 shows many similar forms, and the spectral and temporal modulation tuning of both ferret and model STRFs show similar ranges over the population. In summary, our model represents an alternative to linear approaches for biological auditory encoding while it captures salient data properties and links inhibitory subfields to explaining away effects.
Abdul-Saboor Sheikh, Nicol S. Harper, Jakob Drefs, Yosef Singer, Zhenwen Dai, Richard E. Turner, Jörg Lücke
PLoS Comput. Biol.1
2018 A hierarchical bayesian model for size recommendation in fashion
abstract
We introduce a hierarchical Bayesian approach to tackle the challenging problem of size recommendation in e-commerce fashion. Our approach jointly models a size purchased by a customer, and its possible return event: 1. no return, 2. returned too small 3. returned too big. Those events are drawn following a multinomial distribution parameterized on the joint probability of each event, built following a hierarchy combining priors. Such a model allows us to incorporate extended domain expertise and article characteristics as prior knowledge, which in turn makes it possible for the underlying parameters to emerge thanks to sufficient data. Experiments are presented on real (anonymized) data from millions of customers along with a detailed discussion on the efficiency of such an approach within a large scale production system.
Romain Guigourès, Yuen King Ho, Evgenii Koriagin, Abdul-Saboor Sheikh, Urs Bergmann, Reza Shirvany
RecSys4
2018 Neural Simpletrons: Learning in the Limit of Few Labels with Directed Generative Networks
abstract
We explore classifier training for data sets with very few labels. We investigate this task using a neural network for nonnegative data. The network is derived from a hierarchical normalized Poisson mixture model with one observed and two hidden layers. With the single objective of likelihood optimization, both labeled and unlabeled data are naturally incorporated into learning. The neural activation and learning equations resulting from our derivation are concise and local. As a consequence, the network can be scaled using standard deep learning tools for parallelized GPU implementation. Using standard benchmarks for nonnegative data, such as text document representations, MNIST, and NIST SD19, we study the classification performance when very few labels are used for training. In different settings, the network's performance is compared to standard and recently suggested semisupervised classifiers. While other recent approaches are more competitive for many labels or fully labeled data sets, we find that the network studied here can be applied to numbers of few labels where no other system has been reported to operate so far.
Dennis Forster, Abdul-Saboor Sheikh, Jörg Lücke
Neural Comput.2
2016 Select-and-Sample for Spike-and-Slab Sparse Coding
abstract
Probabilistic inference serves as a popular model for neural processing. It is still unclear, however, how approximate probabilistic inference can be accurate and scalable to very high-dimensional continuous latent spaces. Especially as typical posteriors for sensory data can be expected to exhibit complex latent dependencies including multiple modes. Here, we study an approach that can efficiently be scaled while maintaining a richly structured posterior approximation under these conditions. As example model we use spike-and-slab sparse coding for V1 processing, and combine latent subspace selection with Gibbs sampling (select-and-sample). Unlike factored variational approaches, the method can maintain large numbers of posterior modes and complex latent dependencies. Unlike pure sampling, the method is scalable to very high-dimensional latent spaces. Among all sparse coding approaches with non-trivial posterior approximations (MAP or ICA-like models), we report the largest-scale results. In applications we firstly verify the approach by showing competitiveness in standard denoising benchmarks. Secondly, we use its scalability to, for the first time, study highly-overcomplete settings for V1 encoding using sophisticated posterior representations. More generally, our study shows that very accurate probabilistic inference for multi-modal posteriors with complex dependencies is tractable, functionally desirable and consistent with models for neural inference.
Abdul-Saboor Sheikh, Jörg Lücke
NIPS1
2014 A truncated EM approach for spike-and-slab sparse coding
Abdul-Saboor Sheikh, Jacquelyn Shelton, Jörg Lücke
J. Mach. Learn. Res.1
2012 Why MCA? Nonlinear sparse coding with spike-and-slab prior for neurally plausible image encoding
abstract
Modelling natural images with sparse coding (SC) has faced two main challenges: flexibly representing varying pixel intensities and realistically representing low- level image components. This paper proposes a novel multiple-cause generative model of low-level image statistics that generalizes the standard SC model in two crucial points: (1) it uses a spike-and-slab prior distribution for a more realistic representation of component absence/intensity, and (2) the model uses the highly nonlinear combination rule of maximal causes analysis (MCA) instead of a lin- ear combination. The major challenge is parameter optimization because a model with either (1) or (2) results in strongly multimodal posteriors. We show for the first time that a model combining both improvements can be trained efficiently while retaining the rich structure of the posteriors. We design an exact piece- wise Gibbs sampling method and combine this with a variational method based on preselection of latent dimensions. This combined training scheme tackles both analytical and computational intractability and enables application of the model to a large number of observed and hidden dimensions. Applying the model to image patches we study the optimal encoding of images by simple cells in V1 and compare the model’s predictions with in vivo neural recordings. In contrast to standard SC, we find that the optimal prior favors asymmetric and bimodal ac- tivity of simple cells. Testing our model for consistency we find that the average posterior is approximately equal to the prior. Furthermore, we find that the model predicts a high percentage of globular receptive fields alongside Gabor-like fields. Similarly high percentages are observed in vivo. Our results thus argue in favor of improvements of the standard sparse coding model for simple cells by using flexible priors and nonlinear combinations.
Jacquelyn Shelton, Philip Sterne, Jörg Bornschein, Abdul-Saboor Sheikh, Jörg Lücke
NIPS4
2011 Select and Sample - A Model of Efficient Neural Inference and Learning
abstract
An increasing number of experimental studies indicate that perception encodes a posterior probability distribution over possible causes of sensory stimuli, which is used to act close to optimally in the environment. One outstanding difficulty with this hypothesis is that the exact posterior will in general be too complex to be represented directly, and thus neurons will have to represent an approximation of this distribution. Two influential proposals of efficient posterior representation by neural populations are: 1) neural activity represents samples of the underlying distribution, or 2) they represent a parametric representation of a variational approximation of the posterior. We show that these approaches can be combined for an inference scheme that retains the advantages of both: it is able to represent multiple modes and arbitrary correlations, a feature of sampling methods, and it reduces the represented space to regions of high probability mass, a strength of variational approximations. Neurally, the combined method can be interpreted as a feed-forward preselection of the relevant state space, followed by a neural dynamics implementation of Markov Chain Monte Carlo (MCMC) to approximate the posterior over the relevant states. We demonstrate the effectiveness and efficiency of this approach on a sparse coding model. In numerical experiments on artificial data and image patches, we compare the performance of the algorithms to that of exact EM, variational state space selection alone, MCMC alone, and the combined select and sample approach. The select and sample approach integrates the advantages of the sampling and variational approximations, and forms a robust, neurally plausible, and very efficient model of processing and learning in cortical networks. For sparse coding we show applications easily exceeding a thousand observed and a thousand hidden dimensions.
Jacquelyn Shelton, Jörg Bornschein, Abdul-Saboor Sheikh, Pietro Berkes, Jörg Lücke
NIPS3
2010 Structured literature image finder: Parsing text and figures in biomedical literature
Amr Ahmed 0001, Andrew Arnold, Luís Pedro Coelho, Joshua D. Kangas, Abdul-Saboor Sheikh, Eric P. Xing, William W. Cohen, Robert F. Murphy
J. Web Semant.5