VLDB 2026 Research / reviewers in the wild / expert
Novi Quadrianto
dblp:06/580
· DBLP profile ↗
39ranked-venue papers
14as first author
8since 2021 · last 2026
0000-0001-8819-306XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 14 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
29 papers |
Trustworthy machine learning · 53% Probabilistic and Bayesian machine learning · 16% Efficient and distributed learning · 6% | |
| Databases, data mining, and information retrieval
9 papers |
Data mining · 76% Information retrieval · 18% Data integration and cleaning · 6% | |
| Theoretical computer science
7 papers |
Computational geometry · 47% Mathematical optimization · 27% Graph algorithms and graph theory · 21% |
Topics — the 30 heaviest of 84, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning
fairness |
3.1 | 5 | 2026 | Safe Fairness Guarantees Without Demographics in Classification: Spectral Uncertainty Set Perspective · IEEE Trans. Pattern Anal. Mach. Intell. 2026 Revisiting (Un)Fairness in Recourse by Minimizing Worst-Case Social Burden · AAAI 2026 Null-Sampling for Interpretable and Fair Representations · ECCV (26) 2020 |
Machine learning › Trustworthy machine learning
robustness |
1.8 | 3 | 2026 | Safe Fairness Guarantees Without Demographics in Classification: Spectral Uncertainty Set Perspective · IEEE Trans. Pattern Anal. Mach. Intell. 2026 Are Compressed Language Models Less Subgroup Robust? · EMNLP 2023 Okapi: Generalising Better by Making Statistical Matches Match · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › interpretability › counterfactual explanation
algorithmic recourse |
1.0 | 1 | 2026 | Revisiting (Un)Fairness in Recourse by Minimizing Worst-Case Social Burden · AAAI 2026 |
Machine learning › Trustworthy machine learning › robustness
distributionally robust optimization |
1.0 | 1 | 2026 | Safe Fairness Guarantees Without Demographics in Classification: Spectral Uncertainty Set Perspective · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Trustworthy machine learning › fairness › algorithmic fairness
fairness without sensitive attributes |
1.0 | 1 | 2026 | Safe Fairness Guarantees Without Demographics in Classification: Spectral Uncertainty Set Perspective · IEEE Trans. Pattern Anal. Mach. Intell. 2026 |
Machine learning › Trustworthy machine learning › fairness
fair representation learning |
0.8 | 2 | 2020 | Null-Sampling for Interpretable and Fair Representations · ECCV (26) 2020 Discovering Fair Representations in the Data Domain · CVPR 2019 |
Machine learning › Efficient and distributed learning
model compression |
0.7 | 1 | 2023 | Are Compressed Language Models Less Subgroup Robust? · EMNLP 2023 |
Machine learning › Trustworthy machine learning › robustness › distributional robustness
subgroup robustness |
0.7 | 1 | 2023 | Are Compressed Language Models Less Subgroup Robust? · EMNLP 2023 |
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization |
0.6 | 1 | 2022 | Okapi: Generalising Better by Making Statistical Matches Match · NeurIPS 2022 |
Machine learning › Efficient and distributed learning › model reuse
model patching |
0.6 | 1 | 2022 | RealPatch: A Statistical Matching Framework for Model Patching with Real Samples · ECCV (25) 2022 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › low-resource domain adaptation
semi-supervised domain adaptation |
0.6 | 1 | 2022 | Okapi: Generalising Better by Making Statistical Matches Match · NeurIPS 2022 |
Data mining
clustering |
0.5 | 2 | 2017 | Composing Tree Graphical Models with Persistent Homology Features for Clustering Mixed-Type Data · ICML 2017 Clustering High Dimensional Categorical Data via Topographical Features · ICML 2016 |
Machine learning › Learning paradigms › supervised learning
learning using privileged information |
0.5 | 3 | 2017 | Learning from the Mistakes of Others: Matching Errors in Cross-Dataset Learning · CVPR 2016 Learning to Rank Using Privileged Information · ICCV 2013 Recycling Privileged Learning and Distribution Matching for Fairness · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
discrete latent variable model |
0.4 | 1 | 2020 | Low-Variance Black-Box Gradient Estimates for the Plackett-Luce Distribution · AAAI 2020 |
Machine learning › Trustworthy machine learning
interpretability |
0.4 | 1 | 2020 | Null-Sampling for Interpretable and Fair Representations · ECCV (26) 2020 |
Machine learning › Trustworthy machine learning › interpretability › explainable AI
interpretable representation |
0.4 | 1 | 2020 | Null-Sampling for Interpretable and Fair Representations · ECCV (26) 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.4 | 1 | 2020 | Low-Variance Black-Box Gradient Estimates for the Plackett-Luce Distribution · AAAI 2020 |
Machine learning › Optimization for machine learning
stochastic gradient methods |
0.4 | 1 | 2020 | Low-Variance Black-Box Gradient Estimates for the Plackett-Luce Distribution · AAAI 2020 |
Machine learning › Optimization for machine learning
variance reduction |
0.4 | 1 | 2020 | Low-Variance Black-Box Gradient Estimates for the Plackett-Luce Distribution · AAAI 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.3 | 1 | 2017 | Composing Tree Graphical Models with Persistent Homology Features for Clustering Mixed-Type Data · ICML 2017 |
Machine learning › Trustworthy machine learning › fairness › group fairness › subgroup fairness
multi-attribute fairness |
0.3 | 1 | 2017 | Recycling Privileged Learning and Distribution Matching for Fairness · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
tree-structured graphical models |
0.3 | 1 | 2017 | Composing Tree Graphical Models with Persistent Homology Features for Clustering Mixed-Type Data · ICML 2017 |
Data mining › clustering
mixed data clustering |
0.3 | 1 | 2017 | Composing Tree Graphical Models with Persistent Homology Features for Clustering Mixed-Type Data · ICML 2017 |
Computational geometry › topological data analysis
persistent homology |
0.3 | 1 | 2017 | Composing Tree Graphical Models with Persistent Homology Features for Clustering Mixed-Type Data · ICML 2017 |
Computational geometry
topological data analysis |
0.3 | 1 | 2017 | Composing Tree Graphical Models with Persistent Homology Features for Clustering Mixed-Type Data · ICML 2017 |
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process |
0.3 | 2 | 2014 | Scalable Gaussian Process Structured Prediction for Grid Factor Graph Applications · ICML 2014 Kernel Conditional Quantile Estimation via Reduction Revisited · ICDM 2009 |
Machine learning › Trustworthy machine learning
annotator disagreement |
0.2 | 1 | 2016 | Ambiguity Helps: Classification with Disagreements in Crowdsourced Annotations · CVPR 2016 |
Machine learning › Transfer learning and domain adaptation › cross-domain learning
cross-dataset learning |
0.2 | 1 | 2016 | Learning from the Mistakes of Others: Matching Errors in Cross-Dataset Learning · CVPR 2016 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
feature selection |
0.2 | 1 | 2016 | Learning Using Unselected Features (LUFe) · IJCAI 2016 |
Computer vision › Image recognition and object detection
image classification |
0.2 | 1 | 2016 | Ambiguity Helps: Classification with Disagreements in Crowdsourced Annotations · CVPR 2016 |
Methods — techniques the papers use, named apart from their topics
theoretical characterization · 1.0spectral uncertainty set · 1.0minimax optimization · 1.0fourier feature mapping · 1.0privileged information · 0.8worst-group performance analysis · 0.7BERT · 0.7distribution matching · 0.6statistical matching · 0.6persistent homology · 0.6density estimation · 0.6consistency loss · 0.6topological data analysis · 0.5lagrangian relaxation · 0.3integer linear programming · 0.3sampling · 0.3graphical models · 0.3graphical model · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting (Un)Fairness in Recourse by Minimizing Worst-Case Social BurdenabstractMachine learning based predictions are increasingly used in sensitive decision-making applications that directly affect our lives. This has led to extensive research into ensuring the fairness of classifiers. Beyond just fair classification, emerging legislation now mandates that when a classifier delivers a negative decision, it must also offer actionable steps an individual can take to reverse that outcome. This concept is known as algorithmic recourse. Nevertheless, many researchers have expressed concerns about the fairness guarantees within the recourse process itself. In this work, we provide a theoretical characterization of unfairness in algorithmic recourse, formally linking fairness guarantees in recourse and classification, and highlighting limitations of the standard equal cost paradigm. We then introduce a novel fairness framework based on social burden, along with a practical algorithm (MISOB), broadly applicable under real-world conditions. Empirical results on real-world datasets show that MISOB reduces the social burden across all groups without compromising overall classifier accuracy. Ainhize Barrainkua, Giovanni De Toni, José Antonio Lozano 0001, Novi Quadrianto |
AAAI | 4 |
| 2026 | Safe Fairness Guarantees Without Demographics in Classification: Spectral Uncertainty Set PerspectiveabstractAs automated classification systems become increasingly prevalent, concerns have emerged over their potential to reinforce and amplify existing societal biases. In the light of this issue, many methods have been proposed to enhance the fairness guarantees of classifiers. Most of the existing interventions assume access to group information for all instances, a requirement rarely met in practice. Fairness without access to demographic information has often been approached through robust optimization techniques, which target worst-case outcomes over a set of plausible distributions known as the uncertainty set. However, their effectiveness is strongly influenced by the chosen uncertainty set. In fact, existing approaches often overemphasize outliers or overly pessimistic scenarios, compromising both overall performance and fairness. To overcome these limitations, we introduce SPECTRE, a minimax-fair method that adjusts the spectrum of a simple Fourier feature mapping and constrains the extent to which the worst-case distribution can deviate from the empirical distribution. We perform extensive experiments on the American Community Survey datasets involving 20 states. The safeness of SPECTRE comes as it provides the highest average values on fairness guarantees together with the smallest interquartile range in comparison to state-of-the-art approaches, even compared to those with access to demographic group information. In addition, we provide a theoretical analysis that derives computable bounds on the worst-case error for both individual groups and the overall population, as well as characterizes the worst-case distributions responsible for these extremal performances. Ainhize Barrainkua, Santiago Mazuelas, Novi Quadrianto, José Antonio Lozano 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Uncertainty Matters: Stable Conclusions under Unstable Assessment of Fairness ResultsabstractRecent studies highlight the effectiveness of Bayesian methods in assessing algorithm performance, particularly in fairness and bias evaluation. We present Uncertainty Matters, a multi-objective uncertainty-aware algorithmic comparison framework. In fairness focused scenarios, it models sensitive group confusion matrices using Bayesian updates and facilitates joint comparison of performance (e.g., accuracy) and fairness metrics (e.g., true positive rate parity). Our approach works seamlessly with common evaluation methods like K-fold cross-validation, effectively addressing dependencies among the K posterior metric distributions. The integration of correlated information is carried out through a procedure tailored to the classifier’s complexity. Experiments demonstrate that the insights derived from algorithmic comparisons employing the Uncertainty Matters approach are more informative, reliable, and less influenced by particular data partitions. Code for the paper is publicly available at \url{https://github.com/abarrainkua/UncertaintyMatters}. Ainhize Barrainkua, Paula Gordaliza, José Antonio Lozano 0001, Novi Quadrianto |
AISTATS | 4 |
| 2023 | Are Compressed Language Models Less Subgroup Robust?abstractTo reduce the inference cost of large language models, model compression is increasingly used to create smaller scalable models.However, little is known about their robustness to minority subgroups defined by the labels and attributes of a dataset.In this paper, we investigate the effects of 18 different compression methods and settings on the subgroup robustness of BERT language models.We show that worst-group performance does not depend on model size alone, but also on the compression method used.Additionally, we find that model compression does not always worsen the performance on minority subgroups.Altogether, our analysis serves to further research into the subgroup robustness of model compression. Leonidas Gee, Andrea Zugarini, Novi Quadrianto |
EMNLP | 3 |
| 2023 | Water Physics Aware Semantic Segmentation through Texture-Biased U-Net ArchitecturesabstractReliably identifying water bodies is an important step in automating the identification of potable water. This work therefore investigates water scene segmentation, with a focus on water’s physical properties which give it features that distinguish it from other elements. We propose a physics-aware water segmentation method, in which we adapt both a U-Net and MACUNet model so that they are biased towards texture information by using a Gabor convolutional layer as the first layer, combined with a mixture of average and maximum pooling layers in the encoder. To train the networks, a dataset of annotated water images was created, comprising water bodies from diverse light, atmospheric, geographic, and environmental conditions. We show that the physics-aware, texture-biased models result in effective water segmentation. We then test the texture-biased models using 3 standard aerial scene segmentation benchmarks and show that in all cases they outperform the standard U-Net or MACUNet models. We suggest this is because the new models are sensitive to small variations in texture, meaning they can extract information from scenes affected by light, canopy or shadows. Georgios Voulgaris, Andrew Philippides, Novi Quadrianto |
IGARSS | 3 |
| 2022 | RealPatch: A Statistical Matching Framework for Model Patching with Real Samples
Sara Romiti, Christopher Inskip, Viktoriia Sharmanska, Novi Quadrianto |
ECCV (25) | 4 |
| 2022 | Deep Learning Robustness to Domain Shifts During Seasonal VariationsabstractIn certain geographic locations like South Asia, the landscape changes dramatically between dry and wet seasons. The main factor responsible for this variation is the flora that trans-forms the landscape between seasons. These transformations can affect the performance of deep learning models trained to analyse satellite images, especially if there are domain shifts between training and testing data distributions. The current work shows that an architecture which employs a Gabor convolutional layer as the first layer of a deep network input fo-cuses on more salient parts of the image than one which uses a standard convolutional layer meaning that removing colour information is less damaging than for the standard network. Further we show that the proposed architecture is robust in the presence of domain shifts due to seasonal data variations. Georgios Voulgaris, Andrew Philippides, Novi Quadrianto |
IGARSS | 3 |
| 2022 | Okapi: Generalising Better by Making Statistical Matches MatchabstractWe propose Okapi, a simple, efficient, and general method for robust semi-supervised learning based on online statistical matching. Our method uses a nearest-neighbours-based matching procedure to generate cross-domain views for a consistency loss, while eliminating statistical outliers. In order to perform the online matching in a runtime- and memory-efficient way, we draw upon the self-supervised literature and combine a memory bank with a slow-moving momentum encoder. The consistency loss is applied within the feature space, rather than on the predictive distribution, making the method agnostic to both the modality and the task in question. We experiment on the WILDS 2.0 datasets Sagawa et al., which significantly expands the range of modalities, applications, and shifts available for studying and benchmarking real-world unsupervised adaptation. Contrary to Sagawa et al., we show that it is in fact possible to leverage additional unlabelled data to improve upon empirical risk minimisation (ERM) results with the right method. Our method outperforms the baseline methods in terms of out-of-distribution (OOD) generalisation on the iWildCam (a multi-class classification task) and PovertyMap (a regression task) image datasets as well as the CivilComments (a binary classification task) text dataset. Furthermore, from a qualitative perspective, we show the matches obtained from the learned encoder are strongly semantically related. Code for our paper is publicly available at https://github.com/wearepal/okapi/. Myles Bartlett, Sara Romiti, Viktoriia Sharmanska, Novi Quadrianto |
NeurIPS | 4 |
| 2020 | Low-Variance Black-Box Gradient Estimates for the Plackett-Luce DistributionabstractLearning models with discrete latent variables using stochastic gradient descent remains a challenge due to the high variance of gradient estimates. Modern variance reduction techniques mostly consider categorical distributions and have limited applicability when the number of possible outcomes becomes large. In this work, we consider models with latent permutations and propose control variates for the Plackett-Luce distribution. In particular, the control variates allow us to optimize black-box functions over permutations using stochastic gradient descent. To illustrate the approach, we consider a variety of causal structure learning tasks for continuous and discrete data. We show that our method outperforms competitive relaxation-based optimization methods and is also applicable to non-differentiable score functions. Artyom Gadetsky, Kirill Struminsky, Christopher Robinson, Novi Quadrianto, Dmitry P. Vetrov |
AAAI | 4 |
| 2020 | Null-Sampling for Interpretable and Fair Representations
Thomas Kehrenberg, Myles Bartlett, Oliver Thomas 0002, Novi Quadrianto |
ECCV (26) | 4 |
| 2019 | Quantification under class-conditional dataset shiftabstractQuantification is the estimation of class proportions in a dataset. It is important in a wide range of fields, such as the proportion of positive reviews in sentiment analysis or the age and gender distribution of respondents in market research. David Spence, Christopher Inskip, Novi Quadrianto, David Weir |
ASONAM | 3 |
| 2019 | Discovering Fair Representations in the Data DomainabstractInterpretability and fairness are critical in computer vision and machine learning applications, in particular when dealing with human outcomes, e.g. inviting or not inviting for a job interview based on application materials that may include photographs. One promising direction to achieve fairness is by learning data representations that remove the semantics of protected characteristics, and are therefore able to mitigate unfair outcomes. All available models however learn latent embeddings which comes at the cost of being uninterpretable. We propose to cast this problem as data-to-data translation, i.e. learning a mapping from an input domain to a fair target domain, where a fairness definition is being enforced. Here the data domain can be images, or any tabular data representation. This task would be straightforward if we had fair target data available, but this is not the case. To overcome this, we learn a highly unconstrained mapping by exploiting statistics of residuals -- the difference between input data and its translated version -- and the protected characteristics. When applied to the CelebA dataset of face images with gender attribute as the protected characteristic, our model enforces equality of opportunity by adjusting the eyes and lips regions. Intriguingly, on the same dataset we arrive at similar conclusions when using semantic attribute representations of images for translation. On face images of the recent DiF dataset, with the same gender attribute, our method adjusts nose regions. In the Adult income dataset, also with protected gender attribute, our model achieves equality of opportunity by, among others, obfuscating the wife and husband relationship. Analyzing those systematic changes will allow us to scrutinize the interplay of fairness criterion, chosen protected characteristics, and prediction performance. Novi Quadrianto, Viktoriia Sharmanska, Oliver Thomas 0002 |
CVPR | 1 |
| 2017 | Gray-box Inference for Structured Gaussian Process ModelsabstractWe develop an automated variational inference method for Bayesian structured prediction problems with Gaussian process (GP) priors and linear-chain likelihoods. Our approach does not need to know the details of the structured likelihood model and can scale up to a large number of observations. Furthermore, we show that the required expected likelihood term and its gradients in the variational objective (ELBO) can be estimated efficiently by using expectations over very low-dimensional Gaussian distributions. Optimization of the ELBO is fully parallelizable over sequences and amenable to stochastic optimization, which we use along with control variate techniques to make our framework useful in practice. Results on a set of natural language processing tasks show that our method can be as good as (and sometimes better than, in particular with respect to expected log-likelihood) hard-coded approaches including SVM-struct and CRF, and overcomes the scalability limitations of previous inference algorithms based on sampling. Overall, this is a fundamental step to developing automated inference methods for Bayesian structured prediction. Pietro Galliani, Amir Dezfouli, Edwin V. Bonilla, Novi Quadrianto |
AISTATS | 4 |
| 2017 | Composing Tree Graphical Models with Persistent Homology Features for Clustering Mixed-Type DataabstractClustering data with both continuous and discrete attributes is a challenging task. Existing methods lack a principled probabilistic formulation. In this paper, we propose a clustering method based on a tree-structured graphical model to describe the generation process of mixed-type data. Our tree-structured model factorized into a product of pairwise interactions, and thus localizes the interaction between feature variables of different types. To provide a robust clustering method based on the tree-model, we adopt a topographical view and compute peaks of the density function and their attractive basins for clustering. Furthermore, we leverage the theory from topology data analysis to adaptively merge trivial peaks into large ones in order to achieve meaningful clusterings. Our method outperforms state-of-the-art methods on mixed-type data. Xiuyan Ni, Novi Quadrianto, Yusu Wang 0001, Chao Chen 0012 |
ICML | 2 |
| 2017 | Recycling Privileged Learning and Distribution Matching for FairnessabstractEquipping machine learning models with ethical and legal constraints is a serious issue; without this, the future of machine learning is at risk. This paper takes a step forward in this direction and focuses on ensuring machine learning models deliver fair decisions. In legal scholarships, the notion of fairness itself is evolving and multi-faceted. We set an overarching goal to develop a unified machine learning framework that is able to handle any definitions of fairness, their combinations, and also new definitions that might be stipulated in the future. To achieve our goal, we recycle two well-established machine learning techniques, privileged learning and distribution matching, and harmonize them for satisfying multi-faceted fairness definitions. We consider protected characteristics such as race and gender as privileged information that is available at training but not at test time; this accelerates model training and delivers fairness through unawareness. Further, we cast demographic parity, equalized odds, and equality of opportunity as a classical two-sample problem of conditional distributions, which can be solved in a general form by using distance measures in Hilbert Space. We show several existing models are special cases of ours. Finally, we advocate returning the Pareto frontier of multi-objective minimization of error and unfairness in predictions. This will facilitate decision makers to select an operating point and to be accountable for it. Novi Quadrianto, Viktoriia Sharmanska |
NIPS | 1 |
| 2016 | Ambiguity Helps: Classification with Disagreements in Crowdsourced AnnotationsabstractImagine we show an image to a person and ask her/him to decide whether the scene in the image is warm or not warm, and whether it is easy or not to spot a squirrel in the image. For exactly the same image, the answers to those questions are likely to differ from person to person. This is because the task is inherently ambiguous. Such an ambiguous, therefore challenging, task is pushing the boundary of computer vision in showing what can and can not be learned from visual data. Crowdsourcing has been invaluable for collecting annotations. This is particularly so for a task that goes beyond a clear-cut dichotomy as multiple human judgments per image are needed to reach a consensus. This paper makes conceptual and technical contributions. On the conceptual side, we define disagreements among annotators as privileged information about the data instance. On the technical side, we propose a framework to incorporate annotation disagreements into the classifiers. The proposed framework is simple, relatively fast, and outperforms classifiers that do not take into account the disagreements, especially if tested on high confidence annotations. Viktoriia Sharmanska, Daniel Hernández-Lobato, José Miguel Hernández-Lobato, Novi Quadrianto |
CVPR | 4 |
| 2016 | Learning from the Mistakes of Others: Matching Errors in Cross-Dataset LearningabstractCan we learn about object classes in images by looking at a collection of relevant 3D models? Or if we want to learn about human (inter-)actions in images, can we benefit from videos or abstract illustrations that show these actions? A common aspect of these settings is the availability of additional or privileged data that can be exploited at training time and that will not be available and not of interest at test time. We seek to generalize the learning with privileged information (LUPI) framework, which requires additional information to be defined per image, to the setting where additional information is a data collection about the task of interest. Our framework minimizes the distribution mismatch between errors made in images and in privileged data. The proposed method is tested on four publicly available datasets: Image+ClipArt, Image+3Dobject, and Image+ Video. Experimental results reveal that our new LUPI paradigm naturally addresses the cross-dataset learning. Viktoriia Sharmanska, Novi Quadrianto |
CVPR | 2 |
| 2016 | Clustering High Dimensional Categorical Data via Topographical FeaturesabstractAnalysis of categorical data is a challenging task. In this paper, we propose to compute topographical features of high-dimensional categorical data. We propose an efficient algorithm to extract modes of the underlying distribution and their attractive basins. These topographical features provide a geometric view of the data and can be applied to visualization and clustering of real world challenging datasets. Experiments show that our principled method outperforms state-of-the-art clustering methods while also admits an embarrassingly parallel property. Chao Chen 0012, Novi Quadrianto |
ICML | 2 |
| 2016 | Learning Using Unselected Features (LUFe)
Joseph G. Taylor, Viktoriia Sharmanska, Kristian Kersting, David Weir, Novi Quadrianto |
IJCAI | 5 |
| 2015 | GPstruct: Bayesian Structured Prediction Using Gaussian ProcessesabstractWe introduce a conceptually novel structured prediction model, GPstruct, which is kernelized, non-parametric and Bayesian, by design. We motivate the model with respect to existing approaches, among others, conditional random fields (CRFs), maximum margin Markov networks (M3N), and structured support vector machines (SVMstruct), which embody only a subset of its properties. We present an inference procedure based on Markov Chain Monte Carlo. The framework can be instantiated for a wide range of structured objects such as linear chains, trees, grids, and other general graphs. As a proof of concept, the model is benchmarked on several natural language processing tasks and a video gesture segmentation task involving a linear chain structure. We show prediction accuracies for GPstruct which are comparable to or exceeding those of CRFs and SVMstruct. Sébastien Bratières, Novi Quadrianto, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2015 | A Very Simple Safe-Bayesian Random ForestabstractRandom forests works by averaging several predictions of de-correlated trees. We show a conceptually radical approach to generate a random forest: random sampling of many trees from a prior distribution, and subsequently performing a weighted ensemble of predictive probabilities. Our approach uses priors that allow sampling of decision trees even before looking at the data, and a power likelihood that explores the space spanned by combination of decision trees. While each tree performs Bayesian inference to compute its predictions, our aggregation procedure uses the power likelihood rather than the likelihood and is therefore strictly speaking not Bayesian. Nonetheless, we refer to it as a Bayesian random forest but with a built-in safety. The safeness comes as it has good predictive performance even if the underlying probabilistic model is wrong. We demonstrate empirically that our Safe-Bayesian random forest outperforms MCMC or SMC based Bayesian decision trees in term of speed and accuracy, and achieves competitive performance to entropy or Gini optimised random forest, yet is very simple to construct. Novi Quadrianto, Zoubin Ghahramani |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Scalable Gaussian Process Structured Prediction for Grid Factor Graph ApplicationsabstractStructured prediction is an important and well studied problem with many applications across machine learning. GPstruct is a recently proposed structured prediction model that offers appealing properties such as being kernelised, non-parametric, and supporting Bayesian inference (Bratières et al. 2013). The model places a Gaussian process prior over energy functions which describe relationships between input variables and structured output variables. However, the memory demand of GPstruct is quadratic in the number of latent variables and training runtime scales cubically. This prevents GPstruct from being applied to problems involving grid factor graphs, which are prevalent in computer vision and spatial statistics applications. Here we explore a scalable approach to learning GPstruct models based on ensemble learning, with weak learners (predictors) trained on subsets of the latent variables and bootstrap data, which can easily be distributed. We show experiments with 4M latent variables on image segmentation. Our method outperforms widely-used conditional random field models trained with pseudo-likelihood. Moreover, in image segmentation problems it improves over recent state-of-the-art marginal optimisation methods in terms of predictive performance and uncertainty calibration. Finally, it generalises well on all training set sizes. Sébastien Bratières, Novi Quadrianto, Sebastian Nowozin, Zoubin Ghahramani |
ICML | 2 |
| 2014 | Mind the Nuisance: Gaussian Process Classification using Privileged Noise
Daniel Hernández-Lobato, Viktoriia Sharmanska, Kristian Kersting, Christoph H. Lampert, Novi Quadrianto |
NIPS | 5 |
| 2013 | Learning to Rank Using Privileged InformationabstractMany computer vision problems have an asymmetric distribution of information between training and test time. In this work, we study the case where we are given additional information about the training data, which however will not be available at test time. This situation is called learning using privileged information (LUPI). We introduce two maximum-margin techniques that are able to make use of this additional source of information, and we show that the framework is applicable to several scenarios that have been studied in computer vision before. Experiments with attributes, bounding boxes, image tags and rationales as additional information in object classification show promising results. Viktoriia Sharmanska, Novi Quadrianto, Christoph H. Lampert |
ICCV | 2 |
| 2013 | The Supervised IBP: Neighbourhood Preserving Infinite Latent Feature Models
Novi Quadrianto, Viktoriia Sharmanska, David A. Knowles, Zoubin Ghahramani |
UAI | 1 |
| 2012 | Beyond Dataset Bias: Multi-task Unaligned Shared Knowledge Transfer
Tatiana Tommasi, Novi Quadrianto, Barbara Caputo, Christoph H. Lampert |
ACCV (1) | 2 |
| 2012 | Augmented Attribute Representations
Viktoriia Sharmanska, Novi Quadrianto, Christoph H. Lampert |
ECCV (5) | 2 |
| 2012 | The Most Persistent Soft-Clique in a Set of Sampled Graphs
Novi Quadrianto, Chao Chen 0012, Christoph H. Lampert |
ICML | 1 |
| 2011 | Learning Multi-View Neighborhood Preserving Projections
Novi Quadrianto, Christoph H. Lampert |
ICML | 1 |
| 2010 | Optimal Web-Scale Tiering as a Flow ProblemabstractWe present a fast online solver for large scale maximum-flow problems as they occur in portfolio optimization, inventory management, computer vision, and logistics. Our algorithm solves an integer linear program in an online fashion. It exploits total unimodularity of the constraint matrix and a Lagrangian relaxation to solve the problem as a convex online game. The algorithm generates approximate solutions of max-flow problems by performing stochastic gradient descent on a set of flows. We apply the algorithm to optimize tier arrangement of over 80 Million web pages on a layered set of caches to serve an incoming query stream optimally. We provide an empirical demonstration of the effectiveness of our method on real query-pages data. Gilbert Leung, Novi Quadrianto, Alexander J. Smola, Kostas Tsioutsiouliklis |
NIPS | 2 |
| 2010 | Multitask Learning without Label CorrespondencesabstractWe propose an algorithm to perform multitask learning where each task has potentially distinct label sets and label correspondences are not readily available. This is in contrast with existing methods which either assume that the label sets shared by different tasks are the same or that there exists a label mapping oracle. Our method directly maximizes the mutual information among the labels, and we show that the resulting objective function can be efficiently optimized using existing algorithms. Our proposed approach has a direct application for data integration with different label spaces for the purpose of classification, such as integrating Yahoo! and DMOZ web directories. Novi Quadrianto, Alexander J. Smola, Tibério S. Caetano, S. V. N. Vishwanathan, James Petterson |
NIPS | 1 |
| 2010 | Kernelized SortingabstractObject matching is a fundamental operation in data analysis. It typically requires the definition of a similarity measure between the classes of objects to be matched. Instead, we develop an approach which is able to perform matching by requiring a similarity measure only within each of the classes. This is achieved by maximizing the dependency between matched pairs of observations by means of the Hilbert-Schmidt Independence Criterion. This problem can be cast as one of maximizing a quadratic assignment problem with special structure and we present a simple algorithm for finding a locally optimal solution. Novi Quadrianto, Alexander J. Smola, Tinne Tuytelaars |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2009 | Learning based automatic face annotation for arbitrary poses and expressions from frontal images onlyabstractStatistical approaches for building non-rigid deformable models, such as the active appearance model (AAM), have enjoyed great popularity in recent years, but typically require tedious manual annotation of training images. In this paper, a learning based approach for the automatic annotation of visually deformable objects from a single annotated frontal image is presented and demonstrated on the example of automatically annotating face images that can be used for building AAMs for fitting and tracking. This approach employs the idea of initially learning the correspondences between landmarks in a frontal image and a set of training images with a face in arbitrary poses. Using this learner, virtual images of unseen faces at any arbitrary pose for which the learner was trained can be reconstructed by predicting the new landmark locations and warping the texture from the frontal image. View-based AAMs are then built from the virtual images and used for automatically annotating unseen images, including images of different facial expressions, at any random pose within the maximum range spanned by the virtually reconstructed images. The approach is experimentally validated by automatically annotating face images from three different databases. Akshay Asthana, Roland Göcke, Novi Quadrianto, Tom Gedeon |
CVPR | 3 |
| 2009 | Kernel Conditional Quantile Estimation via Reduction RevisitedabstractQuantile regression refers to the process of estimating the quantiles of a conditional distribution and has many important applications within econometrics and data mining, among other domains. In this paper, we show how to estimate these conditional quantile functions within a Bayes risk minimization framework using a Gaussian process prior. The resulting non-parametric probabilistic model is easy to implement and allows non-crossing quantile functions to be enforced. Moreover, it can directly be used in combination with tools and extensions of standard Gaussian processes such as principled hyperparameter estimation, sparsification, and quantile regression with input-dependent noise rates. No existing approach enjoys all of these desirable properties. Experiments on benchmark datasets show that our method is competitive with state-of-the-art approaches. Novi Quadrianto, Kristian Kersting, Mark D. Reid, Tibério S. Caetano, Wray L. Buntine |
ICDM | 1 |
| 2009 | Convex Relaxation of Mixture Regression with Efficient AlgorithmsabstractWe develop a convex relaxation of maximum a posteriori estimation of a mixture of regression models. Although our relaxation involves a semidefinite matrix variable, we reformulate the problem to eliminate the need for general semidefinite programming. In particular, we provide two reformulations that admit fast algorithms. The first is a max-min spectral reformulation exploiting quasi-Newton descent. The second is a min-min reformulation consisting of fast alternating steps of closed-form updates. We evaluate the methods against Expectation-Maximization in a real problem of motion segmentation from video data. Novi Quadrianto, Tibério S. Caetano, John Lim, Dale Schuurmans |
NIPS | 1 |
| 2009 | Distribution Matching for TransductionabstractMany transductive inference algorithms assume that distributions over training and test estimates should be related, e.g. by providing a large margin of separation on both sets. We use this idea to design a transduction algorithm which can be used without modification for classification, regression, and structured estimation. At its heart we exploit the fact that for a good learner the distributions over the outputs on training and test sets should match. This is a classical two-sample problem which can be solved efficiently in its most general form by using distance measures in Hilbert Space. It turns out that a number of existing heuristics can be viewed as special cases of our approach. Novi Quadrianto, James Petterson, Alexander J. Smola |
NIPS | 1 |
| 2009 | Estimating Labels from Label Proportions
Novi Quadrianto, Alexander J. Smola, Tibério S. Caetano, Quoc V. Le |
J. Mach. Learn. Res. | 1 |
| 2008 | Estimating labels from label proportionsabstractConsider the following problem: given sets of unlabeled observations, each set with known label proportions, predict the labels of another set of observations, also with known label proportions. This problem appears in areas like e-commerce, spam filtering and improper content detection. We present consistent estimators which can reconstruct the correct labels with high probability in a uniform convergence sense. Experiments show that our method works well in practice. Novi Quadrianto, Alexander J. Smola, Tibério S. Caetano, Quoc V. Le |
ICML | 1 |
| 2008 | Kernelized SortingabstractObject matching is a fundamental operation in data analysis. It typically requires the definition of a similarity measure between the classes of objects to be matched. Instead, we develop an approach which is able to perform matching by requiring a similarity measure only within each of the classes. This is achieved by maximizing the dependency between matched pairs of observations by means of the Hilbert Schmidt Independence Criterion. This problem can be cast as one of maximizing a quadratic assignment problem with special structure and we present a simple algorithm for finding a locally optimal solution. Novi Quadrianto, Alexander J. Smola |
NIPS | 1 |