VLDB 2026 Research / reviewers in the wild / expert
Anton Osokin
dblp:08/8344
· DBLP profile ↗
20ranked-venue papers
7as first author
1since 2021 · last 2021
0000-0002-8807-5132ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 7 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
15 papers |
Probabilistic and Bayesian machine learning · 28% Learning theory · 18% Image recognition and object detection · 15% | |
| Databases, data mining, and information retrieval
2 papers |
Data models and query languages · 46% Knowledge graphs · 46% Data mining · 8% | |
| Theoretical computer science
3 papers |
Mathematical optimization · 80% Graph algorithms and graph theory · 20% |
Topics — the 30 heaviest of 46, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field |
0.6 | 3 | 2015 | Submodular Relaxation for Inference in Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015 Putting MRFs on a Tensor Train · ICML 2014 A Principled Deep Random Field Model for Image Segmentation · CVPR 2013 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.5 | 2 | 2017 | On Structured Prediction Theory with Calibrated Convex Surrogate Losses · NIPS 2017 Minding the Gaps for Block Frank-Wolfe Optimization of Structured SVMs · ICML 2016 |
Data models and query languages › natural language interface
natural language query translation |
0.5 | 1 | 2021 | SPARQLing Database Queries from Intermediate Question Decompositions · EMNLP (1) 2021 |
Knowledge graphs › knowledge graph querying
SPARQL query generation |
0.5 | 1 | 2021 | SPARQLing Database Queries from Intermediate Question Decompositions · EMNLP (1) 2021 |
Computer vision › Image recognition and object detection › object detection
anchor-based detection |
0.4 | 1 | 2020 | OS2D: One-Stage One-Shot Object Detection by Matching Anchor Features · ECCV (15) 2020 |
Computer vision › Image recognition and object detection › object detection › few-shot object detection
one-shot object detection |
0.4 | 1 | 2020 | OS2D: One-Stage One-Shot Object Detection by Matching Anchor Features · ECCV (15) 2020 |
Machine learning › Trustworthy machine learning › uncertainty estimation › probability calibration
calibration analysis |
0.3 | 1 | 2018 | Quantifying Learning Guarantees for Convex but Inconsistent Surrogates · NeurIPS 2018 |
Machine learning › Learning theory › loss function › surrogate loss
consistency of surrogate losses |
0.3 | 1 | 2018 | Quantifying Learning Guarantees for Convex but Inconsistent Surrogates · NeurIPS 2018 |
Machine learning › Learning theory › loss function › surrogate loss
convex surrogate loss |
0.3 | 1 | 2018 | Quantifying Learning Guarantees for Convex but Inconsistent Surrogates · NeurIPS 2018 |
Machine learning › Learning theory
loss function |
0.3 | 1 | 2018 | SEARNN: Training RNNs with global-local losses · ICLR (Poster) 2018 |
Machine learning › Deep learning architectures and training
recurrent neural network |
0.3 | 1 | 2018 | SEARNN: Training RNNs with global-local losses · ICLR (Poster) 2018 |
Machine learning › Deep learning architectures and training › recurrent neural network
recurrent neural network training |
0.3 | 1 | 2018 | SEARNN: Training RNNs with global-local losses · ICLR (Poster) 2018 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
MAP inference |
0.3 | 2 | 2014 | Putting MRFs on a Tensor Train · ICML 2014 Submodular decomposition framework for inference in associative Markov networks with global constraints · CVPR 2011 |
Machine learning › Learning theory › loss function
calibrated surrogate |
0.3 | 1 | 2017 | On Structured Prediction Theory with Calibrated Convex Surrogate Losses · NIPS 2017 |
Machine learning › Generative modeling
generative adversarial network |
0.3 | 1 | 2017 | GANs for Biological Image Synthesis · ICCV 2017 |
Machine learning › Learning theory
statistical learning theory |
0.3 | 1 | 2017 | On Structured Prediction Theory with Calibrated Convex Surrogate Losses · NIPS 2017 |
Mathematical optimization
discrete optimization |
0.3 | 2 | 2012 | Minimizing Sparse High-Order Energies by Submodular Vertex-Cover · NIPS 2012 Submodular decomposition framework for inference in associative Markov networks with global constraints · CVPR 2011 |
Machine learning › Optimization for machine learning › constrained optimization
frank-wolfe algorithm |
0.2 | 1 | 2016 | Minding the Gaps for Block Frank-Wolfe Optimization of Structured SVMs · ICML 2016 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
structured SVM |
0.2 | 1 | 2016 | Minding the Gaps for Block Frank-Wolfe Optimization of Structured SVMs · ICML 2016 |
Computer vision › Image recognition and object detection › object detection
contextual reasoning |
0.2 | 1 | 2015 | Context-Aware CNNs for Person Head Detection · ICCV 2015 |
Computer vision › Image recognition and object detection › object detection › object part detection
head detection |
0.2 | 1 | 2015 | Context-Aware CNNs for Person Head Detection · ICCV 2015 |
Machine learning › Optimization for machine learning › energy minimization
markov random field optimization |
0.2 | 1 | 2015 | Submodular Relaxation for Inference in Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Machine learning › Efficient and distributed learning
model compression |
0.2 | 1 | 2015 | Tensorizing Neural Networks · NIPS 2015 |
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
tensor-train decomposition |
0.2 | 1 | 2015 | Tensorizing Neural Networks · NIPS 2015 |
Mathematical optimization
lagrangian relaxation |
0.2 | 1 | 2015 | Submodular Relaxation for Inference in Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Mathematical optimization › submodular optimization
submodular relaxation |
0.2 | 1 | 2015 | Submodular Relaxation for Inference in Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.2 | 1 | 2014 | Putting MRFs on a Tensor Train · ICML 2014 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
graphical model inference |
0.2 | 1 | 2014 | Putting MRFs on a Tensor Train · ICML 2014 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
partition function estimation |
0.2 | 1 | 2014 | Putting MRFs on a Tensor Train · ICML 2014 |
Machine learning › Deep learning architectures and training › loss function design
perceptual loss |
0.2 | 1 | 2014 | Perceptually Inspired Layout-Aware Losses for Image Segmentation · ECCV (2) 2014 |
Methods — techniques the papers use, named apart from their topics
latent space interpolation · 0.9generative adversarial network · 0.9neural semantic parsing · 0.5intermediate question decomposition · 0.5one-stage detection · 0.4feature matching · 0.4tensor-train decomposition · 0.4global-local losses · 0.3convex surrogate analysis · 0.3calibration function lower bound · 0.3submodular optimization · 0.3message passing · 0.3stochastic gradient descent · 0.3convex surrogate loss · 0.3submodular energy minimization · 0.2lagrangian relaxation · 0.2dual decomposition · 0.2submodular decomposition · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | SPARQLing Database Queries from Intermediate Question DecompositionsabstractTo translate natural language questions into executable database queries, most approaches rely on a fully annotated training set.Annotating a large dataset with queries is difficult as it requires query-language expertise.We reduce this burden using grounded in databases intermediate question representations.These representations are simpler to collect and were originally crowdsourced within the Break dataset (Wolfson et al., 2020).Our pipeline consists of two parts: a neural semantic parser that converts natural language questions into the intermediate representations and a non-trainable transpiler to the SPARQL query language (a standard language for accessing knowledge graphs and semantic web).We chose SPARQL because its queries are structurally closer to our intermediate representations (compared to SQL).We observe that the execution accuracy of queries constructed by our model on the challenging Spider dataset is comparable with the state-of-the-art text-to-SQL methods trained with annotated SQL queries.Our code and data are publicly available.1 Irina Saparina, Anton Osokin |
EMNLP (1) | 2 |
| 2020 | OS2D: One-Stage One-Shot Object Detection by Matching Anchor Features
Anton Osokin, Denis Sumin, Vasily Lomakin |
ECCV (15) | 1 |
| 2018 | SEARNN: Training RNNs with global-local losses
Rémi Leblond, Jean-Baptiste Alayrac, Anton Osokin, Simon Lacoste-Julien |
ICLR (Poster) | 3 |
| 2018 | Quantifying Learning Guarantees for Convex but Inconsistent SurrogatesabstractWe study consistency properties of machine learning methods based on minimizing convex surrogates. We extend the recent framework of Osokin et al. (2017) for the quantitative analysis of consistency properties to the case of inconsistent surrogates. Our key technical contribution consists in a new lower bound on the calibration function for the quadratic surrogate, which is non-trivial (not always zero) for inconsistent cases. The new bound allows to quantify the level of inconsistency of the setting and shows how learning with inconsistent surrogates can have guarantees on sample complexity and optimization difficulty. We apply our theory to two concrete cases: multi-class classification with the tree-structured loss and ranking with the mean average precision loss. The results show the approximation-computation trade-offs caused by inconsistent surrogates and their potential benefits. Kirill Struminsky, Simon Lacoste-Julien, Anton Osokin |
NeurIPS | 3 |
| 2018 | Marginal Weighted Maximum Log-likelihood for Efficient Learning of Perturb-and-Map models
Tatiana Shpakova, Francis R. Bach, Anton Osokin |
UAI | 3 |
| 2017 | GANs for Biological Image SynthesisabstractIn this paper, we propose a novel application of Generative Adversarial Networks (GAN) to the synthesis of cells imaged by fluorescence microscopy. Compared to natural images, cells tend to have a simpler and more geometric global structure that facilitates image generation. However, the correlation between the spatial pattern of different fluorescent proteins reflects important biological functions, and synthesized images have to capture these relationships to be relevant for biological applications. We adapt GANs to the task at hand and propose new models with casual dependencies between image channels that can generate multichannel images, which would be impossible to obtain experimentally. We evaluate our approach using two independent techniques and compare it against sensible baselines. Finally, we demonstrate that by interpolating across the latent space we can mimic the known changes in protein localization that occur through time during the cell cycle, allowing us to predict temporal evolution from static images. Anton Osokin, Anatole Chessel, Rafael Edgardo Carazo-Salas, Federico Vaggi |
ICCV | 1 |
| 2017 | On Structured Prediction Theory with Calibrated Convex Surrogate LossesabstractWe provide novel theoretical insights on structured prediction in the context of efficient convex surrogate loss minimization with consistency guarantees. For any task loss, we construct a convex surrogate that can be optimized via stochastic gradient descent and we prove tight bounds on the so-called "calibration function" relating the excess surrogate risk to the actual risk. In contrast to prior related work, we carefully monitor the effect of the exponential number of classes in the learning guarantees as well as on the optimization complexity. As an interesting consequence, we formalize the intuition that some task losses make learning harder than others, and that the classical 0-1 loss is ill-suited for structured prediction. Anton Osokin, Francis R. Bach, Simon Lacoste-Julien |
NIPS | 1 |
| 2016 | Breaking Sticks and Ambiguities with Adaptive Skip-gramabstractThe recently proposed Skip-gram model is a powerful method for learning high-dimensional word representations that capture rich semantic relationships between words. However, Skip-gram as well as most prior work on learning word representations does not take into account word ambiguity and maintain only single representation per word. Although a number of Skip-gram modifications were proposed to overcome this limitation and learn multi-prototype word representations, they either require a known number of word meanings or learn them using greedy heuristic approaches. In this paper we propose the Adaptive Skip-gram model which is a nonparametric Bayesian extension of Skip-gram capable to automatically learn the required number of representations for all words at desired semantic resolution. We derive efficient online variational learning algorithm for the model and empirically demonstrate its efficiency on word-sense induction task. Sergey Bartunov, Dmitry Kondrashkin, Anton Osokin, Dmitry P. Vetrov |
AISTATS | 3 |
| 2016 | Deep Part-Based Generative Shape Model with Latent Variables
Alexander Kirillov, Mikhail Gavrikov, Ekaterina Lobacheva, Anton Osokin, Dmitry P. Vetrov |
BMVC | 4 |
| 2016 | Minding the Gaps for Block Frank-Wolfe Optimization of Structured SVMsabstractIn this paper, we propose several improvements on the block-coordinate Frank-Wolfe (BCFW) algorithm from Lacoste-Julien et al. (2013) recently used to optimize the structured support vector machine (SSVM) objective in the context of structured prediction, though it has wider applications. The key intuition behind our improvements is that the estimates of block gaps maintained by BCFW reveal the block suboptimality that can be used as an *adaptive* criterion. First, we sample objects at each iteration of BCFW in an adaptive non-uniform way via gap-based sampling. Second, we incorporate pairwise and away-step variants of Frank-Wolfe into the block-coordinate setting. Third, we cache oracle calls with a cache-hit criterion based on the block gaps. Fourth, we provide the first method to compute an approximate regularization path for SSVM. Finally, we provide an exhaustive empirical evaluation of all our methods on four structured prediction datasets. Anton Osokin, Jean-Baptiste Alayrac, Isabella Lukasewitz, Puneet K. Dokania, Simon Lacoste-Julien |
ICML | 1 |
| 2015 | Context-Aware CNNs for Person Head DetectionabstractPerson detection is a key problem for many computer vision tasks. While face detection has reached maturity, detecting people under full variation of camera view-points, human poses, lighting conditions and occlusions is still a difficult challenge. In this work we focus on detecting human heads in natural scenes. Starting from the recent R-CNN object detector, we extend it in two ways. First, we leverage person-scene relations and propose a global CNN model trained to predict positions and scales of heads directly from the full image. Second, we explicitly model pairwise relations among the objects via energy-based model where the potentials are computed with a CNN framework. Our full combined model complements R-CNN with contextual cues derived from the scene. To train and test our model, we introduce a large dataset with 369,846 human heads annotated in 224,740 movie frames. We evaluate our method and demonstrate improvements of person head detection compared to several recent baselines on three datasets. We also show improvements of the detection speed provided by our model. Anton Osokin, Ivan Laptev |
ICCV | 2 |
| 2015 | Tensorizing Neural NetworksabstractDeep neural networks currently demonstrate state-of-the-art performance in several domains.At the same time, models of this class are very demanding in terms of computational resources. In particular, a large amount of memory is required by commonly used fully-connected layers, making it hard to use the models on low-end devices and stopping the further increase of the model size. In this paper we convert the dense weight matrices of the fully-connected layers to the Tensor Train format such that the number of parameters is reduced by a huge factor and at the same time the expressive power of the layer is preserved.In particular, for the Very Deep VGG networks we report the compression factor of the dense weight matrix of a fully-connected layer up to 200000 times leading to the compression factor of the whole network up to 7 times. Alexander Novikov 0001, Dmitry Podoprikhin, Anton Osokin, Dmitry P. Vetrov |
NIPS | 3 |
| 2015 | Submodular Relaxation for Inference in Markov Random FieldsabstractIn this paper we address the problem of finding the most probable state of a discrete Markov random field (MRF), also known as the MRF energy minimization problem. The task is known to be NP-hard in general and its practical importance motivates numerous approximate algorithms. We propose a submodular relaxation approach (SMR) based on a Lagrangian relaxation of the initial problem. Unlike the dual decomposition approach of Komodakis et al. [29] SMR does not decompose the graph structure of the initial problem but constructs a submodular energy that is minimized within the Lagrangian relaxation. Our approach is applicable to both pairwise and high-order MRFs and allows to take into account global potentials of certain types. We study theoretical properties of the proposed approach and evaluate it experimentally. Anton Osokin, Dmitry P. Vetrov |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2014 | Perceptually Inspired Layout-Aware Losses for Image Segmentation
Anton Osokin, Pushmeet Kohli |
ECCV (2) | 1 |
| 2014 | Putting MRFs on a Tensor TrainabstractIn the paper we present a new framework for dealing with probabilistic graphical models. Our approach relies on the recently proposed Tensor Train format (TT-format) of a tensor that while being compact allows for efficient application of linear algebra operations. We present a way to convert the energy of a Markov random field to the TT-format and show how one can exploit the properties of the TT-format to attack the tasks of the partition function estimation and the MAP-inference. We provide theoretical guarantees on the accuracy of the proposed algorithm for estimating the partition function and compare our methods against several state-of-the-art algorithms. Alexander Novikov 0001, Anton Rodomanov, Anton Osokin, Dmitry P. Vetrov |
ICML | 3 |
| 2013 | A Principled Deep Random Field Model for Image SegmentationabstractWe discuss a model for image segmentation that is able to overcome the short-boundary bias observed in standard pairwise random field based approaches. To wit, we show that a random field with multi-layered hidden units can encode boundary preserving higher order potentials such as the ones used in the cooperative cuts model of [11] while still allowing for fast and exact MAP inference. Exact inference allows our model to outperform previous image segmentation methods, and to see the true effect of coupling graph edges. Finally, our model can be easily extended to handle segmentation instances with multiple labels, for which it yields promising results. Pushmeet Kohli, Anton Osokin, Stefanie Jegelka |
CVPR | 2 |
| 2012 | Minimizing Sparse High-Order Energies by Submodular Vertex-CoverabstractInference on high-order graphical models has become increasingly important in recent years. We consider energies with simple 'sparse' high-order potentials. Previous work in this area uses either specialized message-passing or transforms each high-order potential to the pairwise case. We take a fundamentally different approach, transforming the entire original problem into a comparatively small instance of a submodular vertex-cover problem. These vertex-cover instances can then be attacked by standard pairwise methods, where they run much faster (4--15 times) and are often more effective than on the original problem. We evaluate our approach on synthetic data, and we show that our algorithm can be useful in a fast hierarchical clustering and model estimation framework. Andrew Delong, Olga Veksler, Anton Osokin, Yuri Boykov |
NIPS | 3 |
| 2012 | Fast Approximate Energy Minimization with Label Costs
Andrew Delong, Anton Osokin, Hossam Isack, Yuri Boykov |
Int. J. Comput. Vis. | 2 |
| 2011 | Submodular decomposition framework for inference in associative Markov networks with global constraintsabstractIn this paper we address the problem of finding the most probable state of discrete Markov random field (MRF) with associative pairwise terms. Although of practical importance, this problem is known to be NP-hard in general. We propose a new type of MRF decomposition, submodular decomposition (SMD). Unlike existing decomposition approaches SMD decomposes the initial problem into sub-problems corresponding to a specific class label while preserving the graph structure of each subproblem. Such decomposition enables us to take into account several types of global constraints in an efficient manner. We study theoretical properties of the proposed approach and demonstrate its applicability on a number of problems. Anton Osokin, Dmitry P. Vetrov, Vladimir Kolmogorov |
CVPR | 1 |
| 2010 | Fast approximate energy minimization with label costsabstractThe α-expansion algorithm has had a significant impact in computer vision due to its generality, effectiveness, and speed. Thus far it can only minimize energies that involve unary, pairwise, and specialized higher-order terms. Our main contribution is to extend α-expansion so that it can simultaneously optimize “label costs” as well. An energy with label costs can penalize a solution based on the set of labels that appear in it. The simplest special case is to penalize the number of labels in the solution. Our energy is quite general, and we prove optimality bounds for our algorithm. A natural application of label costs is multi-model fitting, and we demonstrate several such applications in vision: homography detection, motion segmentation, and unsupervised image segmentation. Our C++/MATLAB implementation is publicly available. Andrew Delong, Anton Osokin, Hossam Isack, Yuri Boykov |
CVPR | 2 |