Anton Osokin

dblp:08/8344 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
1since 2021 · last 2021
0000-0002-8807-5132ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 7 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
15 papers
Probabilistic and Bayesian machine learning · 28% Learning theory · 18% Image recognition and object detection · 15%
Databases, data mining, and information retrieval
2 papers
Data models and query languages · 46% Knowledge graphs · 46% Data mining · 8%
Theoretical computer science
3 papers
Mathematical optimization · 80% Graph algorithms and graph theory · 20%

Topics — the 30 heaviest of 46, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.632015
Submodular Relaxation for Inference in Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Putting MRFs on a Tensor Train · ICML 2014
A Principled Deep Random Field Model for Image Segmentation · CVPR 2013
Machine learning › Probabilistic and Bayesian machine learning
structured prediction
0.522017
On Structured Prediction Theory with Calibrated Convex Surrogate Losses · NIPS 2017
Minding the Gaps for Block Frank-Wolfe Optimization of Structured SVMs · ICML 2016
Data models and query languages › natural language interface
natural language query translation
0.512021
SPARQLing Database Queries from Intermediate Question Decompositions · EMNLP (1) 2021
Knowledge graphs › knowledge graph querying
SPARQL query generation
0.512021
SPARQLing Database Queries from Intermediate Question Decompositions · EMNLP (1) 2021
Computer vision › Image recognition and object detection › object detection
anchor-based detection
0.412020
OS2D: One-Stage One-Shot Object Detection by Matching Anchor Features · ECCV (15) 2020
Computer vision › Image recognition and object detection › object detection › few-shot object detection
one-shot object detection
0.412020
OS2D: One-Stage One-Shot Object Detection by Matching Anchor Features · ECCV (15) 2020
Machine learning › Trustworthy machine learning › uncertainty estimation › probability calibration
calibration analysis
0.312018
Quantifying Learning Guarantees for Convex but Inconsistent Surrogates · NeurIPS 2018
Machine learning › Learning theory › loss function › surrogate loss
consistency of surrogate losses
0.312018
Quantifying Learning Guarantees for Convex but Inconsistent Surrogates · NeurIPS 2018
Machine learning › Learning theory › loss function › surrogate loss
convex surrogate loss
0.312018
Quantifying Learning Guarantees for Convex but Inconsistent Surrogates · NeurIPS 2018
Machine learning › Learning theory
loss function
0.312018
SEARNN: Training RNNs with global-local losses · ICLR (Poster) 2018
Machine learning › Deep learning architectures and training
recurrent neural network
0.312018
SEARNN: Training RNNs with global-local losses · ICLR (Poster) 2018
Machine learning › Deep learning architectures and training › recurrent neural network
recurrent neural network training
0.312018
SEARNN: Training RNNs with global-local losses · ICLR (Poster) 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
MAP inference
0.322014
Putting MRFs on a Tensor Train · ICML 2014
Submodular decomposition framework for inference in associative Markov networks with global constraints · CVPR 2011
Machine learning › Learning theory › loss function
calibrated surrogate
0.312017
On Structured Prediction Theory with Calibrated Convex Surrogate Losses · NIPS 2017
Machine learning › Generative modeling
generative adversarial network
0.312017
GANs for Biological Image Synthesis · ICCV 2017
Machine learning › Learning theory
statistical learning theory
0.312017
On Structured Prediction Theory with Calibrated Convex Surrogate Losses · NIPS 2017
Mathematical optimization
discrete optimization
0.322012
Minimizing Sparse High-Order Energies by Submodular Vertex-Cover · NIPS 2012
Submodular decomposition framework for inference in associative Markov networks with global constraints · CVPR 2011
Machine learning › Optimization for machine learning › constrained optimization
frank-wolfe algorithm
0.212016
Minding the Gaps for Block Frank-Wolfe Optimization of Structured SVMs · ICML 2016
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
structured SVM
0.212016
Minding the Gaps for Block Frank-Wolfe Optimization of Structured SVMs · ICML 2016
Computer vision › Image recognition and object detection › object detection
contextual reasoning
0.212015
Context-Aware CNNs for Person Head Detection · ICCV 2015
Computer vision › Image recognition and object detection › object detection › object part detection
head detection
0.212015
Context-Aware CNNs for Person Head Detection · ICCV 2015
Machine learning › Optimization for machine learning › energy minimization
markov random field optimization
0.212015
Submodular Relaxation for Inference in Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Machine learning › Efficient and distributed learning
model compression
0.212015
Tensorizing Neural Networks · NIPS 2015
Machine learning › Efficient and distributed learning › model compression › low-rank approximation
tensor-train decomposition
0.212015
Tensorizing Neural Networks · NIPS 2015
Mathematical optimization
lagrangian relaxation
0.212015
Submodular Relaxation for Inference in Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Mathematical optimization › submodular optimization
submodular relaxation
0.212015
Submodular Relaxation for Inference in Markov Random Fields · IEEE Trans. Pattern Anal. Mach. Intell. 2015
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.212014
Putting MRFs on a Tensor Train · ICML 2014
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
graphical model inference
0.212014
Putting MRFs on a Tensor Train · ICML 2014
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
partition function estimation
0.212014
Putting MRFs on a Tensor Train · ICML 2014
Machine learning › Deep learning architectures and training › loss function design
perceptual loss
0.212014
Perceptually Inspired Layout-Aware Losses for Image Segmentation · ECCV (2) 2014

Methods — techniques the papers use, named apart from their topics

latent space interpolation · 0.9generative adversarial network · 0.9neural semantic parsing · 0.5intermediate question decomposition · 0.5one-stage detection · 0.4feature matching · 0.4tensor-train decomposition · 0.4global-local losses · 0.3convex surrogate analysis · 0.3calibration function lower bound · 0.3submodular optimization · 0.3message passing · 0.3stochastic gradient descent · 0.3convex surrogate loss · 0.3submodular energy minimization · 0.2lagrangian relaxation · 0.2dual decomposition · 0.2submodular decomposition · 0.1
YearPublicationVenuePosition
2021 SPARQLing Database Queries from Intermediate Question Decompositions
abstract
To translate natural language questions into executable database queries, most approaches rely on a fully annotated training set.Annotating a large dataset with queries is difficult as it requires query-language expertise.We reduce this burden using grounded in databases intermediate question representations.These representations are simpler to collect and were originally crowdsourced within the Break dataset (Wolfson et al., 2020).Our pipeline consists of two parts: a neural semantic parser that converts natural language questions into the intermediate representations and a non-trainable transpiler to the SPARQL query language (a standard language for accessing knowledge graphs and semantic web).We chose SPARQL because its queries are structurally closer to our intermediate representations (compared to SQL).We observe that the execution accuracy of queries constructed by our model on the challenging Spider dataset is comparable with the state-of-the-art text-to-SQL methods trained with annotated SQL queries.Our code and data are publicly available.1
Irina Saparina, Anton Osokin
EMNLP (1)2
2020 OS2D: One-Stage One-Shot Object Detection by Matching Anchor Features
Anton Osokin, Denis Sumin, Vasily Lomakin
ECCV (15)1
2018 SEARNN: Training RNNs with global-local losses
Rémi Leblond, Jean-Baptiste Alayrac, Anton Osokin, Simon Lacoste-Julien
ICLR (Poster)3
2018 Quantifying Learning Guarantees for Convex but Inconsistent Surrogates
abstract
We study consistency properties of machine learning methods based on minimizing convex surrogates. We extend the recent framework of Osokin et al. (2017) for the quantitative analysis of consistency properties to the case of inconsistent surrogates. Our key technical contribution consists in a new lower bound on the calibration function for the quadratic surrogate, which is non-trivial (not always zero) for inconsistent cases. The new bound allows to quantify the level of inconsistency of the setting and shows how learning with inconsistent surrogates can have guarantees on sample complexity and optimization difficulty. We apply our theory to two concrete cases: multi-class classification with the tree-structured loss and ranking with the mean average precision loss. The results show the approximation-computation trade-offs caused by inconsistent surrogates and their potential benefits.
Kirill Struminsky, Simon Lacoste-Julien, Anton Osokin
NeurIPS3
2018 Marginal Weighted Maximum Log-likelihood for Efficient Learning of Perturb-and-Map models
Tatiana Shpakova, Francis R. Bach, Anton Osokin
UAI3
2017 GANs for Biological Image Synthesis
abstract
In this paper, we propose a novel application of Generative Adversarial Networks (GAN) to the synthesis of cells imaged by fluorescence microscopy. Compared to natural images, cells tend to have a simpler and more geometric global structure that facilitates image generation. However, the correlation between the spatial pattern of different fluorescent proteins reflects important biological functions, and synthesized images have to capture these relationships to be relevant for biological applications. We adapt GANs to the task at hand and propose new models with casual dependencies between image channels that can generate multichannel images, which would be impossible to obtain experimentally. We evaluate our approach using two independent techniques and compare it against sensible baselines. Finally, we demonstrate that by interpolating across the latent space we can mimic the known changes in protein localization that occur through time during the cell cycle, allowing us to predict temporal evolution from static images.
Anton Osokin, Anatole Chessel, Rafael Edgardo Carazo-Salas, Federico Vaggi
ICCV1
2017 On Structured Prediction Theory with Calibrated Convex Surrogate Losses
abstract
We provide novel theoretical insights on structured prediction in the context of efficient convex surrogate loss minimization with consistency guarantees. For any task loss, we construct a convex surrogate that can be optimized via stochastic gradient descent and we prove tight bounds on the so-called "calibration function" relating the excess surrogate risk to the actual risk. In contrast to prior related work, we carefully monitor the effect of the exponential number of classes in the learning guarantees as well as on the optimization complexity. As an interesting consequence, we formalize the intuition that some task losses make learning harder than others, and that the classical 0-1 loss is ill-suited for structured prediction.
Anton Osokin, Francis R. Bach, Simon Lacoste-Julien
NIPS1
2016 Breaking Sticks and Ambiguities with Adaptive Skip-gram
abstract
The recently proposed Skip-gram model is a powerful method for learning high-dimensional word representations that capture rich semantic relationships between words. However, Skip-gram as well as most prior work on learning word representations does not take into account word ambiguity and maintain only single representation per word. Although a number of Skip-gram modifications were proposed to overcome this limitation and learn multi-prototype word representations, they either require a known number of word meanings or learn them using greedy heuristic approaches. In this paper we propose the Adaptive Skip-gram model which is a nonparametric Bayesian extension of Skip-gram capable to automatically learn the required number of representations for all words at desired semantic resolution. We derive efficient online variational learning algorithm for the model and empirically demonstrate its efficiency on word-sense induction task.
Sergey Bartunov, Dmitry Kondrashkin, Anton Osokin, Dmitry P. Vetrov
AISTATS3
2016 Deep Part-Based Generative Shape Model with Latent Variables
Alexander Kirillov, Mikhail Gavrikov, Ekaterina Lobacheva, Anton Osokin, Dmitry P. Vetrov
BMVC4
2016 Minding the Gaps for Block Frank-Wolfe Optimization of Structured SVMs
abstract
In this paper, we propose several improvements on the block-coordinate Frank-Wolfe (BCFW) algorithm from Lacoste-Julien et al. (2013) recently used to optimize the structured support vector machine (SSVM) objective in the context of structured prediction, though it has wider applications. The key intuition behind our improvements is that the estimates of block gaps maintained by BCFW reveal the block suboptimality that can be used as an *adaptive* criterion. First, we sample objects at each iteration of BCFW in an adaptive non-uniform way via gap-based sampling. Second, we incorporate pairwise and away-step variants of Frank-Wolfe into the block-coordinate setting. Third, we cache oracle calls with a cache-hit criterion based on the block gaps. Fourth, we provide the first method to compute an approximate regularization path for SSVM. Finally, we provide an exhaustive empirical evaluation of all our methods on four structured prediction datasets.
Anton Osokin, Jean-Baptiste Alayrac, Isabella Lukasewitz, Puneet K. Dokania, Simon Lacoste-Julien
ICML1
2015 Context-Aware CNNs for Person Head Detection
abstract
Person detection is a key problem for many computer vision tasks. While face detection has reached maturity, detecting people under full variation of camera view-points, human poses, lighting conditions and occlusions is still a difficult challenge. In this work we focus on detecting human heads in natural scenes. Starting from the recent R-CNN object detector, we extend it in two ways. First, we leverage person-scene relations and propose a global CNN model trained to predict positions and scales of heads directly from the full image. Second, we explicitly model pairwise relations among the objects via energy-based model where the potentials are computed with a CNN framework. Our full combined model complements R-CNN with contextual cues derived from the scene. To train and test our model, we introduce a large dataset with 369,846 human heads annotated in 224,740 movie frames. We evaluate our method and demonstrate improvements of person head detection compared to several recent baselines on three datasets. We also show improvements of the detection speed provided by our model.
Anton Osokin, Ivan Laptev
ICCV2
2015 Tensorizing Neural Networks
abstract
Deep neural networks currently demonstrate state-of-the-art performance in several domains.At the same time, models of this class are very demanding in terms of computational resources. In particular, a large amount of memory is required by commonly used fully-connected layers, making it hard to use the models on low-end devices and stopping the further increase of the model size. In this paper we convert the dense weight matrices of the fully-connected layers to the Tensor Train format such that the number of parameters is reduced by a huge factor and at the same time the expressive power of the layer is preserved.In particular, for the Very Deep VGG networks we report the compression factor of the dense weight matrix of a fully-connected layer up to 200000 times leading to the compression factor of the whole network up to 7 times.
Alexander Novikov 0001, Dmitry Podoprikhin, Anton Osokin, Dmitry P. Vetrov
NIPS3
2015 Submodular Relaxation for Inference in Markov Random Fields
abstract
In this paper we address the problem of finding the most probable state of a discrete Markov random field (MRF), also known as the MRF energy minimization problem. The task is known to be NP-hard in general and its practical importance motivates numerous approximate algorithms. We propose a submodular relaxation approach (SMR) based on a Lagrangian relaxation of the initial problem. Unlike the dual decomposition approach of Komodakis et al. [29] SMR does not decompose the graph structure of the initial problem but constructs a submodular energy that is minimized within the Lagrangian relaxation. Our approach is applicable to both pairwise and high-order MRFs and allows to take into account global potentials of certain types. We study theoretical properties of the proposed approach and evaluate it experimentally.
Anton Osokin, Dmitry P. Vetrov
IEEE Trans. Pattern Anal. Mach. Intell.1
2014 Perceptually Inspired Layout-Aware Losses for Image Segmentation
Anton Osokin, Pushmeet Kohli
ECCV (2)1
2014 Putting MRFs on a Tensor Train
abstract
In the paper we present a new framework for dealing with probabilistic graphical models. Our approach relies on the recently proposed Tensor Train format (TT-format) of a tensor that while being compact allows for efficient application of linear algebra operations. We present a way to convert the energy of a Markov random field to the TT-format and show how one can exploit the properties of the TT-format to attack the tasks of the partition function estimation and the MAP-inference. We provide theoretical guarantees on the accuracy of the proposed algorithm for estimating the partition function and compare our methods against several state-of-the-art algorithms.
Alexander Novikov 0001, Anton Rodomanov, Anton Osokin, Dmitry P. Vetrov
ICML3
2013 A Principled Deep Random Field Model for Image Segmentation
abstract
We discuss a model for image segmentation that is able to overcome the short-boundary bias observed in standard pairwise random field based approaches. To wit, we show that a random field with multi-layered hidden units can encode boundary preserving higher order potentials such as the ones used in the cooperative cuts model of [11] while still allowing for fast and exact MAP inference. Exact inference allows our model to outperform previous image segmentation methods, and to see the true effect of coupling graph edges. Finally, our model can be easily extended to handle segmentation instances with multiple labels, for which it yields promising results.
Pushmeet Kohli, Anton Osokin, Stefanie Jegelka
CVPR2
2012 Minimizing Sparse High-Order Energies by Submodular Vertex-Cover
abstract
Inference on high-order graphical models has become increasingly important in recent years. We consider energies with simple 'sparse' high-order potentials. Previous work in this area uses either specialized message-passing or transforms each high-order potential to the pairwise case. We take a fundamentally different approach, transforming the entire original problem into a comparatively small instance of a submodular vertex-cover problem. These vertex-cover instances can then be attacked by standard pairwise methods, where they run much faster (4--15 times) and are often more effective than on the original problem. We evaluate our approach on synthetic data, and we show that our algorithm can be useful in a fast hierarchical clustering and model estimation framework.
Andrew Delong, Olga Veksler, Anton Osokin, Yuri Boykov
NIPS3
2012 Fast Approximate Energy Minimization with Label Costs
Andrew Delong, Anton Osokin, Hossam Isack, Yuri Boykov
Int. J. Comput. Vis.2
2011 Submodular decomposition framework for inference in associative Markov networks with global constraints
abstract
In this paper we address the problem of finding the most probable state of discrete Markov random field (MRF) with associative pairwise terms. Although of practical importance, this problem is known to be NP-hard in general. We propose a new type of MRF decomposition, submodular decomposition (SMD). Unlike existing decomposition approaches SMD decomposes the initial problem into sub-problems corresponding to a specific class label while preserving the graph structure of each subproblem. Such decomposition enables us to take into account several types of global constraints in an efficient manner. We study theoretical properties of the proposed approach and demonstrate its applicability on a number of problems.
Anton Osokin, Dmitry P. Vetrov, Vladimir Kolmogorov
CVPR1
2010 Fast approximate energy minimization with label costs
abstract
The α-expansion algorithm has had a significant impact in computer vision due to its generality, effectiveness, and speed. Thus far it can only minimize energies that involve unary, pairwise, and specialized higher-order terms. Our main contribution is to extend α-expansion so that it can simultaneously optimize “label costs” as well. An energy with label costs can penalize a solution based on the set of labels that appear in it. The simplest special case is to penalize the number of labels in the solution. Our energy is quite general, and we prove optimality bounds for our algorithm. A natural application of label costs is multi-model fitting, and we demonstrate several such applications in vision: homography detection, motion segmentation, and unsupervised image segmentation. Our C++/MATLAB implementation is publicly available.
Andrew Delong, Anton Osokin, Hossam Isack, Yuri Boykov
CVPR2