Anton Milan

dblp:135/4914 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
4since 2021 · last 2025
0000-0002-9719-3616ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 7 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 1 since 2021Systems, architecture and hardware · 3 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
22 papers
Video understanding and tracking · 33% Segmentation and scene understanding · 16% 3D vision · 11%
Theoretical computer science
2 papers
Mathematical optimization · 100%

Topics — the 30 heaviest of 44, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
multi-object tracking
2.382021
MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking · Int. J. Comput. Vis. 2021
Multiple object tracking: A literature review · Artif. Intell. 2021
Online Multi-Target Tracking Using Recurrent Neural Networks · AAAI 2017
Computer vision › 3D vision › geometric deep learning › set learning
set prediction
1.232022
Learn to Predict Sets Using Feed-Forward Neural Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Joint Learning of Set Cardinality and State Distribution · AAAI 2018
DeepSetNet: Predicting Sets with Deep Neural Networks · ICCV 2017
Computer vision › Segmentation and scene understanding
semantic segmentation
1.132020
RefineNet: Multi-Path Refinement Networks for Dense Prediction · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Semantic Segmentation from Limited Training Data · ICRA 2018
RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation · CVPR 2017
Machine learning › Learning paradigms
multi-label classification
0.922022
Learn to Predict Sets Using Feed-Forward Neural Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Joint Learning of Set Cardinality and State Distribution · AAAI 2018
Computer vision › Video understanding and tracking › multi-object tracking
data association
0.942017
Online Multi-Target Tracking Using Recurrent Neural Networks · AAAI 2017
Multi-Target Tracking by Discrete-Continuous Energy Minimization · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Continuous Energy Minimization for Multitarget Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Machine learning › Time series and sequential data
anomaly detection
0.912025
Kaputt: A Large-Scale Dataset for Visual Defect Detection · ICCV 2025
Computer vision › Segmentation and scene understanding › semantic segmentation
high-resolution semantic segmentation
0.722020
RefineNet: Multi-Path Refinement Networks for Dense Prediction · IEEE Trans. Pattern Anal. Mach. Intell. 2020
RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation · CVPR 2017
Computer vision › Image recognition and object detection
object detection
0.742022
NimbRo picking: Versatile part handling for warehouse automation · ICRA 2017
Learn to Predict Sets Using Feed-Forward Neural Networks · IEEE Trans. Pattern Anal. Mach. Intell. 2022
Learning People Detectors for Tracking in Crowded Scenes · ICCV 2013
Computer vision › Video understanding and tracking › multi-object tracking
multi-person pose tracking
0.622018
PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018
PoseTrack: Joint Multi-person Pose Estimation and Tracking · CVPR 2017
Computer vision › Video understanding and tracking › object tracking
tracking benchmark
0.512021
MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking · Int. J. Comput. Vis. 2021
Computer vision › Video understanding and tracking › multi-object tracking
tracking-by-detection
0.422016
Multi-Target Tracking by Discrete-Continuous Energy Minimization · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Continuous Energy Minimization for Multitarget Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Computer vision › 3D vision
depth estimation
0.412020
RefineNet: Multi-Path Refinement Networks for Dense Prediction · IEEE Trans. Pattern Anal. Mach. Intell. 2020
Computer vision › Video understanding and tracking › object tracking
articulated object tracking
0.312018
PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018
Computer vision › Segmentation and scene understanding › semantic segmentation
few-shot segmentation
0.312018
Semantic Segmentation from Limited Training Data · ICRA 2018
Robotics › Robot manipulation
grasping
0.312018
Cartman: The Low-Cost Cartesian Manipulator that Won the Amazon Robotics Challenge · ICRA 2018
Computer vision › Face, body and person analysis
human pose estimation
0.312018
PoseTrack: A Benchmark for Human Pose Estimation and Tracking · CVPR 2018
Robotics › Robot manipulation › grasping
pick-and-place
0.312018
Cartman: The Low-Cost Cartesian Manipulator that Won the Amazon Robotics Challenge · ICRA 2018
Robotics › Robot manipulation
warehouse automation
0.312018
Cartman: The Low-Cost Cartesian Manipulator that Won the Amazon Robotics Challenge · ICRA 2018
Robotics › Robot manipulation › grasping › grasping in clutter
bin picking
0.312017
NimbRo picking: Versatile part handling for warehouse automation · ICRA 2017
Computer vision › Image recognition and object detection › object detection
deep learning object detection
0.312017
NimbRo picking: Versatile part handling for warehouse automation · ICRA 2017
Computer vision › Face, body and person analysis › human pose estimation
multi-person pose estimation
0.312017
PoseTrack: Joint Multi-person Pose Estimation and Tracking · CVPR 2017
Mathematical optimization › combinatorial optimization › assignment problem
quadratic assignment problem
0.312017
Data-Driven Approximations to NP-Hard Problems · AAAI 2017
Mathematical optimization › combinatorial optimization › vehicle routing
traveling salesman problem
0.312017
Data-Driven Approximations to NP-Hard Problems · AAAI 2017
Computer vision › Face, body and person analysis
person re-identification
0.212016
Joint Probabilistic Matching Using m-Best Solutions · CVPR 2016
Computer vision › 3D vision › feature matching
point correspondence
0.212016
Joint Probabilistic Matching Using m-Best Solutions · CVPR 2016
Robotics › Robot navigation and mapping › state estimation
trajectory estimation
0.212016
Multi-Target Tracking by Discrete-Continuous Energy Minimization · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Segmentation and scene understanding
video segmentation
0.212015
Joint tracking and segmentation of multiple targets · CVPR 2015
Machine learning › Optimization for machine learning
energy minimization
0.212014
Continuous Energy Minimization for Multitarget Tracking · IEEE Trans. Pattern Anal. Mach. Intell. 2014
Computer vision › Video understanding and tracking
object tracking
0.212013
Learning People Detectors for Tracking in Crowded Scenes · ICCV 2013
Computer vision › Image recognition and object detection › object detection › category-specific object detection
person detection
0.212013
Learning People Detectors for Tracking in Crowded Scenes · ICCV 2013

Methods — techniques the papers use, named apart from their topics

integer linear programming · 1.0anomaly detection benchmarks · 0.9recurrent neural network · 0.9residual connections · 0.7chained residual pooling · 0.7m-best solutions · 0.7likelihood-based set distribution · 0.6feedforward neural network · 0.6literature review · 0.5multi-path refinement · 0.4evaluation server · 0.3benchmark dataset · 0.3quadratic cost function · 0.2marginal distribution approximation · 0.2
YearPublicationVenuePosition
2025 Kaputt: A Large-Scale Dataset for Visual Defect Detection
abstract
We present a novel large-scale dataset for defect detection in a logistics setting. Recent work on industrial anomaly detection has primarily focused on manufacturing scenarios with highly controlled poses and a limited number of object categories. Existing benchmarks like MVTec-AD [6] and VisA [33] have reached saturation, with state-of-the-art methods achieving up to 99.9% AUROC scores. In contrast to manufacturing, anomaly detection in retail logistics faces new challenges, particularly in the diversity and variability of object pose and appearance. Leading anomaly detection methods fall short when applied to this new setting. To bridge this gap, we introduce a new benchmark that overcomes the current limitations of existing datasets. With over 230,000 images (and more than 29,000 defective instances), it is 40 times larger than MVTec-AD and contains more than 48,000 distinct objects. To validate the difficulty of the problem, we conduct an extensive evaluation of multiple state-of-the-art anomaly detection methods, demonstrating that they do not surpass 56.96% AUROC on our dataset. Further qualitative analysis confirms that existing methods struggle to leverage normal samples under heavy pose and appearance variation. With our large-scale dataset, we set a new benchmark and encourage future research towards solving this challenging problem in retail logistics anomaly detection. The dataset is available for download under https://www.kaputt-dataset.com.
Sebastian Höfer, Dorian Henning, Artemij Amiranashvili, Douglas Morrison, Mariliza Tzes, Ingmar Posner, Marc Matvienko, Alessandro Rennola, Anton Milan
ICCV9
2022 Learn to Predict Sets Using Feed-Forward Neural Networks
abstract
This paper addresses the task of set prediction using deep feed-forward neural networks. A set is a collection of elements which is invariant under permutation and the size of a set is not fixed in advance. Many real-world problems, such as image tagging and object detection, have outputs that are naturally expressed as sets of entities. This creates a challenge for traditional deep neural networks which naturally deal with structured outputs such as vectors, matrices or tensors. We present a novel approach for learning to predict sets with unknown permutation and cardinality using deep neural networks. In our formulation we define a likelihood for a set distribution represented by a) two discrete distributions defining the set cardinally and permutation variables, and b) a joint distribution over set elements with a fixed cardinality. Depending on the problem under consideration, we define different training models for set prediction using deep neural networks. We demonstrate the validity of our set formulations on relevant vision problems such as: 1) multi-label image classification where we outperform the other competing methods on the PASCAL VOC and MS COCO datasets, 2) object detection, for which our formulation outperforms popular state-of-the-art detectors, and 3) a complex CAPTCHA test, where we observe that, surprisingly, our set-based network acquired the ability of mimicking arithmetics without any rules being coded.
Seyed Hamid Rezatofighi, Tianyu Zhu 0001, Roman Kaskman, Farbod T. Motlagh, Qinfeng Shi, Anton Milan, Daniel Cremers, Laura Leal-Taixé, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Multiple object tracking: A literature review
Wenhan Luo, Junliang Xing, Anton Milan, Xiaoqin Zhang 0002, Wei Liu 0005, Tae-Kyun Kim 0001
Artif. Intell.3
2021 MOTChallenge: A Benchmark for Single-Camera Multiple Target Tracking
abstract
Abstract Standardized benchmarks have been crucial in pushing the performance of computer vision algorithms, especially since the advent of deep learning. Although leaderboards should not be over-claimed, they often provide the most objective measure of performance and are therefore important guides for research. We present MOTChallenge , a benchmark for single-camera Multiple Object Tracking (MOT) launched in late 2014, to collect existing and new data and create a framework for the standardized evaluation of multiple object tracking methods. The benchmark is focused on multiple people tracking, since pedestrians are by far the most studied object in the tracking community, with applications ranging from robot navigation to self-driving cars. This paper collects the first three releases of the benchmark: (i) MOT15 , along with numerous state-of-the-art results that were submitted in the last years, (ii) MOT16 , which contains new challenging videos, and (iii) MOT17 , that extends MOT16 sequences with more precise labels and evaluates tracking performance on three different object detectors. The second and third release not only offers a significant increase in the number of labeled boxes, but also provide labels for multiple object classes beside pedestrians, as well as the level of visibility for every single object of interest. We finally provide a categorization of state-of-the-art trackers and a broad error analysis. This will help newcomers understand the related work and research trends in the MOT community, and hopefully shed some light into potential future research directions.
Patrick Dendorfer, Aljosa Osep, Anton Milan, Konrad Schindler, Daniel Cremers, Ian D. Reid 0001, Stefan Roth 0001, Laura Leal-Taixé
Int. J. Comput. Vis.3
2020 RefineNet: Multi-Path Refinement Networks for Dense Prediction
abstract
Recently, very deep convolutional neural networks (CNNs) have shown outstanding performance in object recognition and have also been the first choice for dense prediction problems such as semantic segmentation and depth estimation. However, repeated subsampling operations like pooling or convolution striding in deep CNNs lead to a significant decrease in the initial image resolution. Here, we present RefineNet, a generic multi-path refinement network that explicitly exploits all the information available along the down-sampling process to enable high-resolution prediction using long-range residual connections. In this way, the deeper layers that capture high-level semantic features can be directly refined using fine-grained features from earlier convolutions. The individual components of RefineNet employ residual connections following the identity mapping mindset, which allows for effective end-to-end training. Further, we introduce chained residual pooling, which captures rich background context in an efficient manner. We carry out comprehensive experiments on semantic segmentation which is a dense classification problem and achieve good performance on seven public datasets. We further apply our method for depth estimation and demonstrate the effectiveness of our method on dense regression problems.
Guosheng Lin, Fayao Liu, Anton Milan, Chunhua Shen, Ian D. Reid 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2018 Joint Learning of Set Cardinality and State Distribution
abstract
We present a novel approach for learning to predict sets using deep learning. In recent years, deep neural networks have shown remarkable results in computer vision, natural language processing and other related problems. Despite their success,traditional architectures suffer from a serious limitation in that they are built to deal with structured input and output data,i.e. vectors or matrices. Many real-world problems, however, are naturally described as sets, rather than vectors. Existing techniques that allow for sequential data, such as recurrent neural networks, typically heavily depend on the input and output order and do not guarantee a valid solution. Here, we derive in a principled way, a mathematical formulation for set prediction where the output is permutation invariant. In particular, our approach jointly learns both the cardinality and the state distribution of the target set. We demonstrate the validity of our method on the task of multi-label image classification and achieve a new state of the art on the PASCAL VOC and MS COCO datasets.
Seyed Hamid Rezatofighi, Anton Milan, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001
AAAI2
2018 PoseTrack: A Benchmark for Human Pose Estimation and Tracking
abstract
Existing systems for video-based pose estimation and tracking struggle to perform well on realistic videos with multiple people and often fail to output body-pose trajectories consistent over time. To address this shortcoming this paper introduces PoseTrack which is a new large-scale benchmark for video-based human pose estimation and articulated tracking. Our new benchmark encompasses three tasks focusing on i) single-frame multi-person pose estimation, ii) multi-person pose estimation in videos, and iii) multi-person articulated tracking. To establish the benchmark, we collect, annotate and release a new dataset that features videos with multiple people labeled with person tracks and articulated pose. A public centralized evaluation server is provided to allow the research community to evaluate on a held-out test set. Furthermore, we conduct an extensive experimental study on recent approaches to articulated pose tracking and provide analysis of the strengths and weaknesses of the state of the art. We envision that the proposed benchmark will stimulate productive research both by providing a large and representative training dataset as well as providing a platform to objectively evaluate and compare the proposed methods. The benchmark is freely accessible at https://posetrack.net/.
Mykhaylo Andriluka, Umar Iqbal 0001, Eldar Insafutdinov, Leonid Pishchulin, Anton Milan, Juergen Gall, Bernt Schiele
CVPR5
2018 Semantic Segmentation from Limited Training Data
abstract
We present our approach for robotic perception in cluttered scenes that led to winning the recent Amazon Robotics Challenge (ARC) 2017. Next to small objects with shiny and transparent surfaces, the biggest challenge of the 2017 competition was the introduction of unseen categories. In contrast to traditional approaches which require large collections of annotated data and many hours of training, the task here was to obtain a robust perception pipeline with only few minutes of data acquisition and training time. To that end, we present two strategies that we explored. One is a deep metric learning approach that works in three separate steps: semantic-agnostic boundary detection, patch classification and pixel-wise voting. The other is a fully-supervised semantic segmentation approach with efficient dataset collection. We conduct an extensive analysis of the two methods on our ARC 2017 dataset. Interestingly, only few examples of each class are sufficient to fine-tune even very deep convolutional neural networks for this specific task.
Anton Milan, Trung Pham, Kumar Vijay, Douglas Morrison, Adam W. Tow, Lingqiao Liu, Jordan Erskine, Riccardo Grinover, Alec Gurman, Thomas Hunn, Norton Kelly-Boxall, Darryl Qijun Lee, Matthew McTaggart, Gerald Rallos, Andrew Razjigaev, Thomas James Rowntree, Rohan Smith, Sean Wade-McCue, Zheyu Zhuang, Chris Lehnert, Guosheng Lin, Ian D. Reid 0001, Peter I. Corke, Jürgen Leitner
ICRA1
2018 Cartman: The Low-Cost Cartesian Manipulator that Won the Amazon Robotics Challenge
abstract
The Amazon Robotics Challenge enlisted sixteen teams to each design a pick-and-place robot for autonomous warehousing, addressing development in robotic vision and manipulation. This paper presents the design of our custom-built, cost-effective, Cartesian robot system Cartman, which won first place in the competition finals by stowing 14 (out of 16) and picking all 9 items in 27 minutes, scoring a total of 272 points. We highlight our experience-centred design methodology and key aspects of our system that contributed to our competitiveness. We believe these aspects are crucial to building robust and effective robotic systems.
Douglas Morrison, Adam W. Tow, M. McTaggart, Norton Kelly-Boxall, Sean Wade-McCue, Jordan Erskine, R. Grinover, A. Gurman, T. Hunn, Anton Milan, Trung Pham, G. Rallos, A. Razjigaev, T. Rowntree, K. Vijay, Zheyu Zhuang, Chris Lehnert, Ian D. Reid 0001, Peter I. Corke, Jürgen Leitner
ICRA12
2017 Online Multi-Target Tracking Using Recurrent Neural Networks
abstract
We present a novel approach to online multi-target tracking based on recurrent neural networks (RNNs). Tracking multiple objects in real-world scenes involves many challenges, including a) an a-priori unknown and time-varying number of targets, b) a continuous state estimation of all present targets, and c) a discrete combinatorial problem of data association. Most previous methods involve complex models that require tedious tuning of parameters. Here, we propose for the first time, an end-to-end learning approach for online multi-target tracking. Existing deep learning methods are not designed for the above challenges and cannot be trivially applied to the task. Our solution addresses all of the above points in a principled way. Experiments on both synthetic and real data show promising results obtained at ~300 Hz on a standard CPU, and pave the way towards future research in this direction.
Anton Milan, Seyed Hamid Rezatofighi, Anthony R. Dick, Ian D. Reid 0001, Konrad Schindler
AAAI1
2017 Data-Driven Approximations to NP-Hard Problems
abstract
There exist a number of problem classes for which obtaining the exact solution becomes exponentially expensive with increasing problem size. The quadratic assignment problem (QAP) or the travelling salesman problem (TSP) are just two examples of such NP-hard problems. In practice, approximate algorithms are employed to obtain a suboptimal solution, where one must face a trade-off between computational complexity and solution quality. In this paper, we propose to learn to solve these problem from approximate examples, using recurrent neural networks (RNNs). Surprisingly, such architectures are capable of producing highly accurate solutions at minimal computational cost. Moreover, we introduce a simple, yet effective technique for improving the initial (weak) training set by incorporating the objective cost into the training procedure. We demonstrate the functionality of our approach on three exemplar applications: marginal distributions of a joint matching space, feature point matching and the travelling salesman problem. We show encouraging results on synthetic and real data in all three cases.
Anton Milan, Seyed Hamid Rezatofighi, Ravi Garg, Anthony R. Dick, Ian D. Reid 0001
AAAI1
2017 PoseTrack: Joint Multi-person Pose Estimation and Tracking
abstract
In this work, we introduce the challenging problem of joint multi-person pose estimation and tracking of an unknown number of persons in unconstrained videos. Existing methods for multi-person pose estimation in images cannot be applied directly to this problem, since it also requires to solve the problem of person association over time in addition to the pose estimation for each person. We therefore propose a novel method that jointly models multi-person pose estimation and tracking in a single formulation. To this end, we represent body joint detections in a video by a spatio-temporal graph and solve an integer linear program to partition the graph into sub-graphs that correspond to plausible body pose trajectories for each person. The proposed approach implicitly handles occlusion and truncation of persons. Since the problem has not been addressed quantitatively in the literature, we introduce a challenging Multi-Person PoseTrack dataset, and also propose a completely unconstrained evaluation protocol that does not make any assumptions about the scale, size, location or the number of persons. Finally, we evaluate the proposed approach and several baseline methods on our new dataset.
Umar Iqbal 0001, Anton Milan, Juergen Gall
CVPR2
2017 RefineNet: Multi-path Refinement Networks for High-Resolution Semantic Segmentation
abstract
Recently, very deep convolutional neural networks (CNNs) have shown outstanding performance in object recognition and have also been the first choice for dense classification problems such as semantic segmentation. However, repeated subsampling operations like pooling or convolution striding in deep CNNs lead to a significant decrease in the initial image resolution. Here, we present RefineNet, a generic multi-path refinement network that explicitly exploits all the information available along the down-sampling process to enable high-resolution prediction using long-range residual connections. In this way, the deeper layers that capture high-level semantic features can be directly refined using fine-grained features from earlier convolutions. The individual components of RefineNet employ residual connections following the identity mapping mindset, which allows for effective end-to-end training. Further, we introduce chained residual pooling, which captures rich background context in an efficient manner. We carry out comprehensive experiments and set new state-of-the-art results on seven public datasets. In particular, we achieve an intersection-over-union score of 83.4 on the challenging PASCAL VOC 2012 dataset, which is the best reported result to date.
Guosheng Lin, Anton Milan, Chunhua Shen, Ian D. Reid 0001
CVPR2
2017 DeepSetNet: Predicting Sets with Deep Neural Networks
abstract
This paper addresses the task of set prediction using deep learning. This is important because the output of many computer vision tasks, including image tagging and object detection, are naturally expressed as sets of entities rather than vectors. As opposed to a vector, the size of a set is not fixed in advance, and it is invariant to the ordering of entities within it. We define a likelihood for a set distribution and learn its parameters using a deep neural network. We also derive a loss for predicting a discrete distribution corresponding to set cardinality. Set prediction is demonstrated on the problem of multi-class image classification. Moreover, we show that the proposed cardinality loss can also trivially be applied to the tasks of object counting and pedestrian detection. Our approach outperforms existing methods in all three cases on standard datasets.
Seyed Hamid Rezatofighi, Anton Milan, Ehsan Abbasnejad, Anthony R. Dick, Ian D. Reid 0001
ICCV3
2017 NimbRo picking: Versatile part handling for warehouse automation
abstract
Part handling in warehouse automation is challenging if a large variety of items must be accommodated and items are stored in unordered piles. To foster research in this domain, Amazon holds picking challenges. We present our system which achieved second and third place in the Amazon Picking Challenge 2016 tasks. The challenge required participants to pick a list of items from a shelf or to stow items into the shelf. Using two deep-learning approaches for object detection and semantic segmentation and one item model registration method, our system localizes the requested item. Manipulation occurs using suction on points determined heuristically or from 6D item model registration. Parametrized motion primitives are chained to generate motions. We present a full-system evaluation during the APC 2016 and component-level evaluations of the perception system on an annotated dataset.
Max Schwarz, Anton Milan, Christian Lenz, Aura Munoz, Arul Selvam Periyasamy, Michael Schreiber, Sebastian Schüller, Sven Behnke
ICRA2
2016 Joint Probabilistic Matching Using m-Best Solutions
abstract
Matching between two sets of objects is typically approached by finding the object pairs that collectively maximize the joint matching score. In this paper, we argue that this single solution does not necessarily lead to the optimal matching accuracy and that general one-to-one assignment problems can be improved by considering multiple hypotheses before computing the final similarity measure. To that end, we propose to utilize the marginal distributionsfor each entity. Previously, this idea has been neglected mainly because exact marginalization is intractable due to a combinatorial number of all possible matching permutations. Here, we propose a generic approach to efficiently approximate the marginal distributions by exploiting the m-best solutions of the original problem. This approach not only improves the matching solution, but also provides more accurate ranking of the results, because of the extra information included in the marginal distribution. We validate our claim on two distinct objectives: (i) person re-identification and temporal matching modeled as an integer linear program, and (ii) feature point matching using a quadratic cost function. Our experiments confirm that marginalization indeed leads to superior performance compared to the single (nearly) optimal solution, yielding state-of-the-art results in both applications on standard benchmarks.
Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001
CVPR2
2016 Multi-Target Tracking by Discrete-Continuous Energy Minimization
abstract
The task of tracking multiple targets is often addressed with the so-called tracking-by-detection paradigm, where the first step is to obtain a set of target hypotheses for each frame independently. Tracking can then be regarded as solving two separate, but tightly coupled problems. The first is to carry out data association, i.e., to determine the origin of each of the available observations. The second problem is to reconstruct the actual trajectories that describe the spatio-temporal motion pattern of each individual target. The former is inherently a discrete problem, while the latter should intuitively be modeled in continuous space. Having to deal with an unknown number of targets, complex dependencies, and physical constraints, both are challenging tasks on their own and thus most previous work focuses on one of these subproblems. Here, we present a multi-target tracking approach that explicitly models both tasks as minimization of a unified discrete-continuous energy function. Trajectory properties are captured through global label costs, a recent concept from multi-model fitting, which we introduce to tracking. Specifically, label costs describe physical properties of individual tracks, e.g., linear and angular dynamics, or entry and exit points. We further introduce pairwise label costs to describe mutual interactions between targets in order to avoid collisions. By choosing appropriate forms for the individual energy components, powerful discrete optimization techniques can be leveraged to address data association, while the shapes of individual trajectories are updated by gradient-based continuous energy minimization. The proposed method achieves state-of-the-art results on diverse benchmark sequences.
Anton Milan, Konrad Schindler, Stefan Roth 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Joint tracking and segmentation of multiple targets
abstract
Tracking-by-detection has proven to be the most successful strategy to address the task of tracking multiple targets in unconstrained scenarios [e.g. 40, 53, 55]. Traditionally, a set of sparse detections, generated in a preprocessing step, serves as input to a high-level tracker whose goal is to correctly associate these “dots” over time. An obvious short-coming of this approach is that most information available in image sequences is simply ignored by thresholding weak detection responses and applying non-maximum suppression. We propose a multi-target tracker that exploits low level image information and associates every (super)-pixel to a specific target or classifies it as background. As a result, we obtain a video segmentation in addition to the classical bounding-box representation in unconstrained, real-world videos. Our method shows encouraging results on many standard benchmark sequences and significantly outperforms state-of-the-art tracking-by-detection approaches in crowded scenes with long-term partial occlusions.
Anton Milan, Laura Leal-Taixé, Konrad Schindler, Ian D. Reid 0001
CVPR1
2015 Joint Probabilistic Data Association Revisited
abstract
In this paper, we revisit the joint probabilistic data association (JPDA) technique and propose a novel solution based on recent developments in finding the m-best solutions to an integer linear program. The key advantage of this approach is that it makes JPDA computationally tractable in applications with high target and/or clutter density, such as spot tracking in fluorescence microscopy sequences and pedestrian tracking in surveillance footage. We also show that our JPDA algorithm embedded in a simple tracking framework is surprisingly competitive with state-of-the-art global tracking methods in these two applications, while needing considerably less processing time.
Seyed Hamid Rezatofighi, Anton Milan, Zhen Zhang 0008, Qinfeng Shi, Anthony R. Dick, Ian D. Reid 0001
ICCV2
2014 Continuous Energy Minimization for Multitarget Tracking
abstract
Many recent advances in multiple target tracking aim at finding a (nearly) optimal set of trajectories within a temporal window. To handle the large space of possible trajectory hypotheses, it is typically reduced to a finite set by some form of data-driven or regular discretization. In this work, we propose an alternative formulation of multitarget tracking as minimization of a continuous energy. Contrary to recent approaches, we focus on designing an energy that corresponds to a more complete representation of the problem, rather than one that is amenable to global optimization. Besides the image evidence, the energy function takes into account physical constraints, such as target dynamics, mutual exclusion, and track persistence. In addition, partial image evidence is handled with explicit occlusion reasoning, and different targets are disambiguated with an appearance model. To nevertheless find strong local minima of the proposed nonconvex energy, we construct a suitable optimization scheme that alternates between continuous conjugate gradient descent and discrete transdimensional jump moves. These moves, which are executed such that they always reduce the energy, allow the search to escape weak minima and explore a much larger portion of the search space of varying dimensionality. We demonstrate the validity of our approach with an extensive quantitative evaluation on several public data sets.
Anton Milan, Stefan Roth 0001, Konrad Schindler
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Detection- and Trajectory-Level Exclusion in Multiple Object Tracking
abstract
When tracking multiple targets in crowded scenarios, modeling mutual exclusion between distinct targets becomes important at two levels: (1) in data association, each target observation should support at most one trajectory and each trajectory should be assigned at most one observation per frame, (2) in trajectory estimation, two trajectories should remain spatially separated at all times to avoid collisions. Yet, existing trackers often sidestep these important constraints. We address this using a mixed discrete-continuous conditional random field (CRF) that explicitly models both types of constraints: Exclusion between conflicting observations with super modular pairwise terms, and exclusion between trajectories by generalizing global label costs to suppress the co-occurrence of incompatible labels (trajectories). We develop an expansion move-based MAP estimation scheme that handles both non-sub modular constraints and pairwise global label costs. Furthermore, we perform a statistical analysis of ground-truth trajectories to derive appropriate CRF potentials for modeling data fidelity, target dynamics, and inter-target occlusion.
Anton Milan, Konrad Schindler, Stefan Roth 0001
CVPR1
2013 Learning People Detectors for Tracking in Crowded Scenes
abstract
People tracking in crowded real-world scenes is challenging due to frequent and long-term occlusions. Recent tracking methods obtain the image evidence from object (people) detectors, but typically use off-the-shelf detectors and treat them as black box components. In this paper we argue that for best performance one should explicitly train people detectors on failure cases of the overall tracker instead. To that end, we first propose a novel joint people detector that combines a state-of-the-art single person detector with a detector for pairs of people, which explicitly exploits common patterns of person-person occlusions across multiple viewpoints that are a frequent failure case for tracking in crowded scenes. To explicitly address remaining failure modes of the tracker we explore two methods. First, we analyze typical failures of trackers and train a detector explicitly on these cases. And second, we train the detector with the people tracker in the loop, focusing on the most common tracker failures. We show that our joint multi-person detector significantly improves both detection accuracy as well as tracker performance, improving the state-of-the-art on standard benchmarks.
Siyu Tang 0001, Mykhaylo Andriluka, Anton Milan, Konrad Schindler, Stefan Roth 0001, Bernt Schiele
ICCV3