VLDB 2026 Research / reviewers in the wild / expert
Marius Leordeanu
dblp:21/5985
· DBLP profile ↗
43ranked-venue papers
14as first author
9since 2021 · last 2025
0000-0001-8479-8758ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 40 · 14 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 28 · 10 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Learning a fast 3D spectral approach to object segmentation and tracking over space and time
Elena Burceanu, Marius Leordeanu |
Artif. Intell. | 2 |
| 2025 | TeachText: CrossModal text-video retrieval through generalized distillation
Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Yang Liu 0105, Samuel Albanie |
Artif. Intell. | 3 |
| 2024 | A self-supervised cyclic neural-analytic approach for novel view synthesis and 3D reconstruction
Dragos Costea, Alina Marcu, Marius Leordeanu |
BMVC | 3 |
| 2022 | UFO Depth: Unsupervised learning with flow-based odometry optimization for metric depth estimationabstractWe propose an efficient method for unsupervised learning of metric depth estimation from a single image in the context of unconstrained videos captured from UAVs. We combine the accuracy of an analytical solution based on odometry with the power of deep learning. First, we show how to correct the noisy odometric measurements by optimizing the alignment between the derotated optical flow and the projected linear speed in the image. Then, we detail an analytical depth estimation method based on optical flow and corrected camera velocities. Subsequently, the improved depth and camera veloc-ities obtained analytically are used, as additional cost terms, for training our novel unsupervised learning architecture for metric depth estimation. We extensively test on a recent UAV dataset, which we significantly extend by adding completely novel scenes. We outperform by significant margins different kinds of state-of-the-art approaches, ranging from analytical and unsupervised solutions to transformer-based architectures that require heavy computation and pre-training. The resulting algorithm could be deployed on embedded devices, being a good candidate for practical robotics use cases, such as obstacle avoidance and safe landing for UAV s. Vlad Licaret, Victor Robu, Alina Marcu, Dragos Costea, Emil Slusanschi, Rahul Sukthankar, Marius Leordeanu |
ICRA | 7 |
| 2022 | Iterative Knowledge Exchange Between Deep Learning and Space-Time Spectral Clustering for Unsupervised Segmentation in VideosabstractWe propose a dual system for unsupervised object segmentation in video, which brings together two modules with complementary properties: a space-time graph that discovers objects in videos and a deep network that learns powerful object features. The system uses an iterative knowledge exchange policy. A novel spectral space-time clustering process on the graph produces unsupervised segmentation masks passed to the network as pseudo-labels. The net learns to segment in single frames what the graph discovers in video and passes back to the graph strong image-level features that improve its node-level features in the next iteration. Knowledge is exchanged for several cycles until convergence. The graph has one node per each video pixel, but the object discovery is fast. It uses a novel power iteration algorithm computing the main space-time cluster as the principal eigenvector of a special Feature-Motion matrix without actually computing the matrix. The thorough experimental analysis validates our theoretical claims and proves the effectiveness of the cyclical knowledge exchange. We also perform experiments on the supervised scenario, incorporating features pretrained with human supervision. We achieve state-of-the-art level on unsupervised and supervised scenarios on four challenging datasets: DAVIS, SegTrack, YouTube-Objects, and DAVSOD. We will make our code publicly available. Emanuela Haller, Adina Magda Florea, Marius Leordeanu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Semi-Supervised Learning for Multi-Task Scene Understanding by Neural Graph ConsensusabstractWe address the challenging problem of semi-supervised learning in the context of multiple visual interpretations of the world by finding consensus in a graph of neural networks. Each graph node is a scene interpretation layer, while each edge is a deep net that transforms one layer at one node into another from a different node. During the supervised phase edge networks are trained independently. During the next unsupervised stage edge nets are trained on the pseudo-ground truth provided by consensus among multiple paths that reach the nets' start and end nodes. These paths act as ensemble teachers for any given edge and strong consensus is used for high-confidence supervisory signal. The unsupervised learning process is repeated over several generations, in which each edge becomes a "student" and also part of different ensemble "teachers" for training other students. By optimizing such consensus between different paths, the graph reaches consistency and robustness over multiple interpretations and generations, in the face of unknown labels. We give theoretical justifications of the proposed idea and validate it on a large dataset. We show how prediction of different representations such as depth, semantic segmentation, surface normals and pose from RGB input could be effectively learned through self-supervised consensus in our graph. We also compare to state-of-the-art methods for multi-task and semi-supervised learning and show superior performance. Marius Leordeanu, Mihai Cristian Pîrvu, Dragos Costea, Alina Marcu, Emil Slusanschi, Rahul Sukthankar |
AAAI | 1 |
| 2021 | Self-Supervised Learning in Multi-Task Graphs through Iterative Consensus Shift
Emanuela Haller, Elena Burceanu, Marius Leordeanu |
BMVC | 3 |
| 2021 | TeachText: CrossModal Generalized Distillation for Text-Video RetrievalabstractIn recent years, considerable progress on the task of text-video retrieval has been achieved by leveraging large-scale pretraining on visual and audio datasets to construct powerful video encoders. By contrast, despite the natural symmetry, the design of effective algorithms for exploiting large-scale language pretraining remains under-explored. In this work, we are the first to investigate the design of such algorithms and propose a novel generalized distillation method, TeachText, which leverages complementary cues from multiple text encoders to provide an enhanced supervisory signal to the retrieval model. Moreover, we extend our method to video side modalities and show that we can effectively reduce the number of used modalities at test time without compromising performance. Our approach advances the state of the art on several video retrieval benchmarks by a significant margin and adds no computational overhead at test time. Last but not least, we show an effective application of our method for eliminating noise from retrieval datasets. Code and data can be found at https://www.robots.ox.ac.uk/˜vgg/research/teachtext/. Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu, Hailin Jin, Andrew Zisserman, Samuel Albanie, Yang Liu 0105 |
ICCV | 3 |
| 2021 | Discovering Dynamic Salient Regions for Spatio-Temporal Graph Neural NetworksabstractGraph Neural Networks are perfectly suited to capture latent interactions between various entities in the spatio-temporal domain (e.g. videos). However, when an explicit structure is not available, it is not obvious what atomic elements should be represented as nodes. Current works generally use pre-trained object detectors or fixed, predefined regions to extract graph nodes. Improving upon this, our proposed model learns nodes that dynamically attach to well-delimited salient regions, which are relevant for a higher-level task, without using any object-level supervision. Constructing these localized, adaptive nodes gives our model inductive bias towards object-centric representations and we show that it discovers regions that are well correlated with objects in the video. In extensive ablation studies and experiments on two challenging datasets, we show superior performance to previous graph neural networks models for video classification. Iulia Duta, Andrei Liviu Nicolicioiu, Marius Leordeanu |
NeurIPS | 3 |
| 2020 | Semantics Through Time: Semi-supervised Segmentation of Aerial Videos with Iterative Label Propagation
Alina Marcu, Vlad Licaret, Dragos Costea, Marius Leordeanu |
ACCV (1) | 4 |
| 2020 | A hierarchical approach to vision-based language generation: from simple sentences to complex natural languageabstractAutomatically describing videos in natural language is an ambitious problem, which could bridge our understanding of vision and language.We propose a hierarchical approach, by first generating video descriptions as sequences of simple sentences, followed at the next level by a more complex and fluent description in natural language.While the simple sentences describe simple actions in the form of (subject, verb, object), the second-level paragraph descriptions, indirectly using information from the first-level description, presents the visual content in a more compact, coherent and semantically rich manner.To this end, we introduce the first video dataset in the literature that is annotated with captions at two levels of linguistic complexity.We perform extensive tests that demonstrate that our hierarchical linguistic representation, from simple to complex language, allows us to train a two-stage network that is able to generate significantly more complex paragraphs than current one-stage approaches. Simion-Vlad Bogolin, Ioana Croitoru, Marius Leordeanu |
COLING | 3 |
| 2020 | A 3D Convolutional Approach to Spectral Object Segmentation in Space and TimeabstractWe formulate object segmentation in video as a spectral graph clustering problem in space and time, in which nodes are pixels and their relations form local neighbourhoods. We claim that the strongest cluster in this pixel-level graph represents the salient object segmentation. We compute the main cluster using a novel and fast 3D filtering technique that finds the spectral clustering solution, namely the principal eigenvector of the graph's adjacency matrix, without building the matrix explicitly - which would be intractable. Our method is based on the power iteration which we prove is equivalent to performing a specific set of 3D convolutions in the space-time feature volume. This allows us to avoid creating the matrix and have a fast parallel implementation on GPU. We show that our method is much faster than classical power iteration applied directly on the adjacency matrix. Different from other works, ours is dedicated to preserving object consistency in space and time at the level of pixels. In experiments, we obtain consistent improvement over the top state of the art methods on DAVIS-2016 dataset. We also achieve top results on the well-known SegTrackv2 dataset. Elena Burceanu, Marius Leordeanu |
IJCAI | 2 |
| 2020 | Image Difficulty Curriculum for Generative Adversarial Networks (CuGAN)abstractDespite the significant advances in recent years, Generative Adversarial Networks (GANs) are still notoriously hard to train. In this paper, we propose three novel curriculum learning strategies for training GANs. All strategies are first based on ranking the training images by their difficulty scores, which are estimated by a state-of-the-art image difficulty predictor. Our first strategy is to divide images into gradually more difficult batches. Our second strategy introduces a novel curriculum loss function for the discriminator that takes into account the difficulty scores of the real images. Our third strategy is based on sampling from an evolving distribution, which favors the easier images during the initial training stages and gradually converges to a uniform distribution, in which samples are equally likely, regardless of difficulty. We compare our curriculum learning strategies with the classic training procedure on two tasks: image generation and image translation. Our experiments indicate that all strategies provide faster convergence and superior results. For example, our best curriculum learning strategy applied on spectrally normalized GANs (SNGANs) fooled human annotators in thinking that generated CIFAR-like images are real in 25.0% of the presented cases, while the SNGANs trained using the classic procedure fooled the annotators in only 18.4% cases. Similarly, in image translation, the human annotators preferred the images produced by the Cycle-consistent GAN (CycleGAN) trained using curriculum learning in 40.5% cases and those produced by CycleGAN based on classic training in only 19.8% cases, 39.7% cases being labeled as ties. Petru Soviany, Claudiu Ardei, Radu Tudor Ionescu, Marius Leordeanu |
WACV | 4 |
| 2020 | Reading into the mind's eye: Boosting automatic visual recognition with EEG signalsabstractClassifying visual information is an apparently simple and effortless task in our everyday routine, but can we automatically predict what we see from signals emitted by the brain? While other researchers have already attempted to answer this question, we are the first to show that a commercially available BCI could be effectively used for visual image classification in real-world scenarios – when testing takes place at a completely different time than training data collection. The task is difficult, as it requires relating the noisy and low-level EEG signals to complex and highly semantic visual categories. In this paper, we propose different learning approaches and show that simpler classifiers such as Ridge Regression with Gabor filtering of the input EEG signal could be more effective than the powerful Long Short Term Memory Networks and Convolutional Neural Networks in this case of limited and noisy training data. We analyzed the importance of each electrode for the visual classification task and noticed that the sensors with the highest accuracy were the ones that recorded brain activity from regions known to be correlated more with higher level recognition and cognitive processes and less to lower-level visual signal processing. The result is also in accordance with research in computer vision with deep neural networks, which shows that semantic visual features are learned only at higher levels of neural depth. While EEG signals are weaker by themselves for the task of visual classification, we demonstrate that they could be powerful when combined with deep visual features extracted from the image, improving performance from 91% to over 97% in a multi-class recognition setting. Our tests show that EEG input brings additional information that is not learned by artificial deep networks on the given image training set. Thus, a commercially available BCI could be effectively used in conjunction with a deep learning based vision system to form together a stronger visual recognition system that is suitable for real-world applications. Nicolae Cudlenco, Nirvana Popescu, Marius Leordeanu |
Neurocomputing | 3 |
| 2019 | Shift R-CNN: Deep Monocular 3D Object Detection With Closed-Form Geometric ConstraintsabstractWe propose Shift R-CNN, a hybrid model for monocular 3D object detection, which combines deep learning with the power of geometry. We adapt a Faster R-CNN network for regressing initial 2D and 3D object properties and combine it with a least squares solution for the inverse 2D to 3D geometric mapping problem, using the camera projection matrix. The closed-form solution of the mathematical system, along with the initial output of the adapted Faster R-CNN are then passed through a final ShiftNet network that refines the result using our newly proposed Volume Displacement Loss. Our novel, geometrically constrained deep learning approach to monocular 3D object detection obtains top results on KITTI 3D Object Detection Benchmark [5], being the best among all monocular methods that do not use any pre-trained network for depth estimation. Andretti Naiden, Vlad Paunescu, Gyeongmo Kim, ByeongMoon Jeon, Marius Leordeanu |
ICIP | 5 |
| 2019 | Recurrent Space-time Graph Neural NetworksabstractLearning in the space-time domain remains a very challenging problem in machine learning and computer vision. Current computational models for understanding spatio-temporal visual data are heavily rooted in the classical single-image based paradigm. It is not yet well understood how to integrate information in space and time into a single, general model. We propose a neural graph model, recurrent in space and time, suitable for capturing both the local appearance and the complex higher-level interactions of different entities and objects within the changing world scene. Nodes and edges in our graph have dedicated neural networks for processing information. Nodes operate over features extracted from local parts in space and time and over previous memory states. Edges process messages between connected nodes at different locations and spatial scales or between past and present time. Messages are passed iteratively in order to transmit information globally and establish long range interactions. Our model is general and could learn to recognize a variety of high level spatio-temporal concepts and be applied to different learning tasks. We demonstrate, through extensive experiments and ablation studies, that our model outperforms strong baselines and top published methods on recognizing complex activities in video. Moreover, we obtain state-of-the-art performance on the challenging Something-Something human-object interaction dataset. Andrei Liviu Nicolicioiu, Iulia Duta, Marius Leordeanu |
NeurIPS | 3 |
| 2019 | Unsupervised Learning of Foreground Object SegmentationabstractUnsupervised learning represents one of the most interesting challenges in computer vision today. The task has an immense practical value with many applications in artificial intelligence and emerging technologies, as large quantities of unlabeled images and videos can be collected at low cost. In this paper, we address the unsupervised learning problem in the context of segmenting the main foreground objects in single images. We propose an unsupervised learning system, which has two pathways, the teacher and the student, respectively. The system is designed to learn over several generations of teachers and students. At every generation the teacher performs unsupervised object discovery in videos or collections of images and an automatic selection module picks up good frame segmentations and passes them to the student pathway for training. At every generation multiple students are trained, with different deep network architectures to ensure a better diversity. The students at one iteration help in training a better selection module, forming together a more powerful teacher pathway at the next iteration. In experiments, we show that the improvement in the selection power, the training of multiple students and the increase in unlabeled data significantly improve segmentation accuracy from one generation to the next. Our method achieves top results on three current datasets for object discovery in video, unsupervised image segmentation and saliency detection. At test time, the proposed system is fast, being one to two orders of magnitude faster than published unsupervised methods. We also test the strength of our unsupervised features within a well known transfer learning setup and achieve competitive performance, proving that our unsupervised approach can be reliably used in a variety of computer vision tasks. Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu |
Int. J. Comput. Vis. | 3 |
| 2018 | Mining for meaning: from vision to language through multiple networks consensus
Iulia Duta, Andrei Liviu Nicolicioiu, Simion-Vlad Bogolin, Marius Leordeanu |
BMVC | 4 |
| 2018 | Does automatic game difficulty level adjustment improve acrophobia therapy?: differences from baselineabstractThis paper presents the design and development of a Virtual Reality game for treating acrophobia, as well as a comparative study between the players' performance in the game, under two different conditions - one in which the difficulty levels are adjusted according to the subjects' biophysical data and one in which they are not. The results showed an improvement of the parameters correlated with fear level in the first experiment. Oana Mitrut, Gabriela Moise, Alin Moldoveanu, Florica Moldoveanu, Marius Leordeanu |
VRST | 5 |
| 2017 | Unsupervised Learning from Video to Detect Foreground Objects in Single ImagesabstractUnsupervised learning from visual data is one of the most difficult challenges in computer vision. It is essential for understanding how visual recognition works. Learning from unsupervised input has an immense practical value, as huge quantities of unlabeled videos can be collected at low cost. Here we address the task of unsupervised learning to detect and segment foreground objects in single images. We achieve our goal by training a student pathway, consisting of a deep neural network that learns to predict, from a single input image, the output of a teacher pathway that performs unsupervised object discovery in video. Our approach is different from the published methods that perform unsupervised discovery in videos or in collections of images at test time. We move the unsupervised discovery phase during the training stage, while at test time we apply the standard feed-forward processing along the student pathway. This has a dual benefit: firstly, it allows, in principle, unlimited generalization possibilities during training, while remaining fast at testing. Secondly, the student not only becomes able to detect in single images significantly better than its unsupervised video discovery teacher, but it also achieves state of the art results on two current benchmarks, YouTube Objects and Object Discovery datasets. At test time, our system is two orders of magnitude faster than other previous methods. Ioana Croitoru, Simion-Vlad Bogolin, Marius Leordeanu |
ICCV | 3 |
| 2017 | Unsupervised Object Segmentation in Video by Efficient Selection of Highly Probable Positive Features
Emanuela Haller, Marius Leordeanu |
ICCV | 2 |
| 2016 | Labeling the Features Not the Samples: Efficient Video Classification with Minimal SupervisionabstractFeature selection is essential for effective visual recognition. We propose an efficient joint classifier learning and feature selection method that discovers sparse, compact representations of input features from a vast sea of candidates, with an almost unsupervised formulation. Our method requires only the following knowledge, which we call the feature sign - whether or not a particular feature has on average stronger values over positive samples than over negatives. We show how this can be estimated using as few as a single labeled training sample per class. Then, using these feature signs, we extend an initial supervised learning problem into an (almost) unsupervised clustering formulation that can incorporate new data without requiring ground truth labels. Our method works both as a feature selection mechanism and as a fully competitive classifier. It has important properties, low computational cost annd excellent accuracy, especially in difficult cases of very limited training data. We experiment on large-scale recognition in video and show superior speed and performance to established feature selection approaches such as AdaBoost, Lasso, greedy forward-backward selection, and powerful classifiers such as SVM. Marius Leordeanu, Alexandra Radu, Shumeet Baluja, Rahul Sukthankar |
AAAI | 1 |
| 2016 | Aerial image geolocalization from recognition and matching of roads and intersections
Dragos Costea, Marius Leordeanu |
BMVC | 2 |
| 2016 | How Hard Can It Be? Estimating the Difficulty of Visual Search in an ImageabstractWe address the problem of estimating image difficulty defined as the human response time for solving a visual search task. We collect human annotations of image difficulty for the PASCAL VOC 2012 data set through a crowd-sourcing platform. We then analyze what human interpretable image properties can have an impact on visual search difficulty, and how accurate are those properties for predicting difficulty. Next, we build a regression model based on deep features learned with state of the art convolutional neural networks and show better results for predicting the ground-truth visual search difficulty scores produced by human annotators. Our model is able to correctly rank about 75% image pairs according to their difficulty score. We also show that our difficulty predictor generalizes well to new classes not seen during training. Finally, we demonstrate that our predicted difficulty scores are useful for weakly supervised object localization (8% improvement) and semi-supervised object classification (1% improvement). Radu Tudor Ionescu, Alexe Dumitru-Bogdan, Marius Leordeanu, Marius Popescu, Dim P. Papadopoulos, Vittorio Ferrari |
CVPR | 3 |
| 2015 | Multiple Frames Matching for Object Discovery in VideoabstractAutomatic discovery of foreground objects in video sequences is important in computer vision, with applications to object tracking, video segmentation and weakly supervised learning. This task is related to cosegmentation [4, 5] and weakly supervised localization [2, 6]. We propose an efficient method for the simultaneous discovery of foreground objects in video and their segmentation masks across multiple frames. We offer a graph matching formulation for bounding box selection and refinement using second and higher order terms. It is based on an Integer Quadratic Programming formulation and related to graph matching and MAP inference [3]. We take into consideration local frame-based information as well as spatiotemporal and appearance consistency over multiple frames. Our approach consists of three stages. First, we find an initial pool of candidate boxes using a novel and fast foreground estimation method in video (VideoPCA) based on Principal Component Analysis of the video content. The output of VideoPCA combined with Edge Boxes [8] is then used to produce high quality bounding box proposals. Second, we efficiently match bounding boxes across multiple frames, using the IPFP algorithm [3] with pairwise geometric and appearance terms. Third, we optimize the higher order terms using the Mean-Shift algorithm [1] to refine the box locations and establish appearance regularity over multiple frames. We make the following contributions: Otilia Stretcu, Marius Leordeanu |
BMVC | 2 |
| 2014 | Generalized Boundaries from Multiple Image InterpretationsabstractBoundary detection is a fundamental computer vision problem that is essential for a variety of tasks, such as contour and region segmentation, symmetry detection and object recognition and categorization. We propose a generalized formulation for boundary detection, with closed-form solution, applicable to the localization of different types of boundaries, such as object edges in natural images and occlusion boundaries from video. Our generalized boundary detection method (Gb) simultaneously combines low-level and mid-level image representations in a single eigenvalue problem and solves for the optimal continuous boundary orientation and strength. The closed-form solution to boundary detection enables our algorithm to achieve state-of-the-art results at a significantly lower computational cost than current methods. We also propose two complementary novel components that can seamlessly be combined with Gb: first, we introduce a soft-segmentation procedure that provides region input layers to our boundary detection algorithm for a significant improvement in accuracy, at negligible computational cost; second, we present an efficient method for contour grouping and reasoning, which when applied as a final post-processing stage, further increases the boundary detection performance. Marius Leordeanu, Rahul Sukthankar, Cristian Sminchisescu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Locally Affine Sparse-to-Dense Matching for Motion and Occlusion EstimationabstractEstimating a dense correspondence field between successive video frames, under large displacement, is important in many visual learning and recognition tasks. We propose a novel sparse-to-dense matching method for motion field estimation and occlusion detection. As an alternative to the current coarse-to-fine approaches from the optical flow literature, we start from the higher level of sparse matching with rich appearance and geometric constraints collected over extended neighborhoods, using an occlusion aware, locally affine model. Then, we move towards the simpler, but denser classic flow field model, with an interpolation procedure that offers a natural transition between the sparse and the dense correspondence fields. We experimentally demonstrate that our appearance features and our complex geometric constraints permit the correct motion estimation even in difficult cases of large displacements and significant appearance changes. We also propose a novel classification method for occlusion detection that works in conjunction with the sparse-to-dense matching model. We validate our approach on the newly released Sintel dataset and obtain state-of-the-art results. Marius Leordeanu, Andrei Zanfir, Cristian Sminchisescu |
ICCV | 1 |
| 2013 | The Moving Pose: An Efficient 3D Kinematics Descriptor for Low-Latency Action Recognition and DetectionabstractHuman action recognition under low observational latency is receiving a growing interest in computer vision due to rapidly developing technologies in human-robot interaction, computer gaming and surveillance. In this paper we propose a fast, simple, yet powerful non-parametric Moving Pose (MP) framework for low-latency human action and activity recognition. Central to our methodology is a moving pose descriptor that considers both pose information as well as differential quantities (speed and acceleration) of the human body joints within a short time window around the current frame. The proposed descriptor is used in conjunction with a modified kNN classifier that considers both the temporal location of a particular frame within the action sequence as well as the discrimination power of its moving pose descriptor compared to other frames in the training set. The resulting method is non-parametric and enables low-latency recognition, one-shot learning, and action detection in difficult unsegmented sequences. Moreover, the framework is real-time, scalable, and outperforms more sophisticated approaches on challenging benchmarks like MSR-Action3D or MSR-DailyActivities3D. Mihai Zanfir, Marius Leordeanu, Cristian Sminchisescu |
ICCV | 2 |
| 2012 | Efficient Closed-Form Solution to Generalized Boundary Detection
Marius Leordeanu, Rahul Sukthankar, Cristian Sminchisescu |
ECCV (4) | 1 |
| 2012 | Unsupervised Learning for Graph Matching
Marius Leordeanu, Rahul Sukthankar, Martial Hebert |
Int. J. Comput. Vis. | 1 |
| 2011 | Semi-supervised learning and optimization for hypergraph matchingabstractGraph and hypergraph matching are important problems in computer vision. They are successfully used in many applications requiring 2D or 3D feature matching, such as 3D reconstruction and object recognition. While graph matching is limited to using pairwise relationships, hypergraph matching permits the use of relationships between sets of features of any order. Consequently, it carries the promise to make matching more robust to changes in scale, deformations and outliers. In this paper we make two contributions. First, we present a first semi-supervised algorithm for learning the parameters that control the hypergraph matching model and demonstrate experimentally that it significantly improves the performance of current state-of-the-art methods. Second, we propose a novel efficient hypergraph matching algorithm, which outperforms the state-of-the-art, and, when used in combination with other higher-order matching algorithms, it consistently improves their performance. Marius Leordeanu, Andrei Zanfir, Cristian Sminchisescu |
ICCV | 1 |
| 2009 | Unsupervised learning for graph matchingabstractGraph matching is an important problem in computer vision. It is used in 2D and 3D object matching and recognition. Despite its importance, there is little literature on learning the parameters that control the graph matching problem, even though learning is important for improving the matching rate, as shown by this and other work. In this paper we show for the first time how to perform parameter learning in an unsupervised fashion, that is when no correct correspondences between graphs are given during training. We show empirically that unsupervised learning is comparable in efficiency and quality with the supervised one, while avoiding the tedious manual labeling of ground truth correspondences. We also verify experimentally that this learning method can improve the performance of several state-of-the art graph matching algorithms. Marius Leordeanu, Martial Hebert |
CVPR | 1 |
| 2009 | An Integer Projected Fixed Point Method for Graph Matching and MAP InferenceabstractGraph matching and MAP inference are essential problems in computer vision and machine learning. We introduce a novel algorithm that can accommodate both problems and solve them efficiently. Recent graph matching algorithms are based on a general quadratic programming formulation, that takes in consideration both unary and second-order terms reflecting the similarities in local appearance as well as in the pairwise geometric relationships between the matched features. In this case the problem is NP-hard and a lot of effort has been spent in finding efficiently approximate solutions by relaxing the constraints of the original problem. Most algorithms find optimal continuous solutions of the modified problem, ignoring during the optimization the original discrete constraints. The continuous solution is quickly binarized at the end, but very little attention is put into this final discretization step. In this paper we argue that the stage in which a discrete solution is found is crucial for good performance. We propose an efficient algorithm, with climbing and convergence properties, that optimizes in the discrete domain the quadratic score, and it gives excellent results either by itself or by starting from the solution returned by any graph matching algorithm. In practice it outperforms state-or-the art algorithms and it also significantly improves their performance if used in combination. When applied to MAP inference, the algorithm is a parallel extension of Iterated Conditional Modes (ICM) with climbing and convergence properties that make it a compelling alternative to the sequential ICM. In our experiments on MAP inference our algorithm proved its effectiveness by outperforming ICM and Max-Product Belief Propagation. Marius Leordeanu, Martial Hebert, Rahul Sukthankar |
NIPS | 1 |
| 2008 | Smoothing-based OptimizationabstractWe propose an efficient method for complex optimization problems that often arise in computer vision. While our method is general and could be applied to various tasks, it was mainly inspired from problems in computer vision, and it borrows ideas from scale space theory. One of the main motivations for our approach is that searching for the global maximum through the scale space of a function is equivalent to looking for the maximum of the original function, with the advantage of having to handle fewer local optima. Our method works with any non-negative, possibly non-smooth function, and requires only the ability of evaluating the function at any specific point. The algorithm is based on a growth transformation, which is guaranteed to increase the value of the scale space function at every step, unlike gradient methods. To demonstrate its effectiveness we present its performance on a few computer vision applications, and show that in our experiments it is more effective than some well established methods such as MCMC, Simulated Annealing and the more local Nelder-Mead optimization method. Marius Leordeanu, Martial Hebert |
CVPR | 1 |
| 2008 | Discriminative Sparse Image Models for Class-Specific Edge Detection and Image Interpretation
Julien Mairal, Marius Leordeanu, Francis R. Bach, Martial Hebert, Jean Ponce |
ECCV (3) | 2 |
| 2007 | Beyond Local Appearance: Category Recognition from Pairwise Interactions of Simple FeaturesabstractWe present a discriminative shape-based algorithm for object category localization and recognition. Our method learns object models in a weakly-supervised fashion, without requiring the specification of object locations nor pixel masks in the training data. We represent object models as cliques of fully-interconnected parts, exploiting only the pairwise geometric relationships between them. The use of pairwise relationships enables our algorithm to successfully overcome several problems that are common to previously-published methods. Even though our algorithm can easily incorporate local appearance information from richer features, we purposefully do not use them in order to demonstrate that simple geometric relationships can match (or exceed) the performance of state-of-the-art object recognition algorithms. Marius Leordeanu, Martial Hebert, Rahul Sukthankar |
CVPR | 1 |
| 2006 | Discovering Texture Regularity as a Higher-Order Correspondence Problem
James Hays, Marius Leordeanu, Alexei A. Efros, Yanxi Liu 0001 |
ECCV (2) | 2 |
| 2006 | Efficient MAP approximation for dense energy functionsabstractWe present an efficient method for maximizing energy functions with first and second order potentials, suitable for MAP labeling estimation problems that arise in undirected graphical models. Our approach is to relax the integer constraints on the solution in two steps. First we efficiently obtain the relaxed global optimum following a procedure similar to the iterative power method for finding the largest eigenvector of a matrix. Next, we map the relaxed optimum on a simplex and show that the new energy obtained has a certain optimal bound. Starting from this energy we follow an efficient coordinate ascent procedure that is guaranteed to increase the energy at every step and converge to a solution that obeys the initial integral constraints. We also present a sufficient condition for ascent procedures that guarantees the increase in energy at every step. Marius Leordeanu, Martial Hebert |
ICML | 1 |
| 2005 | Unsupervised Learning of Object Features from Video SequencesabstractWe develop an efficient algorithm for unsupervised learning of object models as constellations of features, from low resolution video sequences. The input images typically contain single or multiple objects that change in pose, scale and degree of occlusion. Also, the objects can move significantly between consecutive frames. The content of an input sequence is unlabeled so the learner has to cluster the data based on the data's implicit coherence over time and space. Our approach takes advantage of the dependent pairwise co-occurrences of objects' features within local neighborhoods vs. the independent behavior of unrelated features. We couple or decouple pairs of features based on a probabilistic interpretation of their pairwise statistics and then extract objects as connected components of features. Marius Leordeanu, Robert T. Collins |
CVPR (1) | 1 |
| 2005 | A Spectral Technique for Correspondence Problems Using Pairwise ConstraintsabstractWe present an efficient spectral method for finding consistent correspondences between two sets of features. We build the adjacency matrix M of a graph whose nodes represent the potential correspondences and the weights on the links represent pairwise agreements between potential correspondences. Correct assignments are likely to establish links among each other and thus form a strongly connected cluster. Incorrect correspondences establish links with the other correspondences only accidentally, so they are unlikely to belong to strongly connected clusters. We recover the correct assignments based on how strongly they belong to the main cluster of M, by using the principal eigenvector of M and imposing the mapping constraints required by the overall correspondence mapping (one-to-one or one-to-many). The experimental evaluation shows that our method is robust to outliers, accurate in terms of matching rate, while being much faster than existing methods Marius Leordeanu, Martial Hebert |
ICCV | 1 |
| 2005 | Online Selection of Discriminative Tracking FeaturesabstractThis paper presents an online feature selection mechanism for evaluating multiple features while tracking and adjusting the set of features used to improve tracking performance. Our hypothesis is that the features that best discriminate between object and background are also best for tracking the object. Given a set of seed features, we compute log likelihood ratios of class conditional sample densities from object and background to form a new set of candidate features tailored to the local object/background discrimination task. The two-class variance ratio is used to rank these new features according to how well they separate sample distributions of object and background pixels. This feature evaluation mechanism is embedded in a mean-shift tracking system that adaptively selects the top-ranked discriminative features for tracking. Examples are presented that demonstrate how this method adapts to changing appearances of both tracked object and scene background. We note susceptibility of the variance ratio feature selection method to distraction by spatially correlated background clutter and develop an additional approach that seeks to minimize the likelihood of distraction. Robert T. Collins, Yanxi Liu 0001, Marius Leordeanu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | Automated Feature-Based Range Registration of Urban Scenes of Large ScaleabstractWe are building a system that can automatically acquire 3D range scans and 2D images to build geometrically and photometrically correct 3D models of urban environments. A major bottleneck in the process is the automated registration of a large number of geometrically complex 3D range scans in a common frame of reference. In this paper we provide a method for the accurate and efficient registration of a large number of complex range scans. The method utilizes range segmentation and feature extraction algorithms. Our algorithm automatically computes pairwise registrations between individual scans, builds a topological graph, and places the scans in the same frame of reference. We present results for building large scale 3D models of historic sites and urban structures. Ioannis Stamos, Marius Leordeanu |
CVPR (2) | 2 |
| 2003 | 3D Modeling of Historic Sites Using Range and Image DataabstractPreserving cultural heritage and historic sites is an important problem. These sites are subject to erosion, vandalism, and as long-lived artifacts, they have gone through many phases of construction, damage and repair. It is important to keep an accurate record of these sites using 3-D model building technology as they currently are, so preservationists can track changes, foresee structural problems, and allow a wider audience to "virtually" see and tour these sites. Due to the complexity of these sites, building 3-D models is time consuming and difficult, usually involving much manual effort. This paper discusses new methods that can reduce the time to build a model using automatic methods. Examples of these methods are shown in reconstructing a model of the Cathedral of Ste. Pierre in Beauvais, France. Peter K. Allen, Ioannis Stamos, Alejandro J. Troccoli, Benjamin Smith 0001, Marius Leordeanu, Y. C. Hsu |
ICRA | 5 |