Nikos Komodakis

dblp:05/3782 · DBLP profile ↗
← Back
86ranked-venue papers
20as first author
13since 2021 · last 2025
0009-0000-6767-5641ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 62 · 16 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 54 · 14 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers
abstract
Semantic future prediction is important for autonomous systems navigating dynamic environments. This paper introduces FUTURIST, a method for multimodal future semantic prediction that uses a unified and efficient visual sequence transformer architecture. Our approach incorporates a multimodal masked visual modeling objective and a novel masking mechanism designed for multimodal training. This allows the model to effectively integrate visible information from various modalities, improving prediction accuracy. Additionally, we propose a VAE-free hierarchical tokenization process, which reduces computational complexity, streamlines the training pipeline, and enables end-to-end training with high-resolution, multimodal inputs. We validate FUTURIST on the Cityscapes dataset, demonstrating state-of-the-art performance in future semantic segmentation for both short- and mid-term forecasting. We provide the implementation code and model weights at https://github.com/Sta8is/FUTURIST.
Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis
CVPR4
2025 EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling
abstract
Latent generative models have emerged as a leading approach for high-quality image synthesis. These models rely on an autoencoder to compress images into a latent space, followed by a generative model to learn the latent distribution. We identify that existing autoencoders lack equivariance to semantic-preserving transformations like scaling and rotation, resulting in complex latent spaces that hinder generative performance. To address this, we propose EQ-VAE, a simple regularization approach that enforces equivariance in the latent space, reducing its complexity without degrading reconstruction quality. By finetuning pre-trained autoencoders with EQ-VAE, we enhance the performance of several state-of-the-art generative models, including DiT, SiT, REPA and MaskGIT, achieving a ×7 speedup on DiT-XL/2 with only five epochs of SD-VAE fine-tuning. EQ-VAE is compatible with both continuous and discrete autoencoders, thus offering a versatile enhancement for a wide range of latent generative models.
Theodoros Kouzelis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis
ICML4
2025 Diffusing Boundaries: CBCT-to-CT Translation with Extended Field of View
Quentin Spinat, Audrey Duran, Olivier Teboul, Nikos Paragios, Nikos Komodakis
MICCAI (13)5
2025 Multi-Token Prediction Needs Registers
abstract
Multi-token prediction has emerged as a promising objective for improving language model pretraining, but its benefits have not consistently generalized to other settings such as fine-tuning. In this paper, we propose MuToR, a simple and effective approach to multi-token prediction that interleaves learnable register tokens into the input sequence, each tasked with predicting future targets. Compared to existing methods, MuToR offers several key advantages: it introduces only a negligible number of additional parameters, requires no architectural changes—ensuring compatibility with off-the-shelf pretrained language models—and remains aligned with the next-token pretraining objective, making it especially well-suited for supervised fine-tuning. Moreover, it naturally supports scalable prediction horizons. We demonstrate the effectiveness and versatility of MuToR across a range of use cases, including supervised fine-tuning, parameter-efficient fine-tuning (PEFT), and pretraining, on challenging generative tasks in both language and vision domains.
Anastasios Gerontopoulos, Spyros Gidaris, Nikos Komodakis
NeurIPS3
2025 DINO-Foresight: Looking into the Future with DINO
abstract
Predicting future dynamics is crucial for applications like autonomous driving and robotics, where understanding the environment is key. Existing pixel-level methods are computationally expensive and often focus on irrelevant details. To address these challenges, we introduce DINO-Foresight, a novel framework that operates in the semantic feature space of pretrained Vision Foundation Models (VFMs). Our approach trains a masked feature transformer in a self-supervised manner to predict the evolution of VFM features over time. By forecasting these features, we can apply off-the-shelf, task-specific heads for various scene understanding tasks. In this framework, VFM features are treated as a latent space, to which different heads attach to perform specific tasks for future-frame analysis. Extensive experiments show the very strong performance, robustness and scalability of our framework.
Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis
NeurIPS4
2025 Boosting Generative Image Modeling via Joint Image-Feature Synthesis
abstract
Latent diffusion models (LDMs) dominate high-quality image generation, yet integrating representation learning with generative modeling remains a challenge. We introduce a novel generative image modeling framework that seamlessly bridges this gap by leveraging a diffusion model to jointly model low-level image latents (from a variational autoencoder) and high-level semantic features (from a pretrained self-supervised encoder like DINO). Our latent-semantic diffusion approach learns to generate coherent image-feature pairs from pure noise, significantly enhancing both generative quality and training efficiency, all while requiring only minimal modifications to standard Diffusion Transformer architectures. By eliminating the need for complex distillation objectives, our unified design simplifies training and unlocks a powerful new inference strategy: Representation Guidance, which leverages learned semantics to steer and refine image generation. Evaluated in both conditional and unconditional settings, our method delivers substantial improvements in image quality and training convergence speed, establishing a new direction for representation-aware generative modeling.
Theodoros Kouzelis, Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gidaris, Nikos Komodakis
NeurIPS5
2025 ReplaceMe: Network Simplification via Depth Pruning and Transformer Block Linearization
abstract
We introduce ReplaceMe, a generalized training-free depth pruning method that effectively replaces transformer blocks with a linear operation, while maintaining high performance for low compression ratios. In contrast to conventional pruning approaches that require additional training or fine-tuning, our approach requires only a small calibration dataset that is used to estimate a linear transformation, which approximates the pruned blocks. The estimated linear mapping can be seam- lessly merged with the remaining transformer blocks, eliminating the need for any additional network parameters. Our experiments show that ReplaceMe consistently outperforms other training-free approaches and remains highly competitive with state-of-the-art pruning methods that involve extensive retraining/fine-tuning and architectural modifications. Applied to several large language models (LLMs), ReplaceMe achieves up to 25% pruning while retaining approximately 90% of the original model’s performance on open benchmarks—without any training or healing steps, resulting in minimal computational overhead. We provide an open- source library implementing ReplaceMe alongside several state-of-the-art depth pruning techniques, available at https://github.com/mts-ai/ReplaceMe.
Dmitriy Shopkhoev, Ammar Ali, Magauiya Zhussip, Valentin Malykh, Stamatios Lefkimmiatis, Nikos Komodakis, Sergey Zagoruyko
NeurIPS6
2025 Deep learning detection of acute and sub-acute lesion activity from single-timepoint conventional brain MRI in multiple sclerosis
Quentin Spinat, Benoît Audelan, Bastien Caba, Alexis Benichoux, Despoina Ioannidou, Olivier Teboul, Nikos Komodakis, Willem Huijbers, Refaat Gabr, Arie Gafson, Colm Elliott, Douglas L. Arnold, Nikos Paragios, Shibeshih Mitiku Belachew
Medical Image Anal.8
2024 SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers
abstract
Unsupervised object-centric learning aims to decompose scenes into interpretable object entities, termed slots. Slot-based auto-encoders stand out as a prominent method for this task. Within them, crucial aspects include guiding the encoder to generate object-specific slots and ensuring the decoder utilizes them during reconstruction. This work introduces two novel techniques, (i) an attention-based self-training approach, which distills superior slot-based attention masks from the decoder to the encoder, enhancing object segmentation, and (ii) an innovative patch-order permutation strategy for autoregressive transformers that strengthens the role of slot vectors in reconstruction. The effectiveness of these strategies is showcased experimentally. The combined approach significantly surpasses prior slot-based autoencoder methods in unsupervised object segmentation, especially with complex real-world images. We provide the implementation code at https://github.com/gkakogeorgiou/spot.
Ioannis Kakogeorgiou, Spyros Gidaris, Konstantinos Karantzalos, Nikos Komodakis
CVPR4
2024 ToNNO: Tomographic Reconstruction of a Neural Network's Output for Weakly Supervised Segmentation of 3D Medical Images
abstract
Annotating lots of 3D medical images for training segmentation models is time-consuming. The goal of weakly supervised semantic segmentation is to train segmentation models without using any ground truth segmentation masks. Our work addresses the case where only image-level categorical labels, indicating the presence or absence of a particular region of interest (such as tumours or lesions), are available. Most existing methods rely on class activation mapping (CAM). We propose a novel approach, ToNNO, which is based on the Tomographic reconstruction of a Neural Network's Output. Our technique extracts stacks of slices with different angles from the input 3D volume, feeds these slices to a 2D encoder, and applies the inverse Radon transform in order to reconstruct a 3D heatmap of the encoder's predictions. This generic method allows to perform dense prediction tasks on 3D volumes using any 2D image encoder. We apply it to weakly supervised medical image segmentation by training the 2D encoder to output high values for slices containing the regions of interest. We test it on four large scale medical image datasets and outperform 2D CAM methods. We then extend ToNNO by combining tomographic reconstruction with CAM methods, proposing Averaged CAM and Tomographic CAM, which obtain even better results.
Marius Schmidt-Mengin, Alexis Benichoux, Shibeshih Mitiku Belachew, Nikos Komodakis, Nikos Paragios
CVPR4
2022 What to Hide from Your Students: Attention-Guided Masked Image Modeling
Ioannis Kakogeorgiou, Spyros Gidaris, Bill Psomas, Yannis Avrithis, Andrei Bursuc, Konstantinos Karantzalos, Nikos Komodakis
ECCV (30)7
2022 Self-supervised learning for medieval handwriting identification: A case study from the Vatican Apostolic Library
Lorenzo Lastilla, Serena Ammirati, Donatella Firmani, Nikos Komodakis, Paolo Merialdo, Simone Scardapane
Inf. Process. Manag.4
2021 OBoW: Online Bag-of-Visual-Words Generation for Self-Supervised Learning
abstract
Learning image representations without human supervision is an important and active research field. Several recent approaches have successfully leveraged the idea of making such a representation invariant under different types of perturbations, especially via contrastive-based instance discrimination training. Although effective visual representations should indeed exhibit such invariances, there are other important characteristics, such as encoding contextual reasoning skills, for which alternative reconstruction-based approaches might be better suited.With this in mind, we propose a teacher-student scheme to learn representations by training a convolutional net to reconstruct a bag-of-visual-words (BoW) representation of an image, given as input a perturbed version of that same image. Our strategy performs an online training of both the teacher network (whose role is to generate the BoW targets) and the student network (whose role is to learn representations), along with an online update of the visual-words vocabulary (used for the BoW targets). This idea effectively enables fully online BoW-guided unsupervised learning. Extensive experiments demonstrate the interest of our BoWbased strategy, which surpasses previous state-of-the-art methods (including contrastive-based ones) in several applications. For instance, in downstream tasks such Pascal object detection, Pascal classification and Places205 classification, our method improves over all prior unsupervised approaches, thus establishing new state-of-the-art results that are also significantly better even than those of supervised pre-training. We provide the implementation code at https://github.com/valeoai/obow.
Spyros Gidaris, Andrei Bursuc, Gilles Puy, Nikos Komodakis, Matthieu Cord, Patrick Pérez
CVPR4
2020 Learning Representations by Predicting Bags of Visual Words
abstract
Self-supervised representation learning targets to learn convnet-based image representations from unlabeled data. Inspired by the success of NLP methods in this area, in this work we propose a self-supervised approach based on spatially dense image descriptions that encode discrete visual concepts, here called visual words. To build such discrete representations, we quantize the feature maps of a first pre-trained self-supervised convnet, over a k-means based vocabulary. Then, as a self-supervised task, we train another convnet to predict the histogram of visual words of an image (i.e., its Bag-of-Words representation) given as input a perturbed version of that image. The proposed task forces the convnet to learn perturbation-invariant and context-aware image features, useful for downstream image understanding tasks. We extensively evaluate our method and demonstrate very strong empirical results, e.g., our pre-trained self-supervised representations transfer better on detection task and similarly on classification over classes "unseen'' during pre-training, when compared to the supervised case. This also shows that the process of image discretization into visual words can provide the basis for very powerful self-supervised approaches in the image domain, thus allowing further connections to be made to related methods from the NLP domain that have been extremely successful so far.
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, Matthieu Cord
CVPR3
2020 QuEST: Quantized Embedding Space for Transferring Knowledge
Himalaya Jain, Spyros Gidaris, Nikos Komodakis, Patrick Pérez, Matthieu Cord
ECCV (21)3
2020 Deep Tone Mapping Operator for High Dynamic Range Images
abstract
A computationally fast tone mapping operator (TMO) that can quickly adapt to a wide spectrum of high dynamic range (HDR) content is quintessential for visualization on varied low dynamic range (LDR) output devices such as movie screens or standard displays. Existing TMOs can successfully tone-map only a limited number of HDR content and require an extensive parameter tuning to yield the best subjective-quality tone-mapped output. In this paper, we address this problem by proposing a fast, parameter-free and scene-adaptable deep tone mapping operator (DeepTMO) that yields a high-resolution and high-subjective quality tone mapped output. Based on conditional generative adversarial network (cGAN), DeepTMO not only learns to adapt to vast scenic-content (e.g., outdoor, indoor, human, structures, etc.) but also tackles the HDR related scene-specific challenges such as contrast and brightness, while preserving the fine-grained details. We explore 4 possible combinations of Generator-Discriminator architectural designs to specifically address some prominent issues in HDR related deep-learning frameworks like blurring, tiling patterns and saturation artifacts. By exploring different influences of scales, loss-functions and normalization layers under a cGAN setting, we conclude with adopting a multi-scale model for our task. To further leverage on the large-scale availability of unlabeled HDR data, we train our network by generating targets using an objective HDR quality metric, namely Tone Mapping Image Quality Index (TMQI). We demonstrate results both quantitatively and qualitatively, and showcase that our DeepTMO generates high-resolution, high-quality output images over a large spectrum of real-world scenes. Finally, we evaluate the perceived quality of our results by conducting a pair-wise subjective study which confirms the versatility of our method.
Aakanksha Rana, Praveer Singh, Giuseppe Valenzise, Frédéric Dufaux, Nikos Komodakis, Aljoscha Smolic
IEEE Trans. Image Process.5
2019 Generating Classification Weights With GNN Denoising Autoencoders for Few-Shot Learning
abstract
Given an initial recognition model already trained on a set of base classes, the goal of this work is to develop a meta-model for few-shot learning. The meta-model, given as input some novel classes with few training examples per class, must properly adapt the existing recognition model into a new model that can correctly classify in a unified way both the novel and the base classes. To accomplish this goal it must learn to output the appropriate classification weight vectors for those two types of classes. To build our meta-model we make use of two main innovations: we propose the use of a Denoising Autoencoder network (DAE) that (during training) takes as input a set of classification weights corrupted with Gaussian noise and learns to reconstruct the target-discriminative classification weights. In this case, the injected noise on the classification weights serves the role of regularizing the weight generating meta-model. Furthermore, in order to capture the co-dependencies between different classes in a given task instance of our meta-model, we propose to implement the DAE model as a Graph Neural Network (GNN). In order to verify the efficacy of our approach, we extensively evaluate it on ImageNet based few-shot benchmarks and we report state-of-the-art results.
Spyros Gidaris, Nikos Komodakis
CVPR2
2019 Boosting Few-Shot Visual Learning With Self-Supervision
abstract
Few-shot learning and self-supervised learning address different facets of the same problem: how to train a model with little or no labeled data. Few-shot learning aims for optimization methods and models that can learn efficiently to recognize patterns in the low data regime. Self-supervised learning focuses instead on unlabeled data and looks into it for the supervisory signal to feed high capacity deep neural networks. In this work we exploit the complementarity of these two domains and propose an approach for improving few-shot learning through self-supervision. We use self-supervision as an auxiliary task in a few-shot learning pipeline, enabling feature extractors to learn richer and more transferable visual representations while still using few annotated samples. Through self-supervision, our approach can be naturally extended towards using diverse unlabeled data from other datasets in the few-shot setting. We report consistent improvements across an array of architectures, datasets and self-supervision techniques. We provide the implementation code at: https://github.com/valeoai/BF3S.
Spyros Gidaris, Andrei Bursuc, Nikos Komodakis, Patrick Pérez, Matthieu Cord
ICCV3
2019 Exploring weight symmetry in deep neural networks
Shell Xu Hu, Sergey Zagoruyko, Nikos Komodakis
Comput. Vis. Image Underst.3
2019 Scattering Networks for Hybrid Representation Learning
abstract
Scattering networks are a class of designed Convolutional Neural Networks (CNNs) with fixed weights. We argue they can serve as generic representations for modelling images. In particular, by working in scattering space, we achieve competitive results both for supervised and unsupervised learning tasks, while making progress towards constructing more interpretable CNNs. For supervised learning, we demonstrate that the early layers of CNNs do not necessarily need to be learned, and can be replaced with a scattering network instead. Indeed, using hybrid architectures, we achieve the best results with predefined representations to-date, while being competitive with end-to-end learned CNNs. Specifically, even applying a shallow cascade of small-windowed scattering coefficients followed by $1\times 1$1×1-convolutions results in AlexNet accuracy on the ILSVRC2012 classification task. Moreover, by combining scattering networks with deep residual networks, we achieve a single-crop top-5 error of 11.4 percent on ILSVRC2012. Also, we show they can yield excellent performance in the small sample regime on CIFAR-10 and STL-10 datasets, exceeding their end-to-end counterparts, through their ability to incorporate geometrical priors. For unsupervised learning, scattering coefficients can be a competitive representation that permits image recovery. We use this fact to train hybrid GANs to generate images. Finally, we empirically analyze several properties related to stability and reconstruction of images from scattering coefficients.
Edouard Oyallon, Sergey Zagoruyko, Gabriel Huang, Nikos Komodakis, Simon Lacoste-Julien, Matthew B. Blaschko, Eugene Belilovsky
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Dynamic Few-Shot Visual Learning Without Forgetting
abstract
The human visual system has the remarkably ability to be able to effortlessly learn novel concepts from only a few examples. Mimicking the same behavior on machine learning vision systems is an interesting and very challenging research problem with many practical advantages on real world vision applications. In this context, the goal of our work is to devise a few-shot visual learning system that during test time it will be able to efficiently learn novel categories from only a few training data while at the same time it will not forget the initial categories on which it was trained (here called base categories). To achieve that goal we propose (a) to extend an object recognition system with an attention based few-shot classification weight generator, and (b) to redesign the classifier of a ConvNet model as the cosine similarity function between feature representations and classification weight vectors. The latter, apart from unifying the recognition of both novel and base categories, it also leads to feature representations that generalize better on "unseen" categories. We extensively evaluate our approach on Mini-ImageNet where we manage to improve the prior state-of-the-art on few-shot recognition (i.e., we achieve 56.20% and 73.00% on the 1-shot and 5-shot settings respectively) while at the same time we do not sacrifice any accuracy on the base categories, which is a characteristic that most prior approaches lack. Finally, we apply our approach on the recently introduced few-shot benchmark of Bharath and Girshick [4] where we also achieve state-of-the-art results.
Spyros Gidaris, Nikos Komodakis
CVPR2
2018 GraphVAE: Towards Generation of Small Graphs Using Variational Autoencoders
Martin Simonovsky, Nikos Komodakis
ICANN (1)2
2018 Unsupervised Representation Learning by Predicting Image Rotations
Spyros Gidaris, Praveer Singh, Nikos Komodakis
ICLR (Poster)3
2018 Effective Building Extraction by Learning to Detect and Correct Erroneous Labels in Segmentation Mask
abstract
Semantic segmentation is pivotal for remote sensing image analysis. Although existing segmentation techniques perform well on similar landscape images, their generalization capability on an entirely different landscape is extremely poor. One of the primary reasons is that they partially or wholly, neglect the underlying relationship that exist in the joint space of input and output variables. Thus, effectively they lack to impose structure in their output predictions which is necessary for successful segmentation. In this paper, we address this problem and propose a novel solution by modeling the joint distribution of input-output variable which in turn enforces some structure in the initial segmentation mask. To this end, we first detect erroneous labels, in the form of Error maps, in the initial building masks. These Error maps are then used to correct the corresponding erroneous labels through a replacement technique. We evaluate our methodology on the benchmark Inria Aerial Image Labeling dataset, which is a large scale high resolution dataset for building footprint segmentation. In contrast to previous methods, our predicted segmentation masks are much closer to ground truth, owning to the fact that they are able to effectively correct both the large errors as well as the blobby effects. We lastly perform on par with other state-of-the-arts, validating the efficacy of our technique.
Praveer Singh, Nikos Komodakis
IGARSS2
2018 Cloud-Gan: Cloud Removal for Sentinel-2 Imagery Using a Cyclic Consistent Generative Adversarial Networks
abstract
Cloud cover is a serious impediment in land surface analysis from Remote Sensing images either causing complete obstruction (thick clouds) with loss of information or blurry effects when being semi-transparent (thin clouds). While thick clouds require complete pixel replacement, thin cloud removal is fairly challenging as the atmospheric and land-cover information is inter-twined. In this paper, we address this problem and propose a Cloud-GAN to learn the mapping between cloudy images and cloud-free images. The adver-sarialloss in the proposed method constrains the distribution of generated images to be close enough to the underlying distribution of the non-cloudy images. An additional cycle consistency loss is used to further restrain the generator to predict cloud-free images only of the same scene as reflected in the cloudy images. Our method not only rejects the necessity of any paired (cloud/cloud-free) training dataset but also avoids the need of any additional (expensive) spectral source of information such as Synthetic Aperture Radar imagery which is cloud penetrable. Lastly, we demonstrate the efficacy of our technique by training on an openly available and fairly new Sentinel-2 Imagery dataset consisting of real clouds. We also show significant improvement in PSNR values after removing clouds on synthetic images thus validating the competency of our methodology.
Praveer Singh, Nikos Komodakis
IGARSS2
2018 Improving Recognition of Complex Aerial Scenes Using a Deep Weakly Supervised Learning Paradigm
abstract
Categorizing highly complex aerial scenes is quite strenuous due to the presence of detailed information with a large number of distinctive objects. Recognition happens by first deriving a joint relationship within all these distinguishing objects, distilling finally to some meaningful knowledge that is subsequently employed to label the scene. However, something intriguing is whether all this captured information is actually relevant to classify such a complex scene? What if some objects just create uncertainty with respect to the target label, thereby causing ambiguity in the decision-making? In this letter, we investigate these questions and analyze as to which regions in an aerial scene are the most relevant and are inhibiting in determining the image label accurately. However, for such aerial scene classification (ASC) task, employing supervised knowledge of experts to annotate these discriminative regions is quite costly and laborious, especially when the data set is huge. To this end, we propose a deep weakly supervised learning (DWSL) technique. Our classification-trained convolutional neural network learns to identify discriminative region localizations in an aerial scene solely by utilizing image labels. Using the DWSL model, we significantly improve the recognition accuracies of highly complex scenes, thus validating that extra information causes uncertainty in decision-making. Moreover, our DWSL methodology can also be leveraged as a novel tool for concrete visualization of the most informative regions relevant to accurately classify an aerial scene. Finally, our proposed framework yields a state-of-the-art performance on the existing ASC data sets.
Praveer Singh, Nikos Komodakis
IEEE Geosci. Remote. Sens. Lett.2
2018 SCOM: Spatiotemporal Constrained Optimization for Salient Object Detection
abstract
This paper presents a novel model for video salient object detection called spatiotemporal constrained optimization model (SCOM), which exploits spatial and temporal cues, as well as a local constraint, to achieve a global saliency optimization. For a robust motion estimation of salient objects, we propose a novel approach to modeling the motion cues from optical flow field, the saliency map of the prior video frame and the motion history of change detection, which is able to distinguish the moving salient objects from diverse changing background regions. Furthermore, an effective objectness measure is proposed with intuitive geometrical interpretation to extract some reliable object and background regions, which provided as the basis to define the foreground potential, background potential, and the constraint to support saliency propagation. These potentials and the constraint are formulated into the proposed SCOM framework to generate an optimal saliency map for each frame in a video. The proposed model is extensively evaluated on the widely used challenging benchmark data sets. Experiments demonstrate that our proposed SCOM substantially outperforms the state-of-the-art saliency models.
Yuhuan Chen, Wenbin Zou, Yi Tang 0008, Xia Li 0006, Chen Xu 0004, Nikos Komodakis
IEEE Trans. Image Process.6
2017 Detect, Replace, Refine: Deep Structured Prediction for Pixel Wise Labeling
abstract
Pixel wise image labeling is an interesting and challenging problem with great significance in the computer vision community. In order for a dense labeling algorithm to be able to achieve accurate and precise results, it has to consider the dependencies that exist in the joint space of both the input and the output variables. An implicit approach for modeling those dependencies is by training a deep neural network that, given as input an initial estimate of the output labels and the input image, it will be able to predict a new refined estimate for the labels. In this context, our work is concerned with what is the optimal architecture for performing the label improvement task. We argue that the prior approaches of either directly predicting new label estimates or predicting residual corrections w.r.t. the initial labels with feed-forward deep network architectures are sub-optimal. Instead, we propose a generic architecture that decomposes the label improvement task to three steps: 1) detecting the initial label estimates that are incorrect, 2) replacing the incorrect labels with new ones, and finally 3) refining the renewed labels by predicting residual corrections w.r.t. them. Furthermore, we explore and compare various other alternative architectures that consist of the aforementioned Detection, Replace, and Refine components. We extensively evaluate the examined architectures in the challenging task of dense disparity estimation (stereo matching) and we report both quantitative and qualitative results on three different datasets. Finally, our dense disparity estimation network that implements the proposed generic architecture, achieves state-of-the-art results in the KITTI 2015 test surpassing prior approaches by a significant margin. We plan to release the Torch code that implements the paper in: https://github.com/gidariss/DRR_struct_pred/.
Spyros Gidaris, Nikos Komodakis
CVPR2
2017 Newton-Type Methods for Inference in Higher-Order Markov Random Fields
abstract
Linear programming relaxations are central to MAP inference in discrete Markov Random Fields. The ability to properly solve the Lagrangian dual is a critical component of such methods. In this paper, we study the benefit of using Newton-type methods to solve the Lagrangian dual of a smooth version of the problem. We investigate their ability to achieve superior convergence behavior and to better handle the ill-conditioned nature of the formulation, as compared to first order methods. We show that it is indeed possible to efficiently apply a trust region Newton method for a broad range of MAP inference problems. In this paper we propose a provably globally efficient framework that includes (i) excellent compromise between computational complexity and precision concerning the Hessian matrix construction, (ii) a damping strategy that aids efficient optimization, (iii) a truncation strategy coupled with a generic pre-conditioner for Conjugate Gradients, (iv) efficient sum-product computation for sparse clique potentials. Results for higher-order Markov Random Fields demonstrate the potential of this approach.
Hariprasad Kannan, Nikos Komodakis, Nikos Paragios
CVPR2
2017 Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs
abstract
A number of problems can be formulated as prediction on graph-structured data. In this work, we generalize the convolution operator from regular grids to arbitrary graphs while avoiding the spectral domain, which allows us to handle graphs of varying size and connectivity. To move beyond a simple diffusion, filter weights are conditioned on the specific edge labels in the neighborhood of a vertex. Together with the proper choice of graph coarsening, we explore constructing deep neural networks for graph classification. In particular, we demonstrate the generality of our formulation in point cloud classification, where we set the new state of the art, and on a graph classification dataset, where we outperform other deep learning approaches.
Martin Simonovsky, Nikos Komodakis
CVPR2
2017 Rotation Equivariant Vector Field Networks
abstract
In many computer vision tasks, we expect a particular behavior of the output with respect to rotations of the input image. If this relationship is explicitly encoded, instead of treated as any other variation, the complexity of the problem is decreased, leading to a reduction in the size of the required model. In this paper, we propose the Rotation Equivariant Vector Field Networks (RotEqNet), a Convolutional Neural Network (CNN) architecture encoding rotation equivariance, invariance and covariance. Each convolutional filter is applied at multiple orientations and returns a vector field representing magnitude and angle of the highest scoring orientation at every spatial location. We develop a modified convolution operator relying on this representation to obtain deep architectures. We test RotEqNet on several problems requiring different responses with respect to the inputs' rotation: image classification, biomedical image segmentation, orientation estimation and patch matching. In all cases, we show that RotEqNet offers extremely compact models in terms of number of parameters and provides results in line to those of networks orders of magnitude larger.
Diego Marcos, Michele Volpi, Nikos Komodakis, Devis Tuia
ICCV3
2017 Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer
Sergey Zagoruyko, Nikos Komodakis
ICLR (Poster)2
2017 Patchwork Stereo: Scalable, Structure-Aware 3D Reconstruction in Man-Made Environments
abstract
In this paper, we address the problem of Multi-View Stereo (MVS) reconstruction of highly regular man-made scenes from calibrated, wide-baseline views and a sparse Structure-from-Motion (SfM) point cloud. We introduce a novel patch-based formulation via energy minimization which combines top-down segmentation hypotheses using appearance and vanishing line detections, as well as an arrangement of creased planar structures which are extracted automatically through a robust analysis of available SfM points and image features. The method produces a compact piecewise-planar depth map and a mesh which are aligned with the scene's structure. Experiments show that our approach not only reaches similar levels of accuracy w.r.t state-of-the-art pixel-based methods while using much fewer images, but also produces a much more compact, structure-aware mesh in a considerably shorter runtime by several of orders of magnitude.
Amine Bourki, Martin de La Gorce, Renaud Marlet, Nikos Komodakis
WACV4
2017 Deep compare: A study on using convolutional neural networks to compare image patches
Sergey Zagoruyko, Nikos Komodakis
Comput. Vis. Image Underst.2
2017 Hierarchical Saliency Detection via Probabilistic Object Boundaries
abstract
Though there are many computational models proposed for saliency detection, few of them take object boundary information into account. This paper presents a hierarchical saliency detection model incorporating probabilistic object boundaries, which is based on the observation that salient objects are generally surrounded by explicit boundaries and show contrast with their surroundings. We perform adaptive thresholding operation on ultrametric contour map, which leads to hierarchical image segmentations, and compute the saliency map for each layer based on the proposed robust center bias, border bias, color dissimilarity and spatial coherence measures. After a linear weighted combination of multi-layer saliency maps, and Bayesian enhancement procedure, the final saliency map is obtained. Extensive experimental results on three challenging benchmark datasets demonstrate that the proposed model outperforms eight state-of-the-art saliency detection models.
Haijun Lei, Hai Xie, Wenbin Zou, Kidiyo Kpalma, Nikos Komodakis
Int. J. Pattern Recognit. Artif. Intell.6
2017 A Robust and Efficient Approach to License Plate Detection
abstract
This paper presents a robust and efficient method for license plate detection with the purpose of accurately localizing vehicle license plates from complex scenes in real time. A simple yet effective image downscaling method is first proposed to substantially accelerate license plate localization without sacrificing detection performance compared with that achieved using the original image. Furthermore, a novel line density filter approach is proposed to extract candidate regions, thereby significantly reducing the area to be analyzed for license plate localization. Moreover, a cascaded license plate classifier based on linear support vector machines using color saliency features is introduced to identify the true license plate from among the candidate regions. For performance evaluation, a data set consisting of 3977 images captured from diverse scenes under different conditions is also presented. Extensive experiments on the widely used Caltech license plate data set and our newly introduced data set demonstrate that the proposed approach substantially outperforms state-of-the-art methods in terms of both detection accuracy and run-time efficiency, increasing the detection ratio from 91.09% to 96.62% while decreasing the run time from 672 to 42 ms for processing an image with a resolution of 1082×728 . The executable code and our collected data set are publicly available.
Yule Yuan, Wenbin Zou, Yong Zhao 0010, Xin'an Wang, Xuefeng Hu, Nikos Komodakis
IEEE Trans. Image Process.6
2016 Attend Refine Repeat: Active Box Proposal Generation via In-Out Localization
Spyros Gidaris, Nikos Komodakis
BMVC2
2016 OnionNet: Sharing Features in Cascaded Deep Classifiers
Martin Simonovsky, Nikos Komodakis
BMVC2
2016 Wide Residual Networks
Sergey Zagoruyko, Nikos Komodakis
BMVC2
2016 LocNet: Improving Localization Accuracy for Object Detection
abstract
We propose a novel object localization methodology with the purpose of boosting the localization accuracy of stateof-the-art object detection systems. Our model, given a search region, aims at returning the bounding box of an object of interest inside this region. To accomplish its goal, it relies on assigning conditional probabilities to each row and column of this region, where these probabilities provide useful information regarding the location of the boundaries of the object inside the search region and allow the accurate inference of the object bounding box under a simple probabilistic framework. For implementing our localization model, we make use of a convolutional neural network architecture that is properly adapted for this task, called LocNet. We show experimentally that LocNet achieves a very significant improvement on the mAP for high IoU thresholds on PASCAL VOC2007 test set and that it can be very easily coupled with recent stateof-the-art object detection systems, helping them to boost their performance. Finally, we demonstrate that our detection approach can achieve high detection accuracy even when it is given as input a set of sliding windows, thus proving that it is independent of box proposal methods.
Spyros Gidaris, Nikos Komodakis
CVPR2
2016 A Deep Metric for Multimodal Registration
abstract
Multimodal registration is a challenging problem due the high variability of tissue appearance under different imaging modalities. The crucial component here is the choice of the right similarity measure. We make a step towards a general learning-based solution that can be adapted to specific situations and present a metric based on a convolutional neural network. Our network can be trained from scratch even from a few aligned image pairs. The metric is validated on intersubject deformable registration on a dataset different from the one used for training, demonstrating good generalization. In this task, we outperform mutual information by a significant margin.
Martin Simonovsky, Benjamín Gutiérrez-Becker, Diana Mateus, Nassir Navab, Nikos Komodakis
MICCAI (3)5
2016 Inference and Learning of Graphical Models: Theory and Applications in Computer Vision and Image Analysis
Chaohui Wang, Nikos Komodakis, Hiroshi Ishikawa 0002, Olga Veksler, Endre Boros
Comput. Vis. Image Underst.2
2016 (Hyper)-graphical models in biomedical image analysis
Nikos Paragios, Enzo Ferrante, Ben Glocker, Nikos Komodakis, Sarah Parisot, Evangelia I. Zacharaki
Medical Image Anal.4
2015 Learning to compare image patches via convolutional neural networks
abstract
In this paper we show how to learn directly from image data (i.e., without resorting to manually-designed features) a general similarity function for comparing image patches, which is a task of fundamental importance for many computer vision problems. To encode such a function, we opt for a CNN-based model that is trained to account for a wide variety of changes in image appearance. To that end, we explore and study multiple neural network architectures, which are specifically adapted to this task. We show that such an approach can significantly outperform the state-of-the-art on several problems and benchmark datasets.
Sergey Zagoruyko, Nikos Komodakis
CVPR2
2015 Object Detection via a Multi-region and Semantic Segmentation-Aware CNN Model
abstract
We propose an object detection system that relies on a multi-region deep convolutional neural network (CNN) that also encodes semantic segmentation-aware features. The resulting CNN-based representation aims at capturing a diverse set of discriminative appearance factors and exhibits localization sensitivity that is essential for accurate object localization. We exploit the above properties of our recognition module by integrating it on an iterative localization mechanism that alternates between scoring a box proposal and refining its location with a deep CNN regression model. Thanks to the efficient use of our modules, we detect objects with very high localization accuracy. On the detection challenges of PASCAL VOC2007 and PASCAL VOC2012 we achieve mAP of 78.2% and 73.9% correspondingly, surpassing any other published work by a significant margin.
Spyros Gidaris, Nikos Komodakis
ICCV2
2015 HARF: Hierarchy-Associated Rich Features for Salient Object Detection
abstract
The state-of-the-art salient object detection models are able to perform well for relatively simple scenes, yet for more complex ones, they still have difficulties in highlighting salient objects completely from background, largely due to the lack of sufficiently robust features for saliency prediction. To address such an issue, this paper proposes a novel hierarchy-associated feature construction framework for salient object detection, which is based on integrating elementary features from multi-level regions in a hierarchy. Furthermore, multi-layered deep learning features are introduced and incorporated as elementary features into this framework through a compact integration scheme. This leads to a rich feature representation, which is able to represent the context of the whole object/background and is much more discriminative as well as robust for salient object detection. Extensive experiments on the most widely used and challenging benchmark datasets demonstrate that the proposed approach substantially outperforms the state-of-the-art on salient object detection.
Wenbin Zou, Nikos Komodakis
ICCV2
2015 Building detection in very high resolution multispectral data with deep learning features
abstract
The automated man-made object detection and building extraction from single satellite images is, still, one of the most challenging tasks for various urban planning and monitoring engineering applications. To this end, in this paper we propose an automated building detection framework from very high resolution remote sensing data based on deep convolutional neural networks. The core of the developed method is based on a supervised classification procedure employing a very large training dataset. An MRF model is then responsible for obtaining the optimal labels regarding the detection of scene buildings. The experimental results and the performed quantitative validation indicate the quite promising potentials of the developed approach.
Maria Vakalopoulou, Konstantinos Karantzalos, Nikos Komodakis, Nikos Paragios
IGARSS3
2015 A Comparative Study of Modern Inference Techniques for Structured Discrete Energy Minimization Problems
Jörg H. Kappes, Bjoern Andres, Fred A. Hamprecht, Christoph Schnörr, Sebastian Nowozin, Dhruv Batra, Sungwoong Kim, Bernhard X. Kausler, Thorben Kröger, Jan Lellmann, Nikos Komodakis, Bogdan Savchynskyy, Carsten Rother
Int. J. Comput. Vis.11
2015 A Framework for Efficient Structured Max-Margin Learning of High-Order MRF Models
abstract
We present a very general algorithm for structured prediction learning that is able to efficiently handle discrete MRFs/CRFs (including both pairwise and higher-order models) so long as they can admit a decomposition into tractable subproblems. At its core, it relies on a dual decomposition principle that has been recently employed in the task of MRF optimization. By properly combining such an approach with a max-margin learning method, the proposed framework manages to reduce the training of a complex high-order MRF to the parallel training of a series of simple slave MRFs that are much easier to handle. This leads to a very efficient and general learning scheme that relies on solid mathematical principles. We thoroughly analyze its theoretical properties, and also show that it can yield learning algorithms of increasing accuracy since it naturally allows a hierarchy of convex relaxations to be used for loss-augmented MAP-MRF inference within a max-margin learning approach. Furthermore, it can be easily adapted to take advantage of the special structure that may be present in a given class of MRFs. We demonstrate the generality and flexibility of our approach by testing it on a variety of scenarios, including training of pairwise and higher-order MRFs, training by using different types of regularizers and/or different types of dissimilarity loss functions, as well as by learning of appropriate models for a variety of vision tasks (including high-order models for compact pose-invariant shape priors, knowledge-based segmentation, image denoising, stereo matching as well as high-order Potts MRFs).
Nikos Komodakis, Bo Xiang, Nikos Paragios
IEEE Trans. Pattern Anal. Mach. Intell.1
2015 Unsupervised Joint Salient Region Detection and Object Segmentation
abstract
This paper presents a novel unsupervised algorithm to detect salient regions and to segment out foreground objects from background. In contrast to previous unidirectional saliency-based object segmentation methods, in which only the detected saliency map is used to guide the object segmentation, our algorithm mutually exploits detection/segmentation cues from each other. To achieve this goal, an initial saliency map is generated by the proposed segmentation driven low-rank matrix recovery model. Such a saliency map is exploited to initialize object segmentation model, which is formulated as energy minimization of Markov random field. Mutually, the quality of saliency map is further improved by the segmentation result, and serves as a new guidance for the object segmentation. The optimal saliency map and the final segmentation are achieved by jointly optimizing the defined objective functions. Extensive evaluations on MSRA-B and PASCAL-1500 datasets demonstrate that the proposed algorithm achieves the state-of-the-art performance for both the salient region detection and the object segmentation.
Wenbin Zou, Zhi Liu 0003, Kidiyo Kpalma, Joseph Ronsin, Yong Zhao 0010, Nikos Komodakis
IEEE Trans. Image Process.6
2014 Learning to Detect Ground Control Points for Improving the Accuracy of Stereo Matching
abstract
While machine learning has been instrumental to the ongoing progress in most areas of computer vision, it has not been applied to the problem of stereo matching with similar frequency or success. We present a supervised learning approach for predicting the correctness of stereo matches based on a random forest and a set of features that capture various forms of information about each pixel. We show highly competitive results in predicting the correctness of matches and in confidence estimation, which allows us to rank pixels according to the reliability of their assigned disparities. Moreover, we show how these confidence values can be used to improve the accuracy of disparity maps by integrating them with an MRF-based stereo algorithm. This is an important distinction from current literature that has mainly focused on sparsification by removing potentially erroneous disparities to generate quasi-dense disparity maps.
Aristotle Spyropoulos, Nikos Komodakis, Philippos Mordohai
CVPR2
2014 A MAP-Estimation Framework for Blind Deblurring Using High-Level Edge Priors
Yipin Zhou, Nikos Komodakis
ECCV (2)2
2014 Discrete Visual Perception
abstract
Computational vision and biomedical image have made tremendous progress of the past decade. This is mostly due the development of efficient learning and inference algorithms which allow better, faster and richer modeling of visual perception tasks. Graph-based representations are among the most prominent tools to address such perception through the casting of perception as a graph optimization problem. In this paper, we briefly introduce the interest of such representations, discuss their strength and limitations and present their application to address a variety of problems in computer vision and biomedical image analysis.
Nikos Paragios, Nikos Komodakis
ICPR2
2014 Inference by Learning: Speeding-up Graphical Model Optimization via a Coarse-to-Fine Cascade of Pruning Classifiers
Bruno Conejo, Nikos Komodakis, Sébastien Leprince, Jean-Philippe Avouac
NIPS2
2013 Interactive Image Segmentation via Graph Clustering and Synthetic Coordinates Modeling
Costas Panagiotakis, Harris Papadakis, Ilias Grinias, Nikos Komodakis, Paraskevi Fragopoulou, Georgios Tziritas
CAIP (1)4
2013 A Comparative Study of Modern Inference Techniques for Discrete Energy Minimization Problems
abstract
Even years ago, Szeliski et al. published an influential study on energy minimization methods for Markov random fields (MRF). This study provided valuable insights in choosing the best optimization technique for certain classes of problems. While these insights remain generally useful today, the phenominal success of random field models means that the kinds of inference problems we solve have changed significantly. Specifically, the models today often include higher order interactions, flexible connectivity structures, large label-spaces of different cardinalities, or learned energy tables. To reflect these changes, we provide a modernized and enlarged study. We present an empirical comparison of 24 state-of-art techniques on a corpus of 2,300 energy minimization instances from 20 diverse computer vision applications. To ensure reproducibility, we evaluate all methods in the OpenGM2 framework and report extensive results regarding runtime and solution quality. Key insights from our study agree with the results of Szeliski et al. for the types of models they studied. However, on new and challenging types of models our findings disagree and suggest that polyhedral methods and integer programming solvers are competitive in terms of runtime and solution quality over a large range of model types.
Jörg H. Kappes, Bjoern Andres, Fred A. Hamprecht, Christoph Schnörr, Sebastian Nowozin, Dhruv Batra, Sungwoong Kim, Bernhard X. Kausler, Jan Lellmann, Nikos Komodakis, Carsten Rother
CVPR10
2013 Markov Random Field modeling, inference & learning in computer vision & image understanding: A survey
Chaohui Wang, Nikos Komodakis, Nikos Paragios
Comput. Vis. Image Underst.2
2013 Interactive image segmentation based on synthetic graph coordinates
Costas Panagiotakis, Harris Papadakis, Ilias Grinias, Nikos Komodakis, Paraskevi Fragopoulou, Georgios Tziritas
Pattern Recognit.4
2012 MRF-Based Blind Image Deconvolution
Nikos Komodakis, Nikos Paragios
ACCV (3)1
2011 Efficient training for pairwise or higher order CRFs via dual decomposition
abstract
We present a very general algorithmic framework for structured prediction learning that is able to efficiently handle both pairwise and higher-order discrete MRFs/CRFs1. It relies on a dual decomposition approach that has been recently proposed for MRF optimization. By properly combining this approach with a max-margin method, our framework manages to reduce the training of a complex high-order MRF to the parallel training of a series of simple slave MRFs that are much easier to handle. This leads to an extremely efficient and general learning scheme. Furthermore, the proposed framework can yield learning algorithms of increasing accuracy since it naturally allows a hierarchy of convex relaxations to be used for MRF inference within a max-margin learning approach. It also offers extreme flexibility and can be easily adapted to take advantage of any special structure of a given class of MRFs. Experimental results demonstrate the great effectiveness of our method.
Nikos Komodakis
CVPR1
2011 Learning to cluster using high order graphical models with latent variables
abstract
This paper proposes a very general max-margin learning framework for distance-based clustering. To this end, it formulates clustering as a high order energy minimization problem with latent variables, and applies a dual decomposition approach for training this model. The resulting framework allows learning a very broad class of distance functions, permits an automatic determination of the number of clusters during testing, and is also very efficient. As an additional contribution, we show how our method can be generalized to handle the training of a very broad class of important models in computer vision: arbitrary high-order latent CRFs. Experimental results verify its effectiveness.
Nikos Komodakis
ICCV1
2011 Special issue on Optimization for vision, graphics and medical imaging: Theory and applications
Nikos Komodakis, Georg Langs, Horst Bischof, Nikos Paragios
Comput. Vis. Image Underst.1
2011 MRF Energy Minimization and Beyond via Dual Decomposition
abstract
This paper introduces a new rigorous theoretical framework to address discrete MRF-based optimization in computer vision. Such a framework exploits the powerful technique of Dual Decomposition. It is based on a projected subgradient scheme that attempts to solve an MRF optimization problem by first decomposing it into a set of appropriately chosen subproblems, and then combining their solutions in a principled way. In order to determine the limits of this method, we analyze the conditions that these subproblems have to satisfy and demonstrate the extreme generality and flexibility of such an approach. We thus show that by appropriately choosing what subproblems to use, one can design novel and very powerful MRF optimization algorithms. For instance, in this manner we are able to derive algorithms that: 1) generalize and extend state-of-the-art message-passing methods, 2) optimize very tight LP-relaxations to MRF optimization, and 3) take full advantage of the special structure that may exist in particular MRFs, allowing the use of efficient inference techniques such as, e.g., graph-cut-based methods. Theoretical analysis on the bounds related with the different algorithms derived from our framework and experimental results/comparisons using synthetic and real data for a variety of tasks in computer vision demonstrate the extreme potentials of our approach.
Nikos Komodakis, Nikos Paragios, Georgios Tziritas
IEEE Trans. Pattern Anal. Mach. Intell.1
2010 Towards More Efficient and Effective LP-Based Algorithms for MRF Optimization
Nikos Komodakis
ECCV (2)1
2010 Flooding and MRF-based Algorithms for Interactive Segmentation
abstract
We propose a method for interactive colour image segmentation. The goal is to detect an object from the background, when some markers on object(s) and the background are given. As features only probability distributions of the data are used. At first, all the labelled seeds are independently propagated for obtaining homogeneous connected components for each of them. Then the image is divided in blocks, which are classified according to their probabilistic distance from the classified regions. A topographic surface for each class is obtained, using Bayesian dissimilarities and a min-max criterion. Two algorithms are proposed: a regularized classification based on the topographic surface and incorporating an MRF model, and a priority multi-label flooding algorithm. Segmentation results on the LHI data set are presented.
Ilias Grinias, Nikos Komodakis, Georgios Tziritas
ICPR2
2010 Linear intensity-based image registration by Markov random fields and discrete optimization
Darko Zikic, Ben Glocker, Oliver Kutter, Martin Groher, Nikos Komodakis, Ali Kamen, Nikos Paragios, Nassir Navab
Medical Image Anal.5
2009 Shape priors and discrete MRFs for knowledge-based segmentation
abstract
In this paper we introduce a new approach to knowledge-based segmentation. Our method consists of a novel representation to model shape variations as well as an efficient inference procedure to fit the model to new data. The considered shape model is similarity-invariant and refers to an incomplete graph that consists of intra and intercluster connections representing the inter-dependencies of control points. The clusters are determined according to the co-dependencies of the deformations of the control points within the training set. The connections between the components of a cluster represent the local structure while the connections between the clusters account for the global structure. The distributions of the normalized distances between the connected control points encode the prior model. During search, this model is used together with a discrete Markov random field (MRF) based segmentation, where the unknown variables are the positions of the control points in the image domain. To encode the image support, a Voronoi decomposition of the domain is considered and regional based statistics are used. The resulting model is computationally efficient, can encode complex statistical models of shape variations and benefits from the image support of the entire spatial domain.
Ahmed Besbes, Nikos Komodakis, Georg Langs, Nikos Paragios
CVPR2
2009 Discrete tracking of parametrized curves
abstract
A novel scheme for deformable tracking of curvilinear structures in image sequences is presented. The approach is based on B-spline snakes defined by a set of control points whose optimal configuration is determined through efficient discrete optimization. Each control point is associated with a discrete random variable in a MAP-MRF formulation where a set of labels captures the deformation space. In such a context, generic terms are encoded within this MRF in the form of pairwise potentials. The use of pairwise potentials along with the B-spline representation offers nearly perfect approximation of the continuous domain. Efficient linear programming is considered to recover the approximate optimal solution. The method is successfully applied to the tracking of guide-wires in fluoroscopic X-ray sequences of several hundred frames which requires extremely robust techniques.
Tim Hauke Heibel, Ben Glocker, Martin Groher, Nikos Paragios, Nikos Komodakis, Nassir Navab
CVPR5
2009 Beyond pairwise energies: Efficient optimization for higher-order MRFs
abstract
In this paper, we introduce a higher-order MRF optimization framework. On the one hand, it is very general; we thus use it to derive a generic optimizer that can be applied to almost any higher-order MRF and that provably optimizes a dual relaxation related to the input MRF problem. On the other hand, it is also extremely flexible and thus can be easily adapted to yield far more powerful algorithms when dealing with subclasses of high-order MRFs. We thus introduce a new powerful class of high-order potentials, which are shown to offer enough expressive power and to be useful for many vision tasks. To address them, we derive, based on the same framework, a novel and extremely efficient message-passing algorithm, which goes beyond the aforementioned generic optimizer and is able to deliver almost optimal solutions of very high quality. Experimental results on vision problems demonstrate the extreme effectiveness of our approach. For instance, we show that in some cases we are even able to compute the global optimum for NP-hard higher-order MRFs in a very efficient manner.
Nikos Komodakis, Nikos Paragios
CVPR1
2009 Graphical Models and Deformable Diffeomorphic Population Registration Using Global and Local Metrics
Aristeidis Sotiras, Nikos Komodakis, Ben Glocker, Jean-François Deux, Nikos Paragios
MICCAI (1)2
2009 Real-time exploration and photorealistic reconstruction of large natural environments
Nikos Komodakis, Georgios Tziritas
Vis. Comput.1
2008 Commentary Paper on "Textural Segmentation of Sidescan Sonar Images Based on Gabor Filters Bank and Active Contours without Edges"
abstract
In this paper the authors present a technique suitable for segmenting high resolution sidescan sonar images, which is a problem that may arise in military or civilian applications. To this end, they use the "active contours without edges" model that was proposed by Chan and Vese.
Nikos Komodakis, George Bebis
AVSS1
2008 Optical flow estimation with uncertainties through dynamic MRFs
abstract
In this paper, we propose a novel dynamic discrete framework to address image morphing with application to optical flow estimation. We reformulate the problem using a number of discrete displacements, and therefore the estimation of the morphing parameters becomes a tractable matching criteria independent combinatorial problem which is solved through the FastPD algorithm. In order to overcome the main limitation of discrete approaches (low dimensionality of the label space is unable to capture the continuous nature of the expected solution), we introduce a dynamic behavior in the model where the plausible discrete deformations (displacements) are varying in space (across the domain) and time (different states of the process - successive morphing states) according to the local uncertainty of the obtained solution.
Ben Glocker, Nikos Paragios, Nikos Komodakis, Georgios Tziritas, Nassir Navab
CVPR3
2008 Beyond Loose LP-Relaxations: Optimizing MRFs by Repairing Cycles
Nikos Komodakis, Nikos Paragios
ECCV (3)1
2008 Deformable Mosaicing for Whole-Body MRI
Christian Wachinger, Ben Glocker, Jochen Zeltner, Nikos Paragios, Nikos Komodakis, Michael Sass Hansen, Nassir Navab
MICCAI (2)5
2008 Clustering via LP-based Stabilities
abstract
A novel center-based clustering algorithm is proposed in this paper. We first formulate clustering as an NP-hard linear integer program and we then use linear programming and the duality theory to derive the solution of this optimization problem. This leads to an efficient and very general algorithm, which works in the dual domain, and can cluster data based on an arbitrary set of distances. Despite its generality, it is independent of initialization (unlike EM-like methods such as K-means), has guaranteed convergence, and can also provide online optimality bounds about the quality of the estimated clustering solutions. To deal with the most critical issue in a center-based clustering algorithm (selection of cluster centers), we also introduce the notion of stability of a cluster center, which is a well defined LP-based quantity that plays a key role to our algorithm's success. Furthermore, we also introduce, what we call, the margins (another key ingredient in our algorithm), which can be roughly thought of as dual counterparts to stabilities and allow us to obtain computationally efficient approximations to the latter. Promising experimental results demonstrate the potentials of our method.
Nikos Komodakis, Nikos Paragios, Georgios Tziritas
NIPS1
2008 Performance vs computational efficiency for optimizing single and dynamic MRFs: Setting the state of the art with primal-dual strategies
Nikos Komodakis, Georgios Tziritas, Nikos Paragios
Comput. Vis. Image Underst.1
2008 Dense image registration through MRFs and efficient linear programming
Ben Glocker, Nikos Komodakis, Georgios Tziritas, Nassir Navab, Nikos Paragios
Medical Image Anal.2
2007 Fast, Approximately Optimal Solutions for Single and Dynamic MRFs
abstract
A new efficient MRF optimization algorithm, called Fast-PD, is proposed, which generalizes α-expansion. One of its main advantages is that it offers a substantial speedup over that method, e.g. it can be at least 3-9 times faster than α-expansion. Its efficiency is a result of the fact that Fast-PD exploits information coming not only from the original MRF problem, but also from a dual problem. Furthermore, besides static MRFs, it can also be used for boosting the performance of dynamic MRFs, i.e. MRFs varying over time. On top of that, Fast-PD makes no compromise about the optimality of its solutions: it can compute exactly the same answer as a-expansion, but, unlike that method, it can also guarantee an almost optimal solution for a much wider class of NP-hard MRF problems. Results on static and dynamic MRFs demonstrate the algorithm's efficiency and power. E.g., Fast-PD has been able to compute disparity for stereoscopic sequences in real time, with the resulting disparity coinciding with that of a-expansion.
Nikos Komodakis, Georgios Tziritas, Nikos Paragios
CVPR1
2007 MRF Optimization via Dual Decomposition: Message-Passing Revisited
abstract
A new message-passing scheme for MRF optimization is proposed in this paper. This scheme inherits better theoretical properties than all other state-of-the-art message passing methods and in practice performs equally well/outperforms them. It is based on the very powerful technique of Dual Decomposition [1] and leads to an elegant and general framework for understanding/designing message-passing algorithms that can provide new insights into existing techniques. Promising experimental results and comparisons with the state of the art demonstrate the extreme theoretical and practical potentials of our approach.
Nikos Komodakis, Nikos Paragios, Georgios Tziritas
ICCV1
2007 Primal/Dual Linear Programming and Statistical Atlases for Cartilage Segmentation
Ben Glocker, Nikos Komodakis, Nikos Paragios, Christian Glaser, Georgios Tziritas, Nassir Navab
MICCAI (2)2
2007 Approximate Labeling via Graph Cuts Based on Linear Programming
abstract
A new framework is presented for both understanding and developing graph-cut-based combinatorial algorithms suitable for the approximate optimization of a very wide class of Markov Random Fields (MRFs) that are frequently encountered in computer vision. The proposed framework utilizes tools from the duality theory of linear programming in order to provide an alternative and more general view of state-of-the-art techniques like the \alpha-expansion algorithm, which is included merely as a special case. Moreover, contrary to \alpha-expansion, the derived algorithms generate solutions with guaranteed optimality properties for a much wider class of problems, for example, even for MRFs with nonmetric potentials. In addition, they are capable of providing per-instance suboptimality bounds in all occasions, including discrete MRFs with an arbitrary potential function. These bounds prove to be very tight in practice (that is, very close to 1), which means that the resulting solutions are almost optimal. Our algorithms' effectiveness is demonstrated by presenting experimental results on a variety of low-level vision tasks, such as stereo matching, image restoration, image completion, and optical flow estimation, as well as on synthetic problems.
Nikos Komodakis, Georgios Tziritas
IEEE Trans. Pattern Anal. Mach. Intell.1
2007 Image Completion Using Efficient Belief Propagation Via Priority Scheduling and Dynamic Pruning
abstract
In this paper, a new exemplar-based framework is presented, which treats image completion, texture synthesis, and image inpainting in a unified manner. In order to be able to avoid the occurrence of visually inconsistent results, we pose all of the above image-editing tasks in the form of a discrete global optimization problem. The objective function of this problem is always well-defined, and corresponds to the energy of a discrete Markov random field (MRF). For efficiently optimizing this MRF, a novel optimization scheme, called priority belief propagation (BP), is then proposed, which carries two very important extensions over the standard BP algorithm: "priority-based message scheduling" and "dynamic label pruning." These two extensions work in cooperation to deal with the intolerable computational cost of BP, which is caused by the huge number of labels associated with our MRF. Moreover, both of our extensions are generic, since they do not rely on the use of domain-specific prior knowledge. They can, therefore, be applied to any MRF, i.e., to a very wide class of problems in image processing and computer vision, thus managing to resolve what is currently considered as one major limitation of the BP algorithm: its inefficiency in handling MRFs with very large discrete state spaces. Experimental results on a wide variety of input images are presented, which demonstrate the effectiveness of our image-completion framework for tasks such as object removal, texture synthesis, text removal, and image inpainting.
Nikos Komodakis, Georgios Tziritas
IEEE Trans. Image Process.1
2006 Image Completion Using Global Optimization
abstract
A new exemplar-based framework unifying image completion, texture synthesis and image inpainting is presented in this work. Contrary to existing greedy techniques, these tasks are posed in the form of a discrete global optimization problem with a well defined objective function. For solving this problem a novel optimization scheme, called Priority- BP, is proposed which carries two very important extensions over standard belief propagation (BP): "prioritybased message scheduling" and "dynamic label pruning". These two extensions work in cooperation to deal with the intolerable computational cost of BP caused by the huge number of existing labels. Moreover, both extensions are generic and can therefore be applied to any MRF energy function as well. The effectiveness of our method is demonstrated on a wide variety of image completion examples.
Nikos Komodakis
CVPR (1)1
2005 A New Framework for Approximate Labeling via Graph Cuts
abstract
A new framework is presented that uses tools from duality theory of linear programming to derive graph-cut based combinatorial algorithms for approximating NP-hard classification problems. The derived algorithms include alpha-expansion graph cut techniques merely as a special case, have guaranteed optimality properties even in cases where alpha-expansion techniques fail to do so and can provide very tight per-instance suboptimality bounds in all occasions
Nikos Komodakis, Georgios Tziritas
ICCV1
2005 3D visual reconstruction of large scale natural sites and their fauna
Nikos Komodakis, Costas Panagiotakis, Georgios Tziritas
Signal Process. Image Commun.1