Tinne Tuytelaars

dblp:79/2382 · DBLP profile ↗
← Back
224ranked-venue papers
13as first author
70since 2021 · last 2026
0000-0003-3307-9723ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 171 · 11 first-author · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 142 · 8 first-author · 38 since 2021Systems, architecture and hardware · 13 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1
YearPublicationVenuePosition
2026 RGS-DR: Deferred Reflections and Residual Shading in 2D Gaussian Splatting
abstract
In this work, we address specular appearance in inverse rendering using 2D Gaussian splatting with deferred shading and argue for a refinement stage to improve specular detail, thereby bridging the gap with reconstruction-only methods. Our pipeline estimates editable material properties and environment illumination while employing a directional residual pass that captures leftover view-dependent effects for further refining novel view synthesis. In contrast to per-Gaussian shading with shortest-axis normals and normal residuals, which tends to result in more noisy geometry and specular appearance, a pixel-deferred surfel formulation with specular residuals yields sharper highlights, cleaner materials, and improved editability. We evaluate our approach on rendering and reconstruction quality on three popular datasets featuring glossy objects, and also demonstrate high-quality relighting and material editing. The source code is available at https://github.com/gkouros/RGS-DR.
Georgios Kouros, Minye Wu, Tinne Tuytelaars
3DV3
2026 OASIS: Online Sample Selection for Continual Instruction Tuning
abstract
In continual instruction tuning (CIT) scenarios, where new instruction tuning data continuously arrive in an online streaming manner, training delays from large-scale data significantly hinder real-time adaptation.Data selection can mitigate this overhead, but existing strategies often rely on pre-trained reference models, which are impractical in CIT setups since future data are unknown.Recent reference model-free online sample selection methods address this, but typically select a fixed number of samples per batch (e.g., top-k), making them vulnerable to distribution shifts where informativeness varies across batches.To address these limitations, we propose OASIS, an adaptive online sample selection approach for CIT that (1) selects informative samples by estimating each sample's informativeness relative to all previously seen data, beyond batch-level constraints, and (2) minimizes informative redundancy of selected samples through iterative selection score updates.Experiments on large foundation models show that OASIS, using only 25% of the data, achieves comparable performance to full-data training and outperforms the state-of-the-art sampling methods.‡
Minhyuk Seo, Tingyu Qu, Tinne Tuytelaars
ACL (1)4
2026 Protecting multimodal large language models against misleading visualizations
abstract
Visualizations play a pivotal role in daily communication in an increasingly data-driven world.Research on multimodal large language models (MLLMs) for automated chart understanding has accelerated massively, with steady improvements on standard benchmarks.However, for MLLMs to be reliable, they must be robust to misleading visualizations, i.e., charts that distort the underlying data, leading readers to draw inaccurate conclusions.Here, we uncover an important vulnerability: MLLM question-answering (QA) accuracy on misleading visualizations drops on average to the level of the random baseline.To address this, we provide the first comparison of six inference-time methods to improve QA performance on misleading visualizations, without compromising accuracy on non-misleading ones.We find that two methods, table-based QA and redrawing the visualization, are effective, with improvements of up to 19.6 percentage points.We make our code and data available.1 What is the proportion of death of coronavirus as a proportion of the number of cases?Around 66%Around 16% Around 6%Around 36%What was the general trend in gun deaths in Florida from 2003 to 2007?then Cannot be inferred then Were there more abortions than cancer screenings in 2011?Yes Cannot be inferred No MisrepresentationThe numerical values are not proportional to the height of the bars. Dual axisCancer screenings and abortions are shown on two different axes.
Jonathan Tonglet, Tinne Tuytelaars, Marie-Francine Moens, Iryna Gurevych
ACL (1)2
2026 Is this chart lying to me? Automating the detection of misleading visualizations
abstract
Misleading visualizations are a potent driver of misinformation on social media and the web. By violating chart design principles, they distort data and lead readers to draw inaccurate conclusions. Prior work has shown that both humans and multimodal large language models (MLLMs) are frequently deceived by such visualizations. Automatically detecting misleading visualizations and identifying the specific design rules they violate could help protect readers and reduce the spread of misinformation. However, the training and evaluation of AI models has been limited by the absence of large, diverse, and openly available datasets. In this work, we introduce Misviz, a benchmark of 2,604 real-world visualizations annotated with 12 types of misleaders. To support model training, we also create Misviz-synth, a synthetic dataset of 57,665 visualizations generated using Matplotlib and based on real-world data tables. We perform a comprehensive evaluation on both datasets using state-of-the-art MLLMs, rule-based systems, and image-axis classifiers. Our results reveal that the task remains highly challenging. We release Misviz, Misviz-synth, and the accompanying code.
Jonathan Tonglet, Jan Zimny, Tinne Tuytelaars, Iryna Gurevych
ACL (1)3
2026 Spec-Gloss Surfels and Normal-Diffuse Priors for Relightable Glossy Objects
abstract
Accurate reconstruction and relighting of glossy objects remains a longstanding challenge, as object shape, material properties, and illumination are inherently difficult to disentangle. Existing neural rendering approaches often rely on simplified BRDF models or parameterizations that couple diffuse and specular components, which restrict faithful material recovery and limit relighting fidelity. We propose a relightable framework that integrates a microfacet BRDF with the specular-glossiness parameterization into 2D Gaussian Splatting with deferred shading. This formulation enables more physically consistent material decomposition, while diffusion-based priors for surface normals and diffuse color guide early-stage optimization and mitigate ambiguity. A coarse-to-fine environment map optimization accelerates convergence, and negative-only environment map clipping preserves high-dynamic-range specular reflections. Extensive experiments on complex, glossy scenes demonstrate that our method achieves high-quality geometry and material reconstruction, delivering substantially more realistic and consistent relighting under novel illumination compared to existing Gaussian splatting methods. The source code is available at https://github.com/gkouros/SpecGloss-GS.
Georgios Kouros, Minye Wu, Tinne Tuytelaars
WACV3
2026 Eff-GRot: Efficient and Generalizable Rotation Estimation with Transformers
abstract
We introduce Eff-GRot, an approach for efficient and generalizable rotation estimation from RGB images. Given a query image and a set of reference images with known orientations, our method directly predicts the object’s rotation in a single forward pass, without requiring object- or category-specific training. At the core of our framework is a transformer that performs a comparison in the latent space, jointly processing rotation-aware representations from multiple references alongside a query. This design enables a favorable balance between accuracy and computational efficiency while remaining simple, scalable, and fully end-to-end. Experimental results show that Eff-GRot offers a promising direction toward more efficient rotation estimation, particularly in latency-sensitive applications. Code is available at: https://github.com/fmathiou/eff-grot.
Fanis Mathioulakis, Gorjan Radevski, Tinne Tuytelaars
WACV3
2026 A Framework for Real-Time Surgical Phase Recognition with Application to Robot-Assisted Partial Nephrectomy
abstract
Surgical practice has increasingly integrated advanced technologies to improve procedural outcomes, efficiency, and safety in modern operating rooms. Within this evolving landscape, Automated Surgical Phase Recognition (SPR) leverages Artificial Intelligence to temporally segment surgical workflows into key events, thereby supporting both real-time decision-making and off-line analysis. Despite the potential of SPR, previous research focused on short and linear surgeries, paying limited attention to the development, assessment, and deployment of real-time systems for complex surgical workflows. This work addresses these gaps by targeting the highly-complex and non linear workflow of Robot-Assisted Partial Nephrectomy (RAPN). We develop a real-time SPR system trained on 143 annotated RAPN surgical videos spanning 15 phases. The system incorporates a trainable canonical calibration error estimator combined with Viterbi decoding for more reliable outcomes. Additionally, we introduce a novel assessment framework designed to simultaneously evaluate offline, real-time, and averaged SPR performance, synthesising historical phase predictions over time. For deployment, we implement the SPR pipeline as an end-to-end application using the NVIDIA Holoscan platform. The system was successfully tested during three live RAPN cases on human patients in a collaborating hospital, achieving an average inference latency of 16.65 ms and an accuracy of 68.2%. Results indicate that Viterbi decoding boosts performance in this complex surgery, while canonical calibration does not significantly increase overall performance but enhances classification reliability. We show the feasibility of deploying a real-time SPR pipeline for RAPN, which holds promise for optimising OR planning. The application is available at https://github.com/nvidia-holoscan/holohub/tree/main/applications/orsi
Marco Mezzina, Tom Vercauteren, Tinne Tuytelaars, Matthew B. Blaschko
WACV3
2025 Charm: The Missing Piece in ViT Fine-Tuning for Image Aesthetic Assessment
abstract
The capacity of Vision transformers (ViTs) to handle variable-sized inputs is often constrained by computational complexity and batch processing limitations. Consequently, ViTs are typically trained on small, fixed-size images obtained through downscaling or cropping. While reducing computational burden, these methods result in significant information loss, negatively affecting tasks like image aesthetic assessment. We introduce Charm, a novel tokenization approach that preserves Composition, High-resolution, Aspect Ratio, and Multi-scale information simultaneously. Charm prioritizes high-resolution details in specific regions while downscaling others, enabling shorter fixed-size input sequences for ViTs while incorporating essential information. Charm is designed to be compatible with pre-trained ViTs and their learned positional embeddings. By providing multiscale input and introducing variety to input tokens, Charm improves ViT performance and generalizability for image aesthetic assessment. We avoid cropping or changing the aspect ratio to further preserve information. Extensive experiments demonstrate significant performance improvements on various image aesthetic and quality assessment datasets (up to 8.1 %) using a lightweight ViT backbone. Code and pre-trained models are available at https://github.com/FBehrad/Charm.
Fatemeh Behrad, Tinne Tuytelaars, Johan Wagemans
CVPR2
2025 BG-Triangle: Bezier Gaussian Triangle for 3D Vectorization and Rendering
abstract
Differentiable rendering enables efficient optimization by allowing gradients to be computed through the rendering process, facilitating 3D reconstruction, inverse rendering and neural scene representation learning. To ensure differentiability, existing solutions approximate or reformulate traditional rendering operations using smooth, probabilistic proxies such as volumes or Gaussian primitives. Consequently, they struggle to preserve sharp edges due to the lack of explicit boundary definitions. We present a novel hybrid representation, Bézier Gaussian Triangle (BG-Triangle), that combines Bézier triangle-based vector graphics primitives with Gaussian-Based probabilistic models, to maintain accurate shape modeling while conducting resolution-independent differentiable rendering. We present a robust and effective discontinuity-aware rendering technique to reduce uncertainties at object boundaries. We also employ an adaptive densification and pruning scheme for efficient training while reliably handling level-of-detail (LoD) variations. Experiments show that BG-Triangle achieves comparable rendering quality as 3DGS [27] but with superior boundary preservation. More importantly, BG-Triangle uses a much smaller number of primitives than its alternatives, showcasing the benefits of vectorized graphics primitives and the potential to bridge the gap between classic and emerging representations.
Minye Wu, Haizhao Dai, Kaixin Yao, Tinne Tuytelaars, Jingyi Yu 0001
CVPR4
2025 Self-Incremental Training for Personalized Voice Command Recognition in a Wireless Audio Sensor Network
abstract
This paper studies self-incremental training in the context of personalized Deep Neural Networks (DNNs) for voice command recognition tailored for resource-constrained sensor nodes. The learning task runs when new unsupervised data becomes available within a Wireless Audio Sensor Network (WASN). After collecting a new multi-sensor dataset of voice commands, we experimentally investigate network-level policies to assign pseudo-labels to the new data. Our baseline analysis shows an accuracy improvement of up to +15% with respect to models pretrained on a large keyword corpus dataset. The multi-sensor labeling strategy closely approximates the performance achieved in a single-sensor scenario providing a clean signal, while we observe +4.7% compared to other sensors with degraded signal quality.
Manuele Rusci, Hugo Van hamme, Tinne Tuytelaars
ICASSP3
2025 Object-Centric Pretraining via Target Encoder Bootstrapping
abstract
Object-centric representation learning has recently been successfully applied to real-world datasets. This success can be attributed to pretrained non-object-centric foundation models, whose features serve as reconstruction targets for slot attention. However, targets must remain frozen throughout the training, which sets an upper bound on the performance object-centric models can attain. Attempts to update the target encoder by bootstrapping result in large performance drops, which can be attributed to its lack of object-centric inductive biases, causing the object-centric model’s encoder to drift away from representations useful as reconstruction targets. To address these limitations, we propose **O**bject-**CE**ntric Pretraining by Target Encoder **BO**otstrapping, a self-distillation setup for training object-centric models from scratch, on real-world data, for the first time ever. In OCEBO, the target encoder is updated as an exponential moving average of the object-centric model, thus explicitly being enriched with object-centric inductive biases introduced by slot attention while removing the upper bound on performance present in other models. We mitigate the slot collapse caused by random initialization of the target encoder by introducing a novel cross-view patch filtering approach that limits the supervision to sufficiently informative patches. When pretrained on 241k images from COCO, OCEBO achieves unsupervised object discovery performance comparable to that of object-centric models with frozen non-object-centric target encoders pretrained on hundreds of millions of images. The code and pretrained models are publicly available at https://github.com/djukicn/ocebo.
Nikola Dukic, Tim Lebailly, Tinne Tuytelaars
ICLR3
2025 A Simple Framework for Open-Vocabulary Zero-Shot Segmentation
abstract
Zero-shot classification capabilities naturally arise in models trained within a vision-language contrastive framework. Despite their classification prowess, these models struggle in dense tasks like zero-shot open-vocabulary segmentation. This deficiency is often attributed to the absence of localization cues in captions and the intertwined nature of the learning process, which encompasses both image/text representation learning and cross-modality alignment. To tackle these issues, we propose SimZSS, a $\textbf{Sim}$ple framework for open-vocabulary $\textbf{Z}$ero-$\textbf{S}$hot $\textbf{S}$egmentation. The method is founded on two key principles: i) leveraging frozen vision-only models that exhibit spatial awareness while exclusively aligning the text encoder and ii) exploiting the discrete nature of text and linguistic knowledge to pinpoint local concepts within captions. By capitalizing on the quality of the visual representations, our method requires only image-caption pair datasets and adapts to both small curated and large-scale noisy datasets. When trained on COCO Captions across 8 GPUs, SimZSS achieves state-of-the-art results on 7 out of 8 benchmark datasets in less than 15 minutes. Our code and pretrained models are publicly available at https://github.com/tileb1/simzss.
Thomas Stegmüller, Tim Lebailly, Nikola Dukic, Behzad Bozorgtabar, Tinne Tuytelaars, Jean-Philippe Thiran
ICLR5
2025 Predicting the Susceptibility of Examples to Catastrophic Forgetting
abstract
Catastrophic forgetting -- the tendency of neural networks to forget previously learned data when learning new information -- remains a central challenge in continual learning. In this work, we adopt a behavioral approach, observing a connection between learning speed and forgetting: examples learned more quickly are less prone to forgetting. Focusing on replay-based continual learning, we show that the composition of the replay buffer -- specifically, whether it contains quickly or slowly learned examples -- has a significant effect on forgetting. Motivated by this insight, we introduce Speed-Based Sampling (SBS), a simple yet general strategy that selects replay examples based on their learning speed. SBS integrates easily into existing buffer-based methods and improves performance across a wide range of competitive continual learning benchmarks, advancing state-of-the-art results. Our findings underscore the value of accounting for the forgetting dynamics when designing continual learning algorithms.
Guy Hacohen, Tinne Tuytelaars
ICML2
2025 Collapse-Proof Non-Contrastive Self-Supervised Learning
abstract
We present a principled and simplified design of the projector and loss function for non-contrastive self-supervised learning based on hyperdimensional computing. We theoretically demonstrate that this design introduces an inductive bias that encourages representations to be simultaneously decorrelated and clustered, without explicitly enforcing these properties. This bias provably enhances generalization and suffices to avoid known training failure modes, such as representation, dimensional, cluster, and intracluster collapses. We validate our theoretical findings on image datasets, including SVHN, CIFAR-10, CIFAR-100, and ImageNet-100. Our approach effectively combines the strengths of feature decorrelation and cluster-based self-supervised learning methods, overcoming training failure modes while achieving strong generalization in clustering and linear classification tasks.
Emanuele Sansone, Tim Lebailly, Tinne Tuytelaars
ICML3
2025 DAVE: Diagnostic benchmark for Audio Visual Evaluation
abstract
Audio-visual understanding is a rapidly evolving field that seeks to integrate and interpret information from both auditory and visual modalities. Despite recent advances in multi-modal learning, existing benchmarks often suffer from strong visual bias -- when answers can be inferred from visual data alone -- and provide only aggregate scores that conflate multiple sources of error. This makes it difficult to determine whether models struggle with visual understanding, audio interpretation, or audio-visual alignment. In this work, we introduce DAVE: Diagnostic Audio Visual Evaluation, a novel benchmark dataset designed to systematically evaluate audio-visual models across controlled settings. DAVE alleviates existing limitations by (i) ensuring both modalities are necessary to answer correctly and (ii) decoupling evaluation into atomic subcategories. Our detailed analysis of state-of-the-art models reveals specific failure modes and provides targeted insights for improvement. By offering this standardized diagnostic framework, we aim to facilitate more robust development of audio-visual models.Dataset: https://huggingface.co/datasets/gorjanradevski/daveCode: https://github.com/gorjanradevski/dave
Gorjan Radevski, Teodora Popordanoska, Matthew B. Blaschko, Tinne Tuytelaars
NeurIPS4
2025 DM-Align: Leveraging the power of natural language instructions to make changes to images
abstract
sponsorship: This project was funded by the European Research Council (ERC) Advanced Grant CALCULUS (grant agreement No. 788506) . (European Research Council (ERC)|788506, European Research Council (ERC)|788506)
Maria Mihaela Trusca, Tinne Tuytelaars, Marie-Francine Moens
Comput. Vis. Image Underst.2
2025 Instruction-guided path planning with 3D semantic maps for vision-language navigation
Mingxiao Li 0002, Minye Wu, Marie-Francine Moens, Tinne Tuytelaars
Neurocomputing5
2025 Self-Learning for Personalized Keyword Spotting on Ultralow-Power Audio Sensors
abstract
This article proposes a self-learning method to incrementally train (fine-tune) a personalized keyword spotting (KWS) model after the deployment on ultralow power smart audio sensors. We address the fundamental problem of the absence of labeled training data by assigning pseudo-labels to the new recorded audio frames based on a similarity score with respect to few user recordings. By experimenting with multiple KWS models with a number of parameters up to 0.5 M on two public datasets, we show an accuracy improvement of up to +19.2% and +16.0% versus the initial models pretrained on a large set of generic keywords. The labeling task is demonstrated on a sensor system composed of a low-power microphone and an energy-efficient microcontroller (MCU). By efficiently exploiting the heterogeneous processing engines of the MCU, the always-on labeling task runs in real-time with an average power cost of up to 8.2 mW. On the same platform, we estimate an energy cost for on-device training$10\times $lower than the labeling energy if sampling a new utterance every 6.1 or 18.8 s with a DS-CNN-S or a DS-CNN-M model. Our empirical result paves the way to self-adaptive personalized KWS sensors at the extreme edge.
Manuele Rusci, Francesco Paci, Marco Fariselli, Eric Flamand, Tinne Tuytelaars
IEEE Internet Things J.5
2024 NeVRF: Neural Video-Based Radiance Fields for Long-Duration Sequences
abstract
Adopting Neural Radiance Fields (NeRF) to long-duration dynamic sequences has been challenging. Existing methods struggle to balance between quality and storage size and encounter difficulties with complex scene changes such as topological changes and large motions. To tackle these issues, we propose a novel neural video-based radiance fields (NeVRF) representation. NeVRF marries neural radiance field with image-based rendering to support photo-realistic novel view synthesis on long-duration dynamic inward-looking scenes. We introduce a novel multi-view radiance blending approach to predict radiance directly from multi-view videos. By incorporating continual learning techniques, NeVRF can efficiently reconstruct frames from sequential data without revisiting previous frames, enabling long-duration free-viewpoint video. Furthermore, with a tailored compression approach, NeVRF can compactly represent dynamic scenes, making dynamic radiance fields more practical in real-world scenarios. Our extensive experiments demonstrate the effectiveness of NeVRF in enabling long-duration sequence rendering, sequential data reconstruction, and compact data storage.
Minye Wu, Tinne Tuytelaars
3DV2
2024 TeTriRF: Temporal Tri-Plane Radiance Fields for Efficient Free-Viewpoint Video
abstract
Neural Radiance Fields (NeRF) revolutionize the realm of visual media by providing photorealistic Free-Viewpoint Video (FVV) experiences, offering viewers unparalleled immersion and interactivity. However, the technology's significant storage requirements and the computational complexity involved in generation and rendering currently limit its broader application. To close this gap, this paper presents Temporal Tri-Plane Radiance Fields (TeTriRF), a novel technology that significantly reduces the storage size for Free-Viewpoint Video (FVV) while maintaining low-cost generation and rendering. TeTriRF introduces a hybrid representation with tri-planes and voxel grids to support scaling up to long-duration sequences and scenes with complex motions or rapid changes. We propose a group training scheme tailored to achieving high training efficiency and yielding temporally consistent, low-entropy scene representations on feature domain. Leveraging these properties of the representations, we introduce a compression pipeline with off-the-shelf video codecs, achieving an order of magnitude less storage size compared to the state-of-the-art. Our experiments demonstrate that TeTriRF can achieve competitive quality with a higher compression rate.
Minye Wu, Georgios Kouros, Tinne Tuytelaars
CVPR4
2024 Animate Your Motion: Turning Still Images into Dynamic Videos
Mingxiao Li 0002, Marie-Francine Moens, Tinne Tuytelaars
ECCV (66)4
2024 Introducing Routing Functions to Vision-Language Parameter-Efficient Fine-Tuning with Low-Rank Bottlenecks
Tingyu Qu, Tinne Tuytelaars, Marie-Francine Moens
ECCV (88)2
2024 Prediction Error-based Classification for Class-Incremental Learning
abstract
Class-incremental learning (CIL) is a particularly challenging variant of continual learning, where the goal is to learn to discriminate between all classes presented in an incremental fashion. Existing approaches often suffer from excessive forgetting and imbalance of the scores assigned to classes that have not been seen together during training. In this study, we introduce a novel approach, Prediction Error-based Classification (PEC), which differs from traditional discriminative and generative classification paradigms. PEC computes a class score by measuring the prediction error of a model trained to replicate the outputs of a frozen random neural network on data from that class. The method can be interpreted as approximating a classification rule based on Gaussian Process posterior variance. PEC offers several practical advantages, including sample efficiency, ease of tuning, and effectiveness even when data are presented one class at a time. Our empirical results show that PEC performs strongly in single-pass-through-data CIL, outperforming other rehearsal-free baselines in all cases and rehearsal-based methods with moderate replay buffer size in most cases across multiple benchmarks.
Michal Zajac 0005, Tinne Tuytelaars, Gido M. van de Ven
ICLR2
2024 CrIBo: Self-Supervised Learning via Cross-Image Object-Level Bootstrapping
abstract
Leveraging nearest neighbor retrieval for self-supervised representation learning has proven beneficial with object-centric images. However, this approach faces limitations when applied to scene-centric datasets, where multiple objects within an image are only implicitly captured in the global representation. Such global bootstrapping can lead to undesirable entanglement of object representations. Furthermore, even object-centric datasets stand to benefit from a finer-grained bootstrapping approach. In response to these challenges, we introduce a novel $\textbf{Cr}$oss-$\textbf{I}$mage Object-Level $\textbf{Bo}$otstrapping method tailored to enhance dense visual representation learning. By employing object-level nearest neighbor bootstrapping throughout the training, CrIBo emerges as a notably strong and adequate candidate for in-context learning, leveraging nearest neighbor retrieval at test time. CrIBo shows state-of-the-art performance on the latter task while being highly competitive in more standard downstream segmentation tasks. Our code and pretrained models are publicly available at https://github.com/tileb1/CrIBo.
Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran, Tinne Tuytelaars
ICLR5
2024 Driving from Vision through Differentiable Optimal Control
abstract
This paper proposes DriViDOC: a framework for Driving from Vision through Differentiable Optimal Control, and its application to learn autonomous driving controllers from human demonstrations. DriViDOC combines the automatic inference of relevant features from camera frames with the properties of nonlinear model predictive control (NMPC), such as constraint satisfaction. Our approach leverages the differentiability of parametric NMPC, allowing for end-to-end learning of the driving model from images to control. The model is trained on an offline dataset comprising various human demonstrations collected on a motion-base driving simulator. During online testing, the model demonstrates successful imitation of different driving styles, and the interpreted NMPC parameters provide insights into the achievement of specific driving behaviors. Our experimental results show that DriViDOC outperforms other methods involving NMPC and neural networks, exhibiting an average improvement of 20% in imitation scores.
Flavia Sofia Acerbo, Jan Swevers, Tinne Tuytelaars, Tong Duy Son
IROS3
2024 Visually-Aware Context Modeling for News Image Captioning
abstract
Tingyu Qu, Tinne Tuytelaars, Marie-Francine Moens. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Tingyu Qu, Tinne Tuytelaars, Marie-Francine Moens
NAACL-HLT2
2024 LaSCal: Label-Shift Calibration without target labels
abstract
When machine learning systems face dataset shift, model calibration plays a pivotal role in ensuring their reliability. Calibration error (CE) provides insights into the alignment between the predicted confidence scores and the classifier accuracy. While prior works have delved into the implications of dataset shift on calibration, existing CE estimators either (i) assume access to labeled data from the target domain, often unavailable in practice, or (ii) are derived under a covariate shift assumption. In this work we propose a novel, label-free, consistent CE estimator under label shift. Label shift is characterized by changes in the marginal label distribution p(Y), with a constant conditional p(X|Y) distribution between the source and target. We introduce a novel calibration method, called LaSCal, which uses the estimator in conjunction with a post-hoc calibration strategy, to perform unsupervised calibration on the target distribution. Our thorough empirical analysis demonstrates the effectiveness and reliability of the proposed approach across different modalities, model architectures and label shift intensities.
Teodora Popordanoska, Gorjan Radevski, Tinne Tuytelaars, Matthew B. Blaschko
NeurIPS3
2024 Contrastive Learning for Multi-Object Tracking with Transformers
abstract
The DEtection TRansformer (DETR) opened new possibilities for object detection by modeling it as a translation task: converting image features into object-level representations. Previous works typically add expensive modules to DETR to perform Multi-Object Tracking (MOT), resulting in more complicated architectures. We instead show how DETR can be turned into a MOT model by employing an instance-level contrastive loss, a revised sampling strategy and a lightweight assignment method. Our training scheme learns object appearances while preserving detection capabilities and with little overhead. Its performance surpasses the previous state-of-the-art by +2.6 mMOTA on the challenging BDD100K dataset and is comparable to existing transformer-based methods on the MOT17 dataset.
Pierre-François De Plaen, Nicola Marinello, Marc Proesmans, Tinne Tuytelaars, Luc Van Gool
WACV4
2024 Exploiting CLIP for Zero-shot HOI Detection Requires Knowledge Distillation at Multiple Levels
abstract
In this paper, we investigate the task of zero-shot human-object interaction (HOI) detection, a novel paradigm for identifying HOIs without the need for task-specific annotations. To address this challenging task, we employ CLIP, a large-scale pre-trained vision-language model (VLM), for knowledge distillation on multiple levels. Specifically, we design a multi-branch neural network that leverages CLIP for learning HOI representations at various levels, including global images, local union regions encompassing human-object pairs, and individual instances of humans or objects. To train our model, CLIP is utilized to generate HOI scores for both global images and local union regions that serve as supervision signals. The extensive experiments demonstrate the effectiveness of our novel multi-level CLIP knowledge integration strategy. Notably, the model achieves strong performance, which is even comparable with some fully-supervised and weakly-supervised methods on the public HICO-DET benchmark. Code is available at https://github.com/bobwan1995/Zeroshot-HOI-with-CLIP.
Tinne Tuytelaars
WACV2
2024 Continual pre-training mitigates forgetting in language and vision
abstract
Pre-trained models are commonly used in Continual Learning to initialize the model before training on the stream of non-stationary data. However, pre-training is rarely applied during Continual Learning. We investigate the characteristics of the Continual Pre-Training scenario, where a model is continually pre-trained on a stream of incoming data and only later fine-tuned to different downstream tasks. We introduce an evaluation protocol for Continual Pre-Training which monitors forgetting against a Forgetting Control dataset not present in the continual stream. We disentangle the impact on forgetting of 3 main factors: the input modality (NLP, Vision), the architecture type (Transformer, ResNet) and the pre-training protocol (supervised, self-supervised). Moreover, we propose a Sample-Efficient Pre-training method (SEP) that speeds up the pre-training phase. We show that the pre-training protocol is the most important factor accounting for forgetting. Surprisingly, we discovered that self-supervised continual pre-training in both NLP and Vision is sufficient to mitigate forgetting without the use of any Continual Learning strategy. Other factors, like model depth, input modality and architecture type are not as crucial. • Continual Pre-Training incrementally acquires knowledge from unstructured data streams. • Self-Supervised Continual Pre-Training effectively mitigates forgetting. • The representation drift is reduced by Self-Supervised Continual Pre-Training. • Performance on domain-specific tasks can be improved with a limited amount of data.
Andrea Cossu, Antonio Carta, Lucia C. Passaro, Vincenzo Lomonaco, Tinne Tuytelaars, Davide Bacciu
Neural Networks5
2024 Editorial: Learning With Fewer Labels in Computer Vision
abstract
Undoubtedly, Deep Neural Networks (DNNs), from AlexNet to ResNet to Transformer, have sparked revolutionary advancements in diverse computer vision tasks. The scale of DNNs has grown exponentially due to the rapid development of computational resources. Despite the tremendous success, DNNs typically depend on massive amounts of training data (especially the recent various foundation models) to achieve high performance and are brittle in that their performance can degrade severely with small changes in their operating environment. Generally, collecting massive-scale training datasets is costly or even infeasible, as for certain fields, only very limited or no examples at all can be gathered. Nevertheless, collecting, labeling, and vetting massive amounts of practical training data is certainly difficult and expensive, as it requires the painstaking efforts of experienced human annotators or experts, and in many cases, prohibitively costly or impossible due to some reason, such as privacy, safety or ethic issues.
Li Liu 0002, Timothy M. Hospedales, Yann LeCun, Mingsheng Long, Jiebo Luo 0001, Wanli Ouyang, Matti Pietikäinen, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.8
2024 SNIPPET: A Framework for Subjective Evaluation of Visual Explanations Applied to DeepFake Detection
abstract
Explainable Artificial Intelligence (XAI) attempts to help humans understand machine learning decisions better and has been identified as a critical component toward increasing the trustworthiness of complex black-box systems, such as deep neural networks. In this article, we propose a generic and comprehensive framework named SNIPPET and create a user interface for the subjective evaluation of visual explanations, focusing on finding human-friendly explanations. SNIPPET considers human-centered evaluation tasks and incorporates the collection of human annotations. These annotations can serve as valuable feedback to validate the qualitative results obtained from the subjective assessment tasks. Moreover, we consider different user background categories during the evaluation process to ensure diverse perspectives and comprehensive evaluation. We demonstrate SNIPPET on a DeepFake face dataset. Distinguishing real from fake faces is a non-trivial task even for humans that depends on rather subtle features, making it a challenging use case. Using SNIPPET, we evaluate four popular XAI methods which provide visual explanations: Gradient-weighted Class Activation Mapping, Layer-wise Relevance Propagation, attention rollout, and Transformer Attribution. Based on our experimental results, we observe preference variations among different user categories. We find that most people are more favorable to the explanations of rollout. Moreover, when it comes to XAI-assisted understanding, those who have no or lack relevant background knowledge often consider that visual explanations are insufficient to help them understand. We open-source our framework for continued data collection and annotation at https://github.com/XAI-SubjEvaluation/SNIPPET .
Boris Joukovsky, José Oramas M., Tinne Tuytelaars, Nikos Deligiannis
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Layout-Aware Dreamer for Embodied Visual Referring Expression Grounding
abstract
In this work, we study the problem of Embodied Referring Expression Grounding, where an agent needs to navigate in a previously unseen environment and localize a remote object described by a concise high-level natural language instruction. When facing such a situation, a human tends to imagine what the destination may look like and to explore the environment based on prior knowledge of the environmental layout, such as the fact that a bathroom is more likely to be found near a bedroom than a kitchen. We have designed an autonomous agent called Layout-aware Dreamer (LAD), including two novel modules, that is, the Layout Learner and the Goal Dreamer to mimic this cognitive decision process. The Layout Learner learns to infer the room category distribution of neighboring unexplored areas along the path for coarse layout estimation, which effectively introduces layout common sense of room-to-room transitions to our agent. To learn an effective exploration of the environment, the Goal Dreamer imagines the destination beforehand. Our agent achieves new state-of-the-art performance on the public leaderboard of REVERIE dataset in challenging unseen test environments with improvement on navigation success rate (SR) by 4.02% and remote grounding success (RGS) by 3.43% comparing to previous previous state of the art. The code is released at https://github.com/zehao-wang/LAD.
Mingxiao Li 0002, Tinne Tuytelaars, Marie-Francine Moens
AAAI3
2023 Unbalanced Optimal Transport: A Unified Framework for Object Detection
abstract
During training, supervised object detection tries to correctly match the predicted bounding boxes and associated classification scores to the ground truth. This is essential to determine which predictions are to be pushed towards which solutions, or to be discarded. Popular matching strategies include matching to the closest ground truth box (mostly used in combination with anchors), or matching via the Hungarian algorithm (mostly used in anchor free methods). Each of these strategies comes with its own properties, underlying losses, and heuristics. We show how Unbalanced Optimal Transport unifies these different approaches and opens a whole continuum of methods in between. This allows for a finer selection of the desired properties. Experimentally, we show that training an object detection model with Unbalanced Optimal Transport is able to reach the state-of-the-art both in terms of Average Precision and Average Recall as well as to provide a faster initial convergence. The approach is well suited for GPU implementation, which proves to be an advantage for large-scale models.
Henri De Plaen, Pierre-François De Plaen, Johan A. K. Suykens, Marc Proesmans, Tinne Tuytelaars, Luc Van Gool
CVPR5
2023 CrOC: Cross-View Online Clustering for Dense Visual Representation Learning
abstract
Learning dense visual representations without labels is an arduous task and more so from scene-centric data. We propose to tackle this challenging problem by proposing a Cross-view consistency objective with an Online Clustering mechanism (CrOC) to discover and segment the semantics of the views. In the absence of hand-crafted priors, the resulting method is more generalizable and does not require a cumbersome pre-processing step. More importantly, the clustering algorithm conjointly operates on the features of both views, thereby elegantly bypassing the issue of content not represented in both views and the ambiguous matching of objects from one crop to the other. We demonstrate excellent performance on linear and unsupervised segmentation transfer tasks on various datasets and similarly for video object segmentation. Our code and pre-trained models are publicly available at https://github.com/stegmuel/CrOC.
Thomas Stegmüller, Tim Lebailly, Behzad Bozorgtabar, Tinne Tuytelaars, Jean-Philippe Thiran
CVPR4
2023 Neural Residual Radiance Fields for Streamably Free-Viewpoint Videos
abstract
The success of the Neural Radiance Fields (NeRFs) for modeling and free-view rendering static objects has in-spired numerous attempts on dynamic scenes. Current techniques that utilize neural rendering for facilitating free-view videos (FVVs) are restricted to either offline rendering or are capable of processing only brief sequences with minimal motion. In this paper, we present a novel technique, Residual Radiance Field or ReRF, as a highly com-pact neural representation to achieve real-time FVV ren-dering on long-duration dynamic scenes. ReRF explicitly models the residual information between adjacent times-tamps in the spatial-temporal feature space, with a global coordinate-based tiny MLP as the feature decoder. Specif-ically, ReRF employs a compact motion grid along with a residual feature grid to exploit inter-frame feature similar-ities. We show such a strategy can handle large motions without sacrificing quality. We further present a sequential training scheme to maintain the smoothness and the spar-sity of the motion/residual grids. Based on ReRF, we design a special FVV codec that achieves three orders of magni-tudes compression rate and provides a companion ReRF player to support online streaming of long-duration FVVs of dynamic scenes. Extensive experiments demonstrate the effectiveness of ReRF for compactly representing dynamic radiance fields, enabling an unprecedented free-viewpoint viewing experience in speed and quality.
Qiang Hu 0003, Qihan He, Jingyi Yu 0001, Tinne Tuytelaars, Lan Xu 0003, Minye Wu
CVPR6
2023 Adaptive Similarity Bootstrapping for Self-Distillation based Representation Learning
abstract
Most self-supervised methods for representation learning leverage a cross-view consistency objective i.e. they maximize the representation similarity of a given image’s augmented views. Recent work NNCLR goes beyond the cross-view paradigm and uses positive pairs from different images obtained via nearest neighbor bootstrapping in a contrastive setting. We empirically show that as opposed to the contrastive learning setting which relies on negative samples, incorporating nearest neighbor bootstrapping in a self-distillation scheme can lead to a performance drop or even collapse. We scrutinize the reason for this unexpected behavior and provide a solution. We propose to adaptively bootstrap neighbors based on the estimated quality of the latent space. We report consistent improvements compared to the naive bootstrapping approach and the original baselines. Our approach leads to performance improvements for various self-distillation method/backbone combinations and standard downstream tasks. Our code is publicly available at https://github.com/tileb1/AdaSim.
Tim Lebailly, Thomas Stegmüller, Behzad Bozorgtabar, Jean-Philippe Thiran, Tinne Tuytelaars
ICCV5
2023 Multimodal Distillation for Egocentric Action Recognition
abstract
The focal point of egocentric video understanding is modelling hand-object interactions. Standard models, e.g. CNNs or Vision Transformers, which receive RGB frames as input perform well, however, their performance improves further by employing additional input modalities (e.g. object detections, optical flow, audio, etc.) which provide cues complementary to the RGB modality. The added complexity of the modality-specific modules, on the other hand, makes these models impractical for deployment. The goal of this work is to retain the performance of such a multi-modal approach, while using only the RGB frames as input at inference time. We demonstrate that for egocentric action recognition on the Epic-Kitchens and the Something-Something datasets, students which are taught by multi-modal teachers tend to be more accurate and better calibrated than architecturally equivalent models trained on ground truth labels in a unimodal or multimodal fashion. We further adopt a principled multimodal knowledge distillation framework, allowing us to deal with issues which occur when applying multimodal knowledge distillation in a naïve manner. Lastly, we demonstrate the achieved reduction in computational complexity, and show that our approach maintains higher performance with the reduction of the number of input views. We release our code at: https://github.com/gorjanradevski/multimodal-distillation
Gorjan Radevski, Dusan Grujicic, Matthew B. Blaschko, Marie-Francine Moens, Tinne Tuytelaars
ICCV5
2023 Continual evaluation for lifelong learning: Identifying the stability gap
Matthias De Lange, Gido M. van de Ven, Tinne Tuytelaars
ICLR3
2023 Weakly-supervised HOI Detection via Prior-guided Bi-level Representation Learning
Yongfei Liu, Desen Zhou, Tinne Tuytelaars, Xuming He 0001
ICLR4
2023 Continual Learning with Pretrained Backbones by Tuning in the Input Space
abstract
The intrinsic difficulty in adapting deep learning models to non-stationary environments limits the applicability of neural networks to real-world tasks. This issue is critical in practical supervised learning settings, such as the ones in which a pre-trained model computes projections toward a latent space where different task predictors are sequentially learned over time. As a matter of fact, incrementally fine-tuning the whole model to better adapt to new tasks usually results in catastrophic forgetting, with decreasing performance over the past experiences and losing valuable knowledge from the pretraining stage. In this paper, we propose a novel strategy to make the fine-tuning procedure more effective, by avoiding to update the pre-trained part of the network and learning not only the usual classification head, but also a set of newly-introduced learnable parameters that are responsible for transforming the input data. This process allows the network to effectively leverage the pre-training knowledge and find a good trade-off between plasticity and stability with modest computational efforts, thus especially suitable for on-the-edge settings. Our experiments on four image classification problems in a continual learning setting confirm the quality of the proposed approach when compared to several fine-tuning procedures and to popular continual learning methods.
Simone Marullo, Matteo Tiezzi, Marco Gori, Stefano Melacci, Tinne Tuytelaars
IJCNN5
2023 Few-Shot Open-Set Learning for On-Device Customization of KeyWord Spotting Systems
abstract
sponsorship: This work is partly supported by the European Horizon Europe program under grant agreement 101067475. (European Horizon Europe program|101067475, Marie Curie Actions (MSCA)|101067475)
Manuele Rusci, Tinne Tuytelaars
INTERSPEECH2
2023 Revisiting Evaluation Metrics for Semantic Segmentation: Optimization and Evaluation of Fine-grained Intersection over Union
abstract
Semantic segmentation datasets often exhibit two types of imbalance: \textit{class imbalance}, where some classes appear more frequently than others and \textit{size imbalance}, where some objects occupy more pixels than others. This causes traditional evaluation metrics to be biased towards \textit{majority classes} (e.g. overall pixel-wise accuracy) and \textit{large objects} (e.g. mean pixel-wise accuracy and per-dataset mean intersection over union). To address these shortcomings, we propose the use of fine-grained mIoUs along with corresponding worst-case metrics, thereby offering a more holistic evaluation of segmentation techniques. These fine-grained metrics offer less bias towards large objects, richer statistical information, and valuable insights into model and dataset auditing. Furthermore, we undertake an extensive benchmark study, where we train and evaluate 15 modern neural networks with the proposed metrics on 12 diverse natural and aerial segmentation datasets. Our benchmark study highlights the necessity of not basing evaluations on a single metric and confirms that fine-grained mIoUs reduce the bias towards large objects. Moreover, we identify the crucial role played by architecture designs and loss functions, which lead to best practices in optimizing fine-grained metrics. The code is available at \href{https://github.com/zifuwanggg/JDTLosses}{https://github.com/zifuwanggg/JDTLosses}.
Zifu Wang, Maxim Berman, Amal Rannen Triki, Philip Torr 0001, Devis Tuia, Tinne Tuytelaars, Luc Van Gool, Jiaqian Yu, Matthew B. Blaschko
NeurIPS6
2023 Barlow constrained optimization for Visual Question Answering
abstract
Visual question answering is a vision-and-language multimodal task, that aims at predicting answers given samples from the question and image modalities. Most recent methods focus on learning a good joint embedding space of images and questions, either by improving the interaction between these two modalities, or by making it a more discriminant space. However, how informative this joint space is, has not been well explored. In this paper, we propose a novel regularization for VQA models, Constrained Optimization using Barlow’s theory (COB), that improves the information content of the joint space by minimizing the redundancy. It reduces the correlation between the learned feature components and thereby disentangles semantic concepts. Our model also aligns the joint space with the answer embedding space, where we consider the answer and image+question as two different ‘views’ of what in essence is the same semantic information. We propose a constrained optimization policy to balance the categorical and redundancy minimization forces. When built on the state-of-the-art GGE model, the resulting model improves VQA accuracy by 1.4% and 4% on the VQA-CP v2 and VQA v2 datasets respectively. The model also exhibits better interpretability. Code is made available: https://github.com/abskjha/Barlow-constrained-VQA
Abhishek Jha 0001, Badri Narayana Patro, Luc Van Gool, Tinne Tuytelaars
WACV4
2023 SimGlim: Simplifying glimpse based active visual reconstruction
abstract
In active visual exploration, an agent with a limited field of view needs to sample the most informative local observations of an environment in order to model the global context. Current works train this selection strategy by defining a complex architecture built upon features learned through convolutional encoders. In this paper, we first discuss why vision transformers are better suited than CNNs for such an agent. Next, we propose a simple transformer-based active visual sampling model, called "SimGlim", which utilises transformer’s inherent self-attention architecture to sequentially predict the best next location based on the current observable environment. We show the efficacy of our proposed method on the task of image reconstruction in the partial observable setting and compare our model against existing state-of-the-art active visual reconstruction methods. Finally, we provide ablations for the parameters of our design choice to understand their importance in the overall architecture.
Abhishek Jha 0001, Soroush Seifi, Tinne Tuytelaars
WACV3
2023 Global-Local Self-Distillation for Visual Representation Learning
abstract
The downstream accuracy of self-supervised methods is tightly linked to the proxy task solved during training and the quality of the gradients extracted from it. Richer and more meaningful gradients updates are key to allow self-supervised methods to learn better and in a more efficient manner. In a typical self-distillation framework, the representation of two augmented images are enforced to be coherent at the global level. Nonetheless, incorporating local cues in the proxy task can be beneficial and improve the model accuracy on downstream tasks. This leads to a dual objective in which, on the one hand, coherence between global-representations is enforced and on the other, coherence between local-representations is enforced. Unfortunately, an exact correspondence mapping between two sets of local-representations does not exist making the task of matching local-representations from one augmentation to another non-trivial. We propose to leverage the spatial information in the input images to obtain geometric matchings and compare this geometric approach against previous methods based on similarity matchings. Our study shows that not only 1) geometric matchings perform better than similarity based matchings in low-data regimes but also 2) that similarity based matchings are highly hurtful in low-data regimes compared to the vanilla baseline without local self-distillation. The code is available at https://github.com/tileb1/global-local-self-distillation.
Tim Lebailly, Tinne Tuytelaars
WACV2
2023 Weakly Supervised Face Naming with Symmetry-Enhanced Contrastive Loss
abstract
We revisit the weakly supervised cross-modal face-name alignment task; that is, given an image and a caption, we label the faces in the image with the names occurring in the caption. Whereas past approaches have learned the latent alignment between names and faces by uncertainty reasoning over a set of images and their respective captions, in this paper, we rely on appropriate loss functions to learn the alignments in a neural network setting and propose SECLA and SECLA-B.SECLA is a Symmetry-Enhanced Contrastive Learning-based Alignment model that can effectively maximize the similarity scores between corresponding faces and names in a weakly supervised fashion. A variation of the model, SECLA-B, learns to align names and faces as humans do, that is, learning from easy to hard cases to further increase the performance of SECLA. More specifically, SECLA-B applies a two-stage learning framework: (1) Training the model on an easy subset with a few names and faces in each image-caption pair. (2) Leveraging the known pairs of names and faces from the easy cases using a bootstrapping strategy with additional loss to prevent forgetting and learning new alignments at the same time. We achieve state-of-the-art results for both the augmented Labeled Faces in the Wild dataset and the Celebrity Together dataset. In addition, we believe that our methods can be adapted to other multimodal news understanding tasks.
Tingyu Qu, Tinne Tuytelaars, Marie-Francine Moens
WACV2
2023 Spatial Consistency Loss for Training Multi-Label Classifiers from Single-Label Annotations
abstract
Multi-label image classification is more applicable "in the wild" than single-label classification, as natural images usually contain multiple objects. However, exhaustively annotating images with every object of interest is costly and time-consuming. We train multi-label classifiers from datasets where each image is annotated with a single positive label only. As the presence of all other classes is unknown, we propose an Expected Negative loss that builds a set of expected negative labels in addition to the annotated positives. This set is determined based on prediction consistency, by averaging predictions over consecutive training epochs to build robust targets. Moreover, the ‘crop’ data augmentation leads to additional label noise by cropping out the single annotated object. Our novel spatial consistency loss improves supervision and ensures consistency of the spatial feature maps by maintaining per-class running-average heatmaps for each training image. We use MS-COCO, Pascal VOC, NUS-WIDE and CUB-Birds datasets to demonstrate the gains of the Expected Negative loss in combination with consistency and spatial consistency losses. We also demonstrate improved multi-label classification mAP on ImageNet-1K using the ReaL multi-label validation set.
Thomas Verelst, Paul K. Rubenstein, Marcin Eichner, Tinne Tuytelaars, Maxim Berman
WACV4
2023 CLAD: A realistic Continual Learning benchmark for Autonomous Driving
Eli Verwimp, Sarah Parisot, Lanqing Hong, Steven McDonagh 0001, Eduardo Pérez-Pellitero, Matthias De Lange, Tinne Tuytelaars
Neural Networks8
2023 SegBlocks: Block-Based Dynamic Resolution Networks for Real-Time Segmentation
abstract
SegBlocks reduces the computational cost of existing neural networks, by dynamically adjusting the processing resolution of image regions based on their complexity. Our method splits an image into blocks and downsamples blocks of low complexity, reducing the number of operations and memory consumption. A lightweight policy network, selecting the complex regions, is trained using reinforcement learning. In addition, we introduce several modules implemented in CUDA to process images in blocks. Most important, our novel BlockPad module prevents the feature discontinuities at block borders of which existing methods suffer, while keeping memory consumption under control. Our experiments on Cityscapes, Camvid and Mapillary Vistas datasets for semantic segmentation show that dynamically processing images offers a better accuracy versus complexity trade-off compared to static baselines of similar complexity. For instance, our method reduces the number of floating-point operations of SwiftNet-RN18 by 60% and increases the inference speed by 50%, with only 0.3% decrease in mIoU accuracy on Cityscapes.
Thomas Verelst, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Residual Tuning: Toward Novel Category Discovery Without Labels
abstract
Discovering novel visual categories from a set of unlabeled images is a crucial and essential capability for intelligent vision systems since it enables them to automatically learn new concepts with no need for human-annotated supervision anymore. To tackle this problem, existing approaches first pretrain a neural network with a set of labeled images and then fine-tune the network to cluster unlabeled images into a few categorical groups. However, their unified feature representation hits a tradeoff bottleneck between feature preservation on labeled data and feature adaptation on unlabeled data. To circumvent this bottleneck, we propose a residual-tuning approach, which estimates a new residual feature from the pretrained network and adds it with a previous basic feature to compute the clustering objective together. Our disentangled representation approach facilitates adjusting visual representations for unlabeled images and overcoming forgetting old knowledge acquired from labeled images, with no need of replaying the labeled images again. In addition, residual-tuning is an efficient solution, adding few parameters and consuming modest training time. Our results on three common benchmarks show consistent and considerable gains over other state-of-the-art methods, and further reduce the performance gap to the fully supervised learning setup. Moreover, we explore two extended scenarios, including using fewer labeled classes and continually discovering more unlabeled sets, where the results further signify the advantages and effectiveness of our residual-tuning approach against previous approaches. Our code is available at https://github.com/liuyudut/ResTune.
Yu Liu 0012, Tinne Tuytelaars
IEEE Trans. Neural Networks Learn. Syst.2
2022 Category-Level Pose Retrieval with Contrastive Features Learnt with Occlusion Augmentation
Georgios Kouros, Shubham Shrivastava, Cédric Picron, Sushruth Nagesh, Punarjay Chakravarty, Tinne Tuytelaars
BMVC6
2022 Trident Pyramid Networks for Object Detection
Cédric Picron, Tinne Tuytelaars
BMVC2
2022 Re-examining Distillation for Continual Object Detection
Eli Verwimp, Sarah Parisot, Lanqing Hong, Steven McDonagh 0001, Eduardo Pérez-Pellitero, Matthias De Lange, Tinne Tuytelaars
BMVC8
2022 Generative Negative Text Replay for Continual Vision-Language Pretraining
Shipeng Yan, Lanqing Hong, Hang Xu 0004, Jianhua Han, Tinne Tuytelaars, Zhenguo Li, Xuming He 0001
ECCV (36)5
2022 New Insights on Reducing Abrupt Representation Change in Online Continual Learning
Lucas Caccia, Rahaf Aljundi, Nader Asadi, Tinne Tuytelaars, Joelle Pineau, Eugene Belilovsky
ICLR4
2022 Unsupervised Vision-Language Grammar Induction with Shared Structure Modeling
Wenjuan Han, Zilong Zheng, Tinne Tuytelaars
ICLR4
2022 Keep on Learning
Tinne Tuytelaars
ICPRAM1
2022 RARA: Zero-shot Sim2Real Visual Navigation with Following Foreground Cues
abstract
The gap between simulation and the real-world restrains many machine learning breakthroughs in computer vision and reinforcement learning from being applicable in the real world. In this work, we tackle this gap for the specific case of camera-based navigation, formulating it as following a visual cue in the foreground with arbitrary backgrounds. The visual cue in the foreground can often be simulated realistically, such as a line, gate or cone. The challenge then lies in coping with the unknown backgrounds and integrating both. As such, the goal is to train a visual agent on data captured in an empty simulated environment except for this foreground cue and test this model directly in a visually diverse real world. In order to bridge this big gap, we show it's crucial to combine following techniques namely: Randomized augmentation of the fore- and background, regularization with both deep supervision and triplet loss and finally abstraction of the dynamics by using waypoints rather than direct velocity commands. The various techniques are ablated in our experimental results both qualitatively and quantitatively finally demonstrating a successful transfer from simulation to the real world. Code will be made available on publication22Project page: github.com/kkelchte/tgbg.
Klaas Kelchtermans, Tinne Tuytelaars
IROS2
2022 A Continual Learning Survey: Defying Forgetting in Classification Tasks
abstract
Artificial neural networks thrive in solving the classification problem for a particular rigid task, acquiring knowledge through generalized learning behaviour from a distinct training phase. The resulting network resembles a static entity of knowledge, with endeavours to extend this knowledge without targeting the original task resulting in a catastrophic forgetting. Continual learning shifts this paradigm towards networks that can continually accumulate knowledge over different tasks without the need to retrain from scratch. We focus on task incremental classification, where tasks arrive sequentially and are delineated by clear boundaries. Our main contributions concern: (1) a taxonomy and extensive overview of the state-of-the-art; (2) a novel framework to continually determine the stability-plasticity trade-off of the continual learner; (3) a comprehensive experimental comparison of 11 state-of-the-art continual learning methods; and (4) baselines. We empirically scrutinize method strengths and weaknesses on three benchmarks, considering Tiny Imagenet and large-scale unbalanced iNaturalist and a sequence of recognition datasets. We study the influence of model capacity, weight decay and dropout regularization, and the order in which the tasks are presented, and qualitatively compare methods in terms of required memory, computation time, and storage.
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia 0012, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.8
2021 Revisiting spatio-temporal layouts for compositional action recognition
Gorjan Radevski, Marie-Francine Moens, Tinne Tuytelaars
BMVC3
2021 MinMaxCAM: Improving object coverage for CAM-based Weakly Supervised Object Localization
José Oramas M., Tinne Tuytelaars
BMVC3
2021 Continual Prototype Evolution: Learning Online from Non-Stationary Data Streams
abstract
Attaining prototypical features to represent class distributions is well established in representation learning. However, learning prototypes online from streaming data proves a challenging endeavor as they rapidly become outdated, caused by an ever-changing parameter space during the learning process. Additionally, continual learning does not assume the data stream to be stationary, typically resulting in catastrophic forgetting of previous knowledge. As a first, we introduce a system addressing both problems, where prototypes evolve continually in a shared latent space, enabling learning and prediction at any point in time. To facilitate learning, a novel objective function synchronizes the latent space with the continually evolving prototypes. In contrast to the major body of work in continual learning, data streams are processed in an online fashion without task information and can be highly imbalanced, for which we propose an efficient memory scheme. As an additional contribution, we propose the learner-evaluator framework that i) generalizes existing paradigms in continual learning, ii) introduces data incremental learning, and iii) models the bridge between continual learning and concept drift. We obtain state-of-the-art performance by a significant margin on eight benchmarks, including three highly imbalanced data streams. Code is publicly available.1
Matthias De Lange, Tinne Tuytelaars
ICCV2
2021 Glimpse-Attend-and-Explore: Self-Attention for Active Visual Exploration
abstract
Active visual exploration aims to assist an agent with a limited field of view to understand its environment based on partial observations made by choosing the best viewing directions in the scene. Recent methods have tried to ad-dress this problem either by using reinforcement learning, which is difficult to train, or by uncertainty maps, which are task-specific and can only be implemented for dense prediction tasks. In this paper, we propose the Glimpse-Attend-and-Explore model which: (a) employs self-attention to guide the visual exploration instead of task-specific uncertainty maps; (b) can be used for both dense and sparse prediction tasks; and (c) uses a contrastive stream to further improve the representations learned. Unlike previous works, we show the application of our model on multiple tasks like reconstruction, segmentation and classification. Our model provides encouraging results while being less dependent on dataset bias in driving the exploration. We further perform an ablation study to investigate the features and attention learned by our model. Finally, we show that our self-attention module learns to attend different regions of the scene by minimizing the loss on the downstream task. Code: https://github.com/soroushseifi/glimpse-attend-explore.
Soroush Seifi, Abhishek Jha 0001, Tinne Tuytelaars
ICCV3
2021 BlockCopy: High-Resolution Video Processing with Block-Sparse Feature Propagation and Online Policies
abstract
In this paper we propose BlockCopy, a scheme that accelerates pretrained frame-based CNNs to process video more efficiently, compared to standard frame-by-frame processing. To this end, a lightweight policy network determines important regions in an image, and operations are applied on selected regions only, using custom block-sparse convolutions. Features of non-selected regions are simply copied from the preceding frame, reducing the number of computations and latency. The execution policy is trained using reinforcement learning in an online fashion without requiring ground truth annotations. Our universal framework is demonstrated on dense prediction tasks such as pedestrian detection, instance segmentation and semantic segmentation, using both state of the art (Center and Scale Predictor, MGAN, SwiftNet) and standard baseline networks (Mask-RCNN, DeepLabV3+). BlockCopy achieves significant FLOPS savings and inference speedup with minimal impact on accuracy.
Thomas Verelst, Tinne Tuytelaars
ICCV2
2021 Rehearsal revealed: The limits and merits of revisiting samples in continual learning
abstract
Learning from non-stationary data streams and overcoming catastrophic forgetting still poses a serious challenge for machine learning research. Rather than aiming to improve state-of-the-art, in this work we provide insight into the limits and merits of rehearsal, one of continual learning’s most established methods. We hypothesize that models trained sequentially with rehearsal tend to stay in the same low-loss region after a task has finished, but are at risk of overfitting on its sample memory, hence harming generalization. We provide both conceptual and strong empirical evidence on three benchmarks for both behaviors, bringing novel insights into the dynamics of rehearsal and continual learning in general. Finally, we interpret important continual learning works in the light of our findings, allowing for a deeper understanding of their successes.1
Eli Verwimp, Matthias De Lange, Tinne Tuytelaars
ICCV3
2021 What My Motion tells me about Your Pose: A Self-Supervised Monocular 3D Vehicle Detector
abstract
The estimation of the orientation of an observed vehicle relative to an Autonomous Vehicle (AV) from monocular camera data is an important building block in estimating its 6 DoF pose. Current Deep Learning based solutions for placing a 3D bounding box around this observed vehicle are data hungry and do not generalize well. In this paper, we demonstrate the use of monocular visual odometry for the self-supervised fine-tuning of a model for orientation estimation pre-trained on a reference domain. Specifically, while transitioning from a virtual dataset (vKITTI) to nuScenes, we recover up to 70% of the performance of a fully supervised method. We subsequently demonstrate an optimization-based monocular 3D bounding box detector built on top of the self-supervised vehicle orientation estimator without the requirement of expensive labeled data. This allows 3D vehicle detection algorithms to be self-trained from large amounts of monocular camera data from existing commercial vehicle fleets.
Cédric Picron, Punarjay Chakravarty, Tom Roussel, Tinne Tuytelaars
ICRA4
2021 Unsupervised Motion Estimation of Vehicles Using ICP
Tom Roussel, Tinne Tuytelaars, Luc Van Eycken
ICRA2
2021 Processor Architecture Optimization for Spatially Dynamic Neural Networks
abstract
Spatially dynamic neural networks adjust network execution based on the input data, saving computations by skipping non-important image regions. Yet, GPU implementations fail to achieve speedups from these spatially dynamic execution patterns for most neural network architectures. This paper investigates hardware constraints preventing such speedup and proposes and compares novel processor architectures and dataflows enabling latency improvements due to the dynamic execution with minimal loss of utilization. The presented architectures flexibly support spatial execution of a broad range of networks. For the derived architectures, the spatial unrolling for each layer type is optimized and validated making use of the ZigZag design space exploration framework where appropriate. This allows to benchmark and compare the hardware architectures on NNs for classification and human pose estimation, increasing throughput up to $\times 1.9$ and $\times 2.3$ compared to their static executions, respectively. This is the same order of magnitude as other dynamic execution methods, while being complementary to those.
Steven Colleman, Thomas Verelst, Linyan Mei, Tinne Tuytelaars, Marian Verhelst
VLSI-SoC4
2021 Show me where the action is!
abstract
Abstract Reality TV shows have gained popularity, motivating many production houses to bring new variants for us to watch. Compared to traditional TV shows, reality TV shows have spontaneous unscripted footage. Computer vision techniques could partially replace the manual labour needed to record and process this spontaneity. However, automated real-world video recording and editing is a challenging topic. In this paper, we propose a system that utilises state-of-the-art video and audio processing algorithms to, on the one hand, automatically steer cameras, replacing camera operators and on the other hand, detect all audiovisual action cues in the recorded video, to ease the job of the film editor. This publication has hence two main contributions. The first, automating the steering of multiple Pan-Tilt-Zoom PTZ cameras to take aesthetically pleasing medium shots of all the people present. These shots need to comply with the cinematographic rules and are based on the poses acquired by a pose detector. Secondly, when a huge amount of audio-visual data has been collected, it becomes labour intensive for a human editor retrieve the relevant fragments. As a second contribution, we combine state-of-the-art audio and video processing techniques for sound activity detection, action recognition, face recognition, and pose detection to decrease the required manual labour during and after recording. These techniques used during post-processing produce meta-data allowing for footage filtering, decreasing the search space. We extended our system further by producing timelines uniting generated meta-data, allowing the editor to have a quick overview. We evaluated our system on three in-the-wild reality TV recording sessions of 24 hours (× 8 cameras) each taken in real households.
Timothy Callemein, Tom Roussel, Ali Diba, Floris De Feyter, Wim Boes, Luc Van Eycken, Luc Van Gool, Hugo Van hamme, Tinne Tuytelaars, Toon Goedemé
Multim. Tools Appl.9
2020 Learning Multi-instance Sub-pixel Point Localization
Julien Schroeter, Tinne Tuytelaars, Kirill A. Sidorov, David Marshall 0001
ACCV (5)2
2020 MIX'EM: Unsupervised Image Classification Using a Mixture of Embeddings
Ali Varamesh, Tinne Tuytelaars
ACCV (3)2
2020 Multiple Exemplars-Based Hallucination for Face Super-Resolution and Editing
José Oramas M., Tinne Tuytelaars
ACCV (5)3
2020 In Defense of LSTMs for Addressing Multiple Instance Learning Problems
José Oramas M., Tinne Tuytelaars
ACCV (6)3
2020 On the Exploration of Incremental Learning for Fine-grained Image Retrieval
Wei Chen 0072, Yu Liu 0012, Weiping Wang 0002, Tinne Tuytelaars, Erwin M. Bakker, Michael S. Lew
BMVC4
2020 A Novel Baseline for Zero-shot Learning via Adversarial Visual-Semantic Embedding
Yu Liu 0012, Tinne Tuytelaars
BMVC2
2020 Learning to ground medical text in a 3D human atlas
abstract
In this paper, we develop a method for grounding medical text into a physically meaningful and interpretable space corresponding to a human atlas.We build on text embedding architectures such as BERT and introduce a loss function that allows us to reason about the semantic and spatial relatedness of medical texts by learning a projection of the embedding into a 3D space representing the human body.We quantitatively and qualitatively demonstrate that our proposed method learns a context sensitive and spatially aware mapping, in both the inter-organ and intra-organ sense, using a large scale medical text dataset from the "Large-scale online biomedical semantic indexing" track of the 2020 BioASQ challenge.We extend our approach to a self-supervised setting, and find it to be competitive with a classification based method, and a fully supervised variant of approach.
Dusan Grujicic, Gorjan Radevski, Tinne Tuytelaars, Matthew B. Blaschko
CoNLL3
2020 Unsupervised Model Personalization While Preserving Privacy and Scalability: An Open Problem
abstract
This work investigates the task of unsupervised model personalization, adapted to continually evolving, unlabeled local user images. We consider the practical scenario where a high capacity server interacts with a myriad of resource-limited edge devices, imposing strong requirements on scalability and local data privacy. We aim to address this challenge within the continual learning paradigm and provide a novel Dual User-Adaptation framework (DUA) to explore the problem. This framework flexibly disentangles user-adaptation into model personalization on the server and local data regularization on the user device, with desirable properties regarding scalability and privacy constraints. First, on the server, we introduce incremental learning of task-specific expert models, subsequently aggregated using a concealed unsupervised user prior. Aggregation avoids retraining, whereas the user prior conceals sensitive raw user data, and grants unsupervised adaptation. Second, local user-adaptation incorporates a domain adaptation point of view, adapting regularizing batch normalization parameters to the user data. We explore various empirical user configurations with different priors in categories and a tenfold of transforms for MIT Indoor Scene recognition, and classify numbers in a combined MNIST and SVHN setup. Extensive experiments yield promising results for data-driven local adaptation and elicit user priors for server adaptation to depend on the model rather than user data. Hence, although user-adaptation remains a challenging open problem, the DUA framework formalizes a principled foundation for personalizing both on server and user device, while maintaining privacy and scalability.
Matthias De Lange, Xu Jia 0012, Sarah Parisot, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars
CVPR6
2020 Mixture Dense Regression for Object Detection and Human Pose Estimation
abstract
Mixture models are well-established learning approaches that, in computer vision, have mostly been applied to inverse or ill-defined problems. However, they are general-purpose divide-and-conquer techniques, splitting the input space into relatively homogeneous subsets in a data-driven manner. Not only ill-defined but also well-defined complex problems should benefit from them. To this end, we devise a framework for spatial regression using mixture density networks. We realize the framework for object detection and human pose estimation. For both tasks, a mixture model yields higher accuracy and divides the input space into interpretable modes. For object detection, mixture components focus on object scale, with the distribution of components closely following that of ground truth the object scale. This practically alleviates the need for multi-scale testing, providing a superior speed-accuracy trade-off. For human pose estimation, a mixture model divides the data based on viewpoint and uncertainty - namely, front and back views, with back view imposing higher uncertainty. We conduct experiments on the MS COCO dataset and do not face any mode collapse.
Ali Varamesh, Tinne Tuytelaars
CVPR2
2020 Dynamic Convolutions: Exploiting Spatial Sparsity for Faster Inference
abstract
Modern convolutional neural networks apply the same operations on every pixel in an image. However, not all image regions are equally important. To address this inefficiency, we propose a method to dynamically apply convolutions conditioned on the input image. We introduce a residual block where a small gating branch learns which spatial positions should be evaluated. These discrete gating decisions are trained end-to-end using the Gumbel-Softmax trick, in combination with a sparsity criterion. Our experiments on CIFAR, ImageNet, Food-101 and MPII show that our method has better focus on the region of interest and better accuracy than existing methods, at a lower computational complexity. Moreover, we provide an efficient CUDA implementation of our dynamic convolutions using a gather-scatter approach, achieving a significant improvement in inference speed on MobileNetV2 and ShuffleNetV2. On human pose estimation, a task that is inherently spatially sparse, the processing speed is increased by 60% with no loss in accuracy.
Thomas Verelst, Tinne Tuytelaars
CVPR2
2020 More Classifiers, Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning
Yu Liu 0012, Sarah Parisot, Gregory Slabaugh, Xu Jia 0012, Ales Leonardis, Tinne Tuytelaars
ECCV (26)6
2020 Attend and Segment: Attention Guided Active Semantic Segmentation
Soroush Seifi, Tinne Tuytelaars
ECCV (25)2
2020 Unpaired Image-To-Image Shape Translation Across Fashion Data
abstract
We address the problem of unpaired geometric image-to-image translation. Rather than transferring the style of an image as a whole, our goal is to translate the geometry of an object while preserving its appearance. Our model is trained without the need for paired images. It performs all steps of the shape transfer within a single model and without additional post-processing stages. Experiments on clothing-based datasets show the effectiveness of the proposed method.
Liqian Ma, José Oramas M., Luc Van Gool, Tinne Tuytelaars
ICIP5
2020 $\mathbb {H}$H-Patches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors
abstract
In this paper, a novel benchmark is introduced for evaluating local image descriptors. We demonstrate limitations of the commonly used datasets and evaluation protocols, that lead to ambiguities and contradictory results in the literature. Furthermore, these benchmarks are nearly saturated due to the recent improvements in local descriptors obtained by learning from large annotated datasets. To address these issues, we introduce a new large dataset suitable for training and testing modern descriptors, together with strictly defined evaluation protocols in several tasks such as matching, retrieval and verification. This allows for more realistic, thus more reliable comparisons in different application scenarios. We evaluate the performance of several state-of-the-art descriptors and analyse their properties. We show that a simple normalisation of traditional hand-crafted descriptors is able to boost their performance to the level of deep learning based descriptors once realistic benchmarks are considered. Additionally we specify a protocol for learning and evaluating using cross validation. We show that when training state-of-the-art descriptors on this dataset, the traditional verification task is almost entirely saturated.
Vassileios Balntas, Karel Lenc, Andrea Vedaldi, Tinne Tuytelaars, Jiri Matas, Krystian Mikolajczyk
IEEE Trans. Pattern Anal. Mach. Intell.4
2020 A Deep Multi-Modal Explanation Model for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) has attracted significant attention due to its capabilities of classifying new images from unseen classes. To perform the classification task for ZSL, learning visual and semantic embeddings has been the main research approach in existing literature. At the same time, generating complementary explanations to justify the classification decision has remained largely unexplored. In this paper, we propose to address a new and challenging task, namely explainable zero-shot learning (XZSL), which aims to generate visual and textual explanations to support the classification decision. To accomplish this task, we build a novel Deep Multi-modal Explanation (DME) model that incorporates a joint visual-attribute embedding module and a multi-channel explanation module in an end-to-end fashion. In contrast to existing ZSL approaches, our visual-attribute embedding is associated not only with the decision, but also with new visual and textual explanations. For visual explanations, we first capture several attribute activation maps (AAM) and then merge them into a class activation map (CAM) that visually infers which region of an image is relevant to the class. Textual explanations are generated from the multi-channel explanation module, jointly integrating three long short-term memory models (LSTMs) each of which is conditioned on a different feature representation. Additionally, we suggest that the DME model can retain explanatory consistency for similar instances and explanatory diversity for diverse instances. We conduct qualitative and quantitative experiments to assess the model for ZSL classification and explanation. Specifically, the ablation studies verify the effectiveness of the components in our model. Our results on three well-known datasets are competitive with prior approaches. More importantly, the joint training of our embedding and explanation modules demonstrates mutual performance improvements between ZSL classification and explanation. We shed more light on DME to analyze and diagnose its advantages and limitations.
Yu Liu 0012, Tinne Tuytelaars
IEEE Trans. Image Process.2
2019 Task-Free Continual Learning
abstract
Methods proposed in the literature towards continual deep learning typically operate in a task-based sequential learning setup. A sequence of tasks is learned, one at a time, with all data of current task available but not of previous or future tasks. Task boundaries and identities are known at all times. This setup, however, is rarely encountered in practical applications. Therefore we investigate how to transform continual learning to an online setup. We develop a system that keeps on learning over time in a streaming fashion, with data distributions gradually changing and without the notion of separate tasks. To this end, we build on the work on Memory Aware Synapses, and show how this method can be made online by providing a protocol to decide i) when to update the importance weights, ii) which data to use to update them, and iii) how to accumulate the importance weights at each update step. Experimental results show the validity of the approach in the context of two applications: (self-)supervised learning of a face recognition model by watching soap series and learning a robot to avoid collisions.
Rahaf Aljundi, Klaas Kelchtermans, Tinne Tuytelaars
CVPR3
2019 Towards Object Shape Translation Through Unsupervised Generative Deep Models
abstract
This paper focuses on the problem of unsupervised image-to-image translation. More specifically, we aim at finding a translation network such that objects and shapes that only appear in the source domain are translated to objects and shapes only appearing in the target domain, while style color features present in the source domain remain the same. To achieve this, we use a domain-specific variational autoencoder and represent each image in its latent space representation. In a second step, we learn a translation between latent spaces of different domains using generative adversarial networks. We evaluate this framework on multiple datasets and verify the effect of multiple perceptual losses. Experiments on the MNIST and SVHN datasets show the effectiveness of the proposed translation method.
Lies Bollens, Tinne Tuytelaars, José Oramas M.
ICIP2
2019 Selfless Sequential Learning
Rahaf Aljundi, Marcus Rohrbach, Tinne Tuytelaars
ICLR (Poster)3
2019 Visual Explanation by Interpretation: Improving Visual Feedback Capabilities of Deep Neural Networks
José Oramas M., Tinne Tuytelaars
ICLR (Poster)3
2019 Exemplar Guided Unsupervised Image-to-Image Translation with Semantic Consistency
Liqian Ma, Xu Jia 0012, Stamatios Georgoulis, Tinne Tuytelaars, Luc Van Gool
ICLR (Poster)4
2019 Monocular Depth Estimation in New Environments With Absolute Scale
abstract
In this work we propose an unsupervised training method that finetunes a single image depth estimation CNN towards a new environment. The network, which has been pretrained on stereo data, only requires monocular input for finetuning. Unlike other unsupervised methods, it produces depth estimations with absolute scale - a feature that is essential for most practical applications, yet has mostly been overlooked in the literature. First, we show how our method allows adapting a network trained on one dataset (Cityscapes) to another (KITTI). Next, by splitting KITTI in subsets, we show the sensitivity of pretrained models to a domain shift. We then demonstrate that, by finetuning the model using our method, it is possible to improve the performance on the target subset, without using stereo or any form of groundtruth depth and with preservation of the correct absolute scale.
Tom Roussel, Luc Van Eycken, Tinne Tuytelaars
IROS3
2019 Online Continual Learning with Maximal Interfered Retrieval
abstract
Continual learning, the setting where a learning agent is faced with a never-ending stream of data, continues to be a great challenge for modern machine learning systems. In particular the online or "single-pass through the data" setting has gained attention recently as a natural setting that is difficult to tackle. Methods based on replay, either generative or from a stored memory, have been shown to be effective approaches for continual learning, matching or exceeding the state of the art in a number of standard benchmarks. These approaches typically rely on randomly selecting samples from the replay memory or from a generative model, which is suboptimal. In this work, we consider a controlled sampling of memories for replay. We retrieve the samples which are most interfered, i.e. whose prediction will be most negatively impacted by the foreseen parameters update. We show a formulation for this sampling criterion in both the generative replay and the experience replay setting, producing consistent gains in performance and greatly reduced forgetting. We release an implementation of our method at https://github.com/optimass/MaximallyInterferedRetrieval
Rahaf Aljundi, Eugene Belilovsky, Tinne Tuytelaars, Laurent Charlin, Massimo Caccia, Lucas Caccia
NeurIPS3
2018 Exploring the Challenges Towards Lifelong Fact Learning
Mohamed Elhoseiny 0001, Francesca Babiloni, Rahaf Aljundi, Marcus Rohrbach, Manohar Paluri, Tinne Tuytelaars
ACCV (6)6
2018 Memory Aware Synapses: Learning What (not) to Forget
Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny 0001, Marcus Rohrbach, Tinne Tuytelaars
ECCV (3)5
2018 The CAMETRON Lecture Recording System: High Quality Video Recording and Editing with Minimal Human Supervision
Dries Hulens, Bram Aerts, Punarjay Chakravarty, Ali Diba, Toon Goedemé, Tom Roussel, Jeroen Zegers, Tinne Tuytelaars, Luc Van Eycken, Luc Van Gool, Hugo Van hamme, Joost Vennekens
MMM (1)8
2018 Modeling Temporal Structure with LSTM for Online Action Detection
abstract
Online action detection is a challenging problem: a system needs to decide what action is happening at the current frame, based on previous frames only. Fortunately in real-life, human actions are not independent from one another: there are strong (long-term) dependencies between them. An online action detection method should be able to capture these dependencies, to enable a more accurate early detection. At first sight, an LSTM seems very suitable for this problem. It is able to model both short-term and long-term patterns. It takes its input one frame at the time, updates its internal state and gives as output the current class probabilities. In practice, however, the detection results obtained with LSTMs are still quite low. In this work, we start from the hypothesis that it may be too difficult for an LSTM to learn both the interpretation of the input and the temporal patterns at the same time. We propose a two-stream feedback network, where one stream processes the input and the other models the temporal relations. We show improved detection accuracy on an artificial toy dataset and on the Breakfast Dataset [21] and the TVSeries Dataset [7], reallife datasets with inherent temporal dependencies between the actions.
Roeland De Geest, Tinne Tuytelaars
WACV2
2018 From Pixels to Actions: Learning to Drive a Car with Deep Neural Networks
abstract
The promise of self-driving cars promotes several advantages, e.g. they have the ability to outperform human drivers while being safer. Here we take a deeper look into some aspects from algorithms aimed at making this promise a reality. More specifically, we analyze an end-to-end neural network to predict a car's steering actions on a highway based on images taken from a single car-mounted camera. We focus our analysis on several aspects which could have a significant impact on the performance of the system. These aspects are: the input data format, the temporal dependencies between consecutive inputs, and the origin of the data. We show that, for the task at hand, regression networks outperform their classifier counterparts. In addition, there seems to be a small difference between networks that use coloured images and ones that use grayscale images as input. For the second aspect, by feeding the network three concatenated images, we get a significant decrease of 30% in mean squared error. For the third aspect, by using simulation data we are able to train networks that have a performance comparable to networks trained on real-life datasets. We also qualitatively demonstrate that the standard metrics that are used to evaluate networks do not necessarily accurately reflect a system's driving behaviour. We show that a promising confusion matrix may result in poor driving behaviour while a very ill-looking confusion matrix may result in good driving behaviour.
Jonas Heylen, Seppe Iven, Bert De Brabandere, José Oramas M., Luc Van Gool, Tinne Tuytelaars
WACV6
2018 An Analysis of Human-Centered Geolocation
abstract
Online social networks contain a constantly increasing amount of images - most of them focusing on people. Due to cultural and climate factors, fashion trends and physical appearance of individuals differ from city to city. In this paper we investigate to what extent such cues can be exploited in order to infer the geographic location, i.e. the city, where a picture was taken. We conduct a user study, as well as an evaluation of automatic methods based on convolutional neural networks. Experiments on the Fashion 144k and a Pinterest-based dataset show that the automatic methods succeed at this task to a reasonable extent. As a matter of fact, our empirical results suggest that automatic methods can surpass human performance by a large margin. Further inspection of the trained models shows that human-centered characteristics, like clothing style, physical features, and accessories, are informative for the task at hand. Moreover, it reveals that also contextual features, e.g. wall type, natural environment, etc., are taken into account by the automatic methods.
Yu-Hui Huang, José Oramas M., Luc Van Gool, Tinne Tuytelaars
WACV5
2018 Reflectance and Natural Illumination from Single-Material Specular Objects Using Deep Learning
abstract
In this paper, we present a method that estimates reflectance and illumination information from a single image depicting a single-material specular object from a given class under natural illumination. We follow a data-driven, learning-based approach trained on a very large dataset, but in contrast to earlier work we do not assume one or more components (shape, reflectance, or illumination) to be known. We propose a two-step approach, where we first estimate the object's reflectance map, and then further decompose it into reflectance and illumination. For the first step, we introduce a Convolutional Neural Network (CNN) that directly predicts a reflectance map from the input image itself, as well as an indirect scheme that uses additional supervision, first estimating surface orientation and afterwards inferring the reflectance map using a learning-based sparse data interpolation technique. For the second step, we suggest a CNN architecture to reconstruct both Phong reflectance parameters and high-resolution spherical illumination maps from the reflectance map. We also propose new datasets to train these CNNs. We demonstrate the effectiveness of our approach for both steps by extensive quantitative and qualitative evaluation in both synthetic and real data as well as through numerous applications, that show improvements over the state-of-the-art.
Stamatios Georgoulis, Konstantinos Rematas, Tobias Ritschel 0001, Efstratios Gavves, Mario Fritz, Luc Van Gool, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.7
2017 Expert Gate: Lifelong Learning with a Network of Experts
abstract
In this paper we introduce a model of lifelong learning, based on a Network of Experts. New tasks / experts are learned and added to the model sequentially, building on what was learned before. To ensure scalability of this process, data from previous tasks cannot be stored and hence is not available when learning a new task. A critical issue in such context, not addressed in the literature so far, relates to the decision which expert to deploy at test time. We introduce a set of gating autoencoders that learn a representation for the task at hand, and, at test time, automatically forward the test sample to the relevant expert. This also brings memory efficiency as only one expert network has to be loaded into memory at any given time. Further, the autoencoders inherently capture the relatedness of one task to another, based on which the most relevant prior model to be used for training a new expert, with fine-tuning or learning-without-forgetting, can be selected. We evaluate our method on image classification and video prediction problems.
Rahaf Aljundi, Punarjay Chakravarty, Tinne Tuytelaars
CVPR3
2017 What is Around the Camera?
abstract
How much does a single image reveal about the environment it was taken in? In this paper, we investigate how much of that information can be retrieved from a foreground object, combined with the background (i.e. the visible part of the environment). Assuming it is not perfectly diffuse, the foreground object acts as a complexly shaped and far-from-perfect mirror An additional challenge is that its appearance confounds the light coming from the environment with the unknown materials it is made of. We propose a learning-based approach to predict the environment from multiple reflectance maps that are computed from approximate surface normals. The proposed method allows us to jointly model the statistics of environments and material properties. We train our system from synthesized training data, but demonstrate its applicability to real-world data. Interestingly, our analysis shows that the information obtained from objects made out of multiple materials often is complementary and leads to better performance.
Stamatios Georgoulis, Konstantinos Rematas, Tobias Ritschel 0001, Mario Fritz, Tinne Tuytelaars, Luc Van Gool
ICCV5
2017 Encoder Based Lifelong Learning
abstract
This paper introduces a new lifelong learning solution where a single model is trained for a sequence of tasks. The main challenge that vision systems face in this context is catastrophic forgetting: as they tend to adapt to the most recently seen task, they lose performance on the tasks that were learned previously. Our method aims at preserving the knowledge of the previous tasks while learning a new one by using autoencoders. For each task, an under-complete autoencoder is learned, capturing the features that are crucial for its achievement. When a new task is presented to the system, we prevent the reconstructions of the features with these autoencoders from changing, which has the effect of preserving the information on which the previous tasks are mainly relying. At the same time, the features are given space to adjust to the most recent environment as only their projection into a low dimension submanifold is controlled. The proposed system is evaluated on image classification tasks and shows a reduction of forgetting over the state-ofthe-art.
Amal Rannen Triki, Rahaf Aljundi, Matthew B. Blaschko, Tinne Tuytelaars
ICCV4
2017 CNN-based single image obstacle avoidance on a quadrotor
abstract
This paper demonstrates the use of a single forward facing camera for obstacle avoidance on a quadrotor. We train a CNN for estimating depth from a single image. The depth map is then fed to a behaviour arbitration based control algorithm that steers the quadrotor away from obstacles. We conduct experiments with simulated and real drones in a variety of environments.
Punarjay Chakravarty, Klaas Kelchtermans, Tom Roussel, Stijn Wellens, Tinne Tuytelaars, Luc Van Eycken
ICRA5
2017 Pose Guided Person Image Generation
abstract
This paper proposes the novel Pose Guided Person Generation Network (PG$^2$) that allows to synthesize person images in arbitrary poses, based on an image of that person and a novel pose. Our generation framework PG$^2$ utilizes the pose information explicitly and consists of two key stages: pose integration and image refinement. In the first stage the condition image and the target pose are fed into a U-Net-like network to generate an initial but coarse image of the person with the target pose. The second stage then refines the initial and blurry result by training a U-Net-like generator in an adversarial way. Extensive experimental results on both 128$\times$64 re-identification images and 256$\times$256 fashion photos show that our model generates high-quality person images with convincing details.
Liqian Ma, Xu Jia 0012, Qianru Sun, Bernt Schiele, Tinne Tuytelaars, Luc Van Gool
NIPS5
2017 Context-based object viewpoint estimation: A 2D relational approach
José Oramas M., Luc De Raedt, Tinne Tuytelaars
Comput. Vis. Image Underst.3
2017 DeepProposals: Hunting Objects and Actions by Cascading Deep Convolutional Layers
abstract
In this paper, a new method for generating object and action proposals in images and videos is proposed. It builds on activations of different convolutional layers of a pretrained CNN, combining the localization accuracy of the early layers with the high informativeness (and hence recall) of the later layers. To this end, we build an inverse cascade that, going backward from the later to the earlier convolutional layers of the CNN, selects the most promising locations and refines them in a coarse-to-fine manner. The method is efficient, because (i) it re-uses the same features extracted for detection, (ii) it aggregates features using integral images, and (iii) it avoids a dense evaluation of the proposals thanks to the use of the inverse coarse-to-fine cascade. The method is also accurate. We show that DeepProposals outperform most of the previous object proposal and action proposal approaches and, when plugged into a CNN-based object detector, produce state-of-the-art detection performance.
Amir Ghodrati, Ali Diba, Marco Pedersoli, Tinne Tuytelaars, Luc Van Gool
Int. J. Comput. Vis.4
2017 Entity linking across vision and language
Aparna Nurani Venkitasubramanian, Tinne Tuytelaars, Marie-Francine Moens
Multim. Tools Appl.2
2017 Rank Pooling for Action Recognition
abstract
We propose a function-based temporal pooling method that captures the latent structure of the video sequence data - e.g., how frame-level features evolve over time in a video. We show how the parameters of a function that has been fit to the video data can serve as a robust new video representation. As a specific example, we learn a pooling function via ranking machines. By learning to rank the frame-level features of a video in chronological order, we obtain a new representation that captures the video-wide temporal dynamics of a video, suitable for action recognition. Other than ranking functions, we explore different parametric models that could also explain the temporal changes in videos. The proposed functional pooling methods, and rank pooling in particular, is easy to interpret and implement, fast to compute and effective in recognizing a wide variety of actions. We evaluate our method on various benchmarks for generic action, fine-grained action and gesture recognition. Results show that rank pooling brings an absolute improvement of 7-10 average pooling baseline. At the same time, rank pooling is compatible with and complementary to several appearance and local motion based methods and features, such as improved trajectories and deep learning features.
Basura Fernando, Efstratios Gavves, José Oramas M., Amir Ghodrati, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.5
2017 Novel Views of Objects from a Single Image
abstract
Taking an image of an object is at its core a lossy process. The rich information about the three-dimensional structure of the world is flattened to an image plane and decisions such as viewpoint and camera parameters are final and not easily revertible. As a consequence, possibilities of changing viewpoint are limited. Given a single image depicting an object, novel-view synthesis is the task of generating new images that render the object from a different viewpoint than the one given. The main difficulty is to synthesize the parts that are disoccluded; disocclusion occurs when parts of an object are hidden by the object itself under a specific viewpoint. In this work, we show how to improve novel-view synthesis by making use of the correlations observed in 3D models and applying them to new image instances. We propose a technique to use the structural information extracted from a 3D model that matches the image object in terms of viewpoint and shape. For the latter part, we propose an efficient 2D-to-3D alignment method that associates precisely the image appearance with the 3D model geometry with minimal user interaction. Our technique is able to simulate plausible viewpoint changes for a variety of object classes within seconds. Additionally, we show that our synthesized images can be used as additional training data that improves the performance of standard object detectors.
Konstantinos Rematas, Chuong H. Nguyen, Tobias Ritschel 0001, Mario Fritz, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.5
2016 Who's that Actor? Automatic Labelling of Actors in TV Series Starting from IMDB Images
Rahaf Aljundi, Punarjay Chakravarty, Tinne Tuytelaars
ACCV (3)3
2016 Towards Automatic Image Editing: Learning to See another You
Xu Jia 0012, Amir Ghodrati, Marco Pedersoli, Tinne Tuytelaars
BMVC4
2016 Deep Reflectance Maps
abstract
Undoing the image formation process and therefore decomposing appearance into its intrinsic properties is a challenging task due to the under-constrained nature of this inverse problem. While significant progress has been made on inferring shape, materials and illumination from images only, progress in an unconstrained setting is still limited. We propose a convolutional neural architecture to estimate reflectance maps of specular materials in natural lighting conditions. We achieve this in an end-to-end learning formulation that directly predicts a reflectance map from the image itself. We show how to improve estimates by facilitating additional supervision in an indirect scheme that first predicts surface orientation and afterwards predicts the reflectance map by a learning-based sparse data interpolation. In order to analyze performance on this difficult task, we propose a new challenge of Specular MAterials on SHapes with complex IllumiNation (SMASHINg) using both synthetic and real images. Furthermore, we show the application of our method to a range of image editing tasks on real images.
Konstantinos Rematas, Tobias Ritschel 0001, Mario Fritz, Efstratios Gavves, Tinne Tuytelaars
CVPR5
2016 Cross-Modal Supervision for Learning Active Speaker Detection in Video
Punarjay Chakravarty, Tinne Tuytelaars
ECCV (5)2
2016 Online Action Detection
Roeland De Geest, Efstratios Gavves, Amir Ghodrati, Cees Snoek, Tinne Tuytelaars
ECCV (5)6
2016 Pose Estimation Errors, the Ultimate Diagnosis
Carolina Redondo-Cabrera, Roberto Javier López-Sastre, Tinne Tuytelaars, Silvio Savarese
ECCV (7)4
2016 Active speaker detection with audio-visual co-training
abstract
In this work, we show how to co-train a classifier for active speaker detection using audio-visual data. First, audio Voice Activity Detection (VAD) is used to train a personalized video-based active speaker classifier in a weakly supervised fashion. The video classifier is in turn used to train a voice model for each person. The individual voice models are then used to detect active speakers. There is no manual supervision - audio weakly supervises video classification, and the co-training loop is completed by using the trained video classifier to supervise the training of a personalized audio voice classifier.
Punarjay Chakravarty, Jeroen Zegers, Tinne Tuytelaars, Hugo Van hamme
ICMI3
2016 Vision and Language Integration Meets Multimedia Fusion: Proceedings of ACM Multimedia 2016 Workshop
abstract
Multimodal information fusion both at the signal and the semantics levels is a core part in most multimedia applications, including multimedia indexing, retrieval, summarization and others. Early or late fusion of modality-specific processing results has been addressed in multimedia prototypes since their very early days, through various methodologies including rule-based approaches, information-theoretic models and machine learning. Vision and Language are two of the predominant modalities that are being fused and which have attracted special attention in international challenges with a long history of results, such as TRECVid, ImageClef and others. During the last decade, vision-language semantic integration has attracted attention from traditionally non-interdisciplinary research communities, such as Computer Vision and Natural Language Processing. This is due to the fact that one modality can greatly assist the processing of another providing cues for disambiguation, complementary information and noise/error filtering. The latest boom of deep learning methods has opened up new directions in joint modelling of visual and co-occurring verbal information in multimedia discourse. The workshop on Vision and Language Integration Meets Multimedia Fusion has been held during the workshop weekend of the ACM Multimedia 2016 Conference and the European Conference on Computer Vision (ECCV 2016) on October 16, 2016 in Amsterdam, the Netherlands. The proceedings contain seven selected long papers, which have been orally presented at the workshop, and three abstracts of the invited keynote speeches. The papers and abstracts discuss data collection, representation learning, deep learning approaches, matrix and tensor factorization methods and graph based clustering with regard to the fusion of multimedia data. A variety of applications is presented including image captioning, summarization of news, video hyperlinking, sub-shot segmentation of user generated video, cross-modal classification, cross-modal question-answering, and the detection of misleading metadata of user generated video. The workshop is organized and supported by the EU COST action iV&L Net, the European Network on Integrating Vision and Language: Combining Computer Vision and Language Processing for Advanced Search, Retrieval, Annotation and Description of Visual Data (IC 1307--2014-2018).
Marie-Francine Moens, Katerina Pastra, Kate Saenko, Tinne Tuytelaars
ACM Multimedia4
2016 Dynamic Filter Networks
abstract
In a traditional convolutional layer, the learned filters stay fixed after training. In contrast, we introduce a new framework, the Dynamic Filter Network, where filters are generated dynamically conditioned on an input. We show that this architecture is a powerful one, with increased flexibility thanks to its adaptive nature, yet without an excessive increase in the number of model parameters. A wide variety of filtering operation can be learned this way, including local spatial transformations, but also others like selective (de)blurring or adaptive feature extraction. Moreover, multiple such layers can be combined, e.g. in a recurrent architecture. We demonstrate the effectiveness of the dynamic filter network on the tasks of video and stereo prediction, and reach state-of-the-art performance on the moving MNIST dataset with a much smaller model. By visualizing the learned filters, we illustrate that the network has picked up flow information by only looking at unlabelled training data. This suggests that the network can be used to pretrain networks for various supervised tasks in an unsupervised way, like optical flow and depth estimation.
Xu Jia 0012, Bert De Brabandere, Tinne Tuytelaars, Luc Van Gool
NIPS3
2016 Recovering hard-to-find object instances by sampling context-based object proposals
José Oramas M., Tinne Tuytelaars
Comput. Vis. Image Underst.2
2016 Wildlife recognition in nature documentaries with weak supervision from subtitles and external data
Aparna Nurani Venkitasubramanian, Tinne Tuytelaars, Marie-Francine Moens
Pattern Recognit. Lett.2
2016 Scalable Semi-Automatic Annotation for Multi-Camera Person Tracking
abstract
This paper proposes a generic methodology for semi-automatic generation of reliable position annotations for evaluating multi-camera people-trackers on large video datasets. Most of the annotation data is computed automatically, by estimating a consensus tracking result from multiple existing trackers and people detectors and classifying it as either reliable or not. A small subset of the data, composed of tracks with insufficient reliability is verified by a human using a simple binary decision task, a process faster than marking the correct person position. The proposed framework is generic and can handle additional trackers. We present results on a dataset of approximately 6 hours captured by 4 cameras, featuring a person in a holiday flat, performing activities such as walking, cooking, eating, cleaning, and watching TV. When aiming for a tracking accuracy of 60cm, 80% of all video frames are automatically annotated. The annotations for the remaining 20% of the frames were added after human verification of an automatically selected subset of data. This involved about 2.4 hours of manual labour. According to a subsequent comprehensive visual inspection to judge the annotation procedure, we found 99% of the automatically annotated frames to be correct. We provide guidelines on how to apply the proposed methodology to new datasets. We also provide an exploratory study for the multi-target case, applied on existing and new benchmark video sequences.
Jorge Oswaldo Niño Castañeda, Andrés Frias-Velázquez, Nyan Bo Bo, Maarten Slembrouck, Junzhi Guan, Glen Debard, Bart Vanrumste, Tinne Tuytelaars, Wilfried Philips
IEEE Trans. Image Process.8
2016 Example-Based Sketch Segmentation and Labeling Using CRFs
abstract
We introduce a new approach for segmentation and label transfer in sketches that substantially improves the state of the art. We build on successful techniques to find how likely each segment is to belong to a label, and use a Conditional Random Field to find the most probable global configuration. Our method is trained fully on the sketch domain, such that it can handle abstract sketches that are very far from 3D meshes. It also requires a small quantity of annotated data, which makes it easily adaptable to new datasets. The testing phase is completely automatic, and our performance is comparable to state-of-the-art methods that require manual tuning and a considerable amount of previous annotation [Huang et al. 2014].
Rosália G. Schneider, Tinne Tuytelaars
ACM Trans. Graph.2
2015 Spatio-Temporal Object Recognition
Roeland De Geest, Francis Deboeverie, Wilfried Philips, Tinne Tuytelaars
ACIVS4
2015 Subspace Alignment Based Domain Adaptation for RCNN Detector
abstract
In this paper, we propose subspace alignment based domain adaptation of the state of the art RCNN based object detector. The aim is to be able to achieve high quality object detection in novel, real world target scenarios without requiring labels from the target domain. While, unsupervised domain adaptation has been studied in the case of object classification, for object detection it has been relatively unexplored. In subspace based domain adaptation for objects, we need access to source and target subspaces for the bounding box features. The absence of supervision (labels and bounding boxes are absent) makes the task challenging. In this paper, we show that we can still adapt sub- spaces that are localized to the object by obtaining detections from the RCNN detector trained on source and applied on target. Then we form localized subspaces from the detections and show that subspace alignment based adaptation between these subspaces yields improved object detection. This evaluation is done by considering challenging real world datasets of PASCAL VOC as source and validation set of Microsoft COCO dataset as target for various categories.
Anant Raj, Vinay P. Namboodiri, Tinne Tuytelaars
BMVC3
2015 Weakly supervised object detection with convex clustering
abstract
Weakly supervised object detection, is a challenging task, where the training procedure involves learning at the same time both, the model appearance and the object location in each image. The classical approach to solve this problem is to consider the location of the object of interest in each image as a latent variable and minimize the loss generated by such latent variable during learning. However, as learning appearance and localization are two interconnected tasks, the optimization is not convex and the procedure can easily get stuck in a poor local minimum, i.e. the algorithm “misses” the object in some images. In this paper, we help the optimization to get close to the global minimum by enforcing a “soft” similarity between each possible location in the image and a reduced set of “exemplars”, or clusters, learned with a convex formulation in the training images. The help is effective because it comes from a different and smooth source of information that is not directly connected with the main task. Results show that our method improves a strong baseline based on convolutional neural network features by more than 4 points without any additional features or extra computation at testing time but only adding a small increment of the training time due to the convex clustering.
Hakan Bilen, Marco Pedersoli, Tinne Tuytelaars
CVPR3
2015 Modeling video evolution for action recognition
abstract
In this paper we present a method to capture video-wide temporal information for action recognition. We postulate that a function capable of ordering the frames of a video temporally (based on the appearance) captures well the evolution of the appearance within the video. We learn such ranking functions per video via a ranking machine and use the parameters of these as a new video representation. The proposed method is easy to interpret and implement, fast to compute and effective in recognizing a wide variety of actions. We perform a large number of evaluations on datasets for generic action recognition (Hollywood2 and HMDB51), fine-grained actions (MPII- cooking activities) and gestures (Chalearn). Results show that the proposed method brings an absolute improvement of 7–10%, while being compatible with and complementary to further improvements in appearance and local motion based methods.
Basura Fernando, Efstratios Gavves, José Oramas M., Amir Ghodrati, Tinne Tuytelaars
CVPR5
2015 Dataset fingerprints: Exploring image collections through data mining
abstract
As the amount of visual data increases, so does the need for summarization tools that can be used to explore large image collections and to quickly get familiar with their content. In this paper, we propose dataset fingerprints, a new and powerful method based on data mining that extracts meaningful patterns from a set of images. The discovered patterns are compositions of discriminative mid-level features that co-occur in several images. Compared to earlier work, ours stands out because i) it's fully unsupervised, ii) discovered patterns cover large parts of the images, often corresponding to full objects or meaningful parts thereof, and iii) different patterns are connected based on co-occurrence, allowing a user to “browse” the images from one pattern to the next and to group patterns in a semantically meaningful manner.
Konstantinos Rematas, Basura Fernando, Frank Dellaert, Tinne Tuytelaars
CVPR4
2015 Continuous Pose Estimation with a Spatial Ensemble of Fisher Regressors
abstract
In this paper, we treat the problem of continuous pose estimation for object categories as a regression problem on the basis of only 2D training information. While regression is a natural framework for continuous problems, regression methods so far achieved inferior results with respect to 3D-based and 2D-based classification-and-refinement approaches. This may be attributed to their weakness to high intra-class variability as well as to noisy matching procedures and lack of geometrical constraints. We propose to apply regression to Fisher-encoded vectors computed from large cells by learning an array of Fisher regressors. Fisher encoding makes our algorithm flexible to variations in class appearance, while the array structure permits to indirectly introduce spatial context information in the approach. We formulate our problem as a MAP inference problem, where the likelihood function is composed of a generative term based on the prediction error generated by the ensemble of Fisher regressors as well as a discriminative term based on SVM classifiers. We test our algorithm on three publicly available datasets that envisage several difficulties, such as high intra-class variability, truncations, occlusions, and motion blur, obtaining state-of-the-art results.
Michele Fenzi, Laura Leal-Taixé, Jörn Ostermann, Tinne Tuytelaars
ICCV4
2015 Learning to Rank Based on Subsequences
abstract
We present a supervised learning to rank algorithm that effectively orders images by exploiting the structure in image sequences. Most often in the supervised learning to rank literature, ranking is approached either by analysing pairs of images or by optimizing a list-wise surrogate loss function on full sequences. In this work we propose MidRank, which learns from moderately sized sub-sequences instead. These sub-sequences contain useful structural ranking information that leads to better learnability during training and better generalization during testing. By exploiting sub-sequences, the proposed MidRank improves ranking accuracy considerably on an extensive array of image ranking applications and datasets.
Basura Fernando, Efstratios Gavves, Damien Muselet, Tinne Tuytelaars
ICCV4
2015 Active Transfer Learning with Zero-Shot Priors: Reusing Past Datasets for Future Tasks
abstract
How can we reuse existing knowledge, in the form of available datasets, when solving a new and apparently unrelated target task from a set of unlabeled data? In this work we make a first contribution to answer this question in the context of image classification. We frame this quest as an active learning problem and use zero-shot classifiers to guide the learning process by linking the new task to the the existing classifiers. By revisiting the dual formulation of adaptive SVM, we reveal two basic conditions to choose greedily only the most relevant samples to be annotated. On this basis we propose an effective active learning algorithm which learns the best possible target classification model with minimum human labeling effort. Extensive experiments on two challenging datasets show the value of our approach compared to the state-of-the-art active learning methodologies, as well as its potential to reuse past datasets with minimal effort for future tasks.
Efstratios Gavves, Thomas Mensink, Tatiana Tommasi, Cees Snoek, Tinne Tuytelaars
ICCV5
2015 DeepProposal: Hunting Objects by Cascading Deep Convolutional Layers
abstract
In this paper we evaluate the quality of the activation layers of a convolutional neural network (CNN) for the generation of object proposals. We generate hypotheses in a sliding-window fashion over different activation layers and show that the final convolutional layers can find the object of interest with high recall but poor localization due to the coarseness of the feature maps. Instead, the first layers of the network can better localize the object of interest but with a reduced recall. Based on this observation we design a method for proposing object locations that is based on CNN features and that combines the best of both worlds. We build an inverse cascade that, going from the final to the initial convolutional layers of the CNN, selects the most promising object locations and refines their boxes in a coarse-to-fine manner. The method is efficient, because i) it uses the same features extracted for detection, ii) it aggregates features using integral images, and iii) it avoids a dense evaluation of the proposals due to the inverse coarse-to-fine cascade. The method is also accurate, it outperforms most of the previously proposed object proposals approaches and when plugged into a CNN-based detector produces state-of-the-art detection performance.
Amir Ghodrati, Ali Diba, Marco Pedersoli, Tinne Tuytelaars, Luc Van Gool
ICCV4
2015 Guiding the Long-Short Term Memory Model for Image Caption Generation
abstract
In this work we focus on the problem of image caption generation. We propose an extension of the long short term memory (LSTM) model, which we coin gLSTM for short. In particular, we add semantic information extracted from the image as extra input to each unit of the LSTM block, with the aim of guiding the model towards solutions that are more tightly coupled to the image content. Additionally, we explore different length normalization strategies for beam search to avoid bias towards short sentences. On various benchmark datasets such as Flickr8K, Flickr30K and MS COCO, we obtain results that are on par with or better than the current state-of-the-art.
Xu Jia 0012, Efstratios Gavves, Basura Fernando, Tinne Tuytelaars
ICCV4
2015 Learning Where to Position Parts in 3D
abstract
A common issue in deformable object detection is finding a good way to position the parts. This issue is even more outspoken when considering detection and pose estimation for 3D objects, where parts should be placed in a three-dimensional space. Some methods extract the 3D shape of the object from 3D CAD models. This limits their applicability to categories for which such models are available. Others represent the object with a predefined and simple shape (e.g. a cuboid). This extends the applicability of the model, but in many cases the pre-defined shape is too simple to properly represent the object in 3D. In this paper we propose a new method for the detection and pose estimation of 3D objects, that does not use any 3D CAD model or other 3D information. Starting from a simple and general 3D shape, we learn in a weakly supervised manner the 3D part locations that best fit the training data. As this method builds on a iterative estimation of the part locations, we introduce several speedups to make the method fast enough for practical experiments. We evaluate our model for the detection and pose estimation of faces and cars. Our method obtains results comparable with the state of the art, it is faster than most of the other approaches and does not need any additional 3D information.
Marco Pedersoli, Tinne Tuytelaars
ICCV2
2015 Towards sign language recognition based on body parts relations
abstract
Over the years, hand gesture recognition has been mostly addressed considering hand trajectories in isolation. However, in most sign languages, hand gestures are defined on a particular context (body region). We propose a pipeline which models hand movements in the context of other parts of the body captured in the 3D space using the Kinect sensor. In addition, we perform sign recognition based on the different hand postures that occur during a sign. Our experiments show that considering different body parts brings improved performance when compared with methods which only consider global hand trajectories. Finally, we demonstrate that the combination of hand postures features with hand gestures features helps to improve the prediction of a given sign.
Marc Martínez-Camarena, José Oramas M., Tinne Tuytelaars
ICIP3
2015 Who's Speaking?: Audio-Supervised Classification of Active Speakers in Video
abstract
Active speakers have traditionally been identified in video by detecting their moving lips. This paper demonstrates the same using spatio-temporal features that aim to capture other cues: movement of the head, upper body and hands of active speakers. Speaker directional information, obtained using sound source localization from a microphone array is used to supervise the training of these video features.
Punarjay Chakravarty, Sayeh Mirzaei, Tinne Tuytelaars, Hugo Van hamme
ICMI3
2015 Swap Retrieval: Retrieving Images of Cats When the Query Shows a Dog
abstract
Query-by-example remains popular in image retrieval because it can exploit contextual information encoded in the image, that is difficult to express in a traditional textual query. Textual queries, on the other hand, give more flexibility in that it's easy to reformulate and refine a text query based on initial results.
Amir Ghodrati, Xu Jia 0012, Marco Pedersoli, Tinne Tuytelaars
ICMR4
2015 Location recognition over large time lags
Basura Fernando, Tatiana Tommasi, Tinne Tuytelaars
Comput. Vis. Image Underst.3
2015 Local Alignments for Fine-Grained Categorization
Efstratios Gavves, Basura Fernando, Cees Snoek, Arnold W. M. Smeulders, Tinne Tuytelaars
Int. J. Comput. Vis.5
2015 An Elastic Deformation Field Model for Object Detection and Tracking
Marco Pedersoli, Radu Timofte, Tinne Tuytelaars, Luc Van Gool
Int. J. Comput. Vis.3
2015 Joint cross-domain classification and subspace learning for unsupervised adaptation
Basura Fernando, Tatiana Tommasi, Tinne Tuytelaars
Pattern Recognit. Lett.3
2014 A Scalable 3D HOG Model for Fast Object Detection and Viewpoint Estimation
abstract
In this paper we present a scalable way to learn and detect objects using a 3D representation based on HOG patches placed on a 3D cuboid. The model consists of a single 3D representation that is shared among views. Similarly to the work of Fidler et al. [5], at detection time this representation is projected on the image plane over the desired viewpoints. However, whereas in [5] the projection is done at image-level and therefore the computational cost is linear in the number of views, in our model every view is approximated at feature level as a linear combination of the pre-computed fron to-parallel views. As a result, once the fron to-parallel views have been computed, the cost of computing new views is almost negligible. This allows the model to be evaluated on many more viewpoints. In the experimental results we show that the proposed model has a comparable detection and pose estimation performance to standard multiview HOG detectors, but it is faster, it scales very well with the number of views and can better generalize to unseen views. Finally, we also show that with a procedure similar to label propagation it is possible to train the model even without using pose annotations at training time.
Marco Pedersoli, Tinne Tuytelaars
3DV2
2014 Weakly Supervised Detection with Posterior Regularization
Hakan Bilen, Marco Pedersoli, Tinne Tuytelaars
BMVC3
2014 Is 2D Information Enough For Viewpoint Estimation?
Amir Ghodrati, Marco Pedersoli, Tinne Tuytelaars
BMVC3
2014 Scene-driven Cues for Viewpoint Classification for Elongated Object Classes
José Oramas M., Tinne Tuytelaars
BMVC2
2014 All together now: Simultaneous Detection and Continuous Pose Estimation using a Hough Forest with Probabilistic Locally Enhanced Voting
Carolina Redondo-Cabrera, Roberto Javier López-Sastre, Tinne Tuytelaars
BMVC3
2014 Object Classification with Adaptable Regions
abstract
In classification of objects substantial work has gone into improving the low level representation of an image by considering various aspects such as different features, a number of feature pooling and coding techniques and considering different kernels. Unlike these works, in this paper, we propose to enhance the semantic representation of an image. We aim to learn the most important visual components of an image and how they interact in order to classify the objects correctly. To achieve our objective, we propose a new latent SVM model for category level object classification. Starting from image-level annotations, we jointly learn the object class and its context in terms of spatial location (where) and appearance (what). Furthermore, to regularize the complexity of the model we learn the spatial and co-occurrence relations between adjacent regions, such that unlikely configurations are penalized. Experimental results demonstrate that the proposed method can consistently enhance results on the challenging Pascal VOC dataset in terms of classification and weakly supervised detection. We also show how semantic representation can be exploited for finding similar content.
Hakan Bilen, Marco Pedersoli, Vinay P. Namboodiri, Tinne Tuytelaars, Luc Van Gool
CVPR4
2014 Using a Deformation Field Model for Localizing Faces and Facial Points under Weak Supervision
abstract
Face detection and facial points localization are interconnected tasks. Recently it has been shown that solving these two tasks jointly with a mixture of trees of parts (MTP) leads to state-of-the-art results. However, MTP, as most other methods for facial point localization proposed so far, requires a complete annotation of the training data at facial point level. This is used to predefine the structure of the trees and to place the parts correctly. In this work we extend the mixtures from trees to more general loopy graphs. In this way we can learn in a weakly supervised manner (using only the face location and orientation) a powerful deformable detector that implicitly aligns its parts to the detected face in the image. By attaching some reference points to the correct parts of our detector we can then localize the facial points. In terms of detection our method clearly outperforms the state-of-the-art, even if competing with methods that use facial point annotations during training. Additionally, without any facial point annotation at the level of individual training images, our method can localize facial points with an accuracy similar to fully supervised approaches.
Marco Pedersoli, Radu Timofte, Tinne Tuytelaars, Luc Van Gool
CVPR3
2014 Image-Based Synthesis and Re-synthesis of Viewpoints Guided by 3D Models
abstract
We propose a technique to use the structural informa- tion extracted from a set of 3D models of an object class to improve novel-view synthesis for images showing unknown instances of this class. These novel views can be used to "amplify" training image collections that typically contain only a low number of views or lack certain classes of views entirely (e. g. top views). We extract the correlation of position, normal, re- flectance and appearance from computer-generated images of a few exemplars and use this information to infer new appearance for new instances. We show that our approach can improve performance of state-of-the-art detectors using real-world training data. Additional applications include guided versions of inpainting, 2D-to-3D conversion, super- resolution and non-local smoothing.
Konstantinos Rematas, Tobias Ritschel 0001, Mario Fritz, Tinne Tuytelaars
CVPR4
2014 Color features for dating historical color images
abstract
Estimating the age of historical photographs is a challenging task for human beings. Only recently this task has been addressed in computational image analysis perspective. The characteristics of the device used to acquire each photograph are discriminative features for this task. We aim at extracting such characteristics from a historical color photographs. The acquisition device mainly effects two properties of the colors: the distribution of their derivatives and the angles drawn by three consecutive pixels in the RGB space. We propose two color features that take advantage of these observations. We show that these two color descriptors (namely color derivatives and color angles) attain the state-of-the-art in the context of image dating.
Basura Fernando, Damien Muselet, Rahat Khan, Tinne Tuytelaars
ICIP4
2014 Dense interest features for video processing
abstract
We propose two novel feature detection methods for action recognition, based on the dense interest points described by Tuytelaars [1]. The first one is an extension of dense interest points to three dimensions. In the second one, trajectories are constructed starting from dense interest points. We present an analysis of the properties of these methods and conclude that both give higher classification accuracies than dense sampling when less features are used.
Roeland De Geest, Tinne Tuytelaars
ICIP2
2014 The Combinator: Optimal Combination of Multiple Pedestrian Detectors
abstract
In recent years, the accuracy of pedestrian detectors significantly improved. Currently, state-of-the-art pedestrian detectors achieve high accuracy results on challenging datasets. As opposed to refining a single detector, in this paper we propose a different approach to further increase the detection accuracy: combining multiple pedestrian detectors. The most straight-forward way to combine pedestrian detectors would be a naive AND or OR combination. Here, we present a novel generic combination framework in which we exploit specific information from each pedestrian detector to determine the optimal combination parameters. Our main motivation for this approach is based on the fact that several pedestrian detection approaches are based on very different techniques (e.g. a different feature pool), and thus an efficient combination should yield higher accuracy results. Indeed, such a combination is far more powerful, and our experiments indicate that specific (that is, cleverly chosen) combinations outperform existing state-of-the-art pedestrian detection results.
Floris De Smedt, Kristof Van Beeck, Tinne Tuytelaars, Toon Goedemé
ICPR3
2014 Learning Like a Toddler: Watching Television Series to Learn Vocabulary from Images and Audio
abstract
This paper presents the initial findings of our efforts to build an unsupervised multimodal vocabulary learning scheme in a realistic scenario. For this purpose, a new multimodal dataset, called Musti3D, has been created. The Musti3D database contains episodes from an animation series for toddlers. Annotated with audiovisual information, this database is used for the investigation of a non-negative matrix factorization (NMF)-based audiovisual learning technique. The performance of the technique, i.e. correctly matching the audio and visual representations of the objects, has been evaluated by gradually reducing the level of supervision starting from the ground truth transcriptions. Moreover, we have performed experiments using different visual representations and time spans for combining the audiovisual information. The preliminary results show the feasibility of the proposed audiovisual learning framework.
Emre Yilmaz 0001, Konstantinos Rematas, Tinne Tuytelaars, Hugo Van hamme
ACM Multimedia3
2014 Coupling video segmentation and action recognition
abstract
Recently a lot of progress has been made in the field of video segmentation. The question then arises whether and how these results can be exploited for this other video processing challenge, action recognition. In this paper we show that a good segmentation is actually very important for recognition. We propose and evaluate several ways to integrate and combine the two tasks: i) recognition using a standard, bottom-up segmentation, ii) using a top-down segmentation geared towards actions, iii) using a segmentation based on inter-video similarities (co-segmentation), and iv) tight integration of recognition and segmentation via iterative learning. Our results clearly show that, on the one hand, the two tasks are interdependent and therefore an iterative optimization of the two makes sense and gives better results. On the other hand, comparable results can also be obtained with two separate steps but mapping the feature-space with a non-linear kernel.
Amir Ghodrati, Marco Pedersoli, Tinne Tuytelaars
WACV3
2014 Towards cautious collective inference for object verification
abstract
It is by now generally accepted that reasoning about the relationships between objects (and object hypotheses) can improve the accuracy of object detection methods. Relations between objects allow to reject inconsistent hypotheses and reduce the uncertainty of the initial hypotheses. However, most methods to date reason about object relations in a relatively crude way. In this paper we propose an alternative using cautious inference. Building on ideas from Collective Classification, we favor the most confident hypotheses as sources of contextual information and give higher relevance to the object relations observed during training. Additionally, we propose to cluster the pairwise relations into relationships. Our experiments on part of the KITTI data benchmark and the MIT StreetScenes dataset show that both steps improve the performance of relational classifiers.
José Oramas M., Luc De Raedt, Tinne Tuytelaars
WACV3
2014 Action in chains: A chains model for action localization and classification
abstract
In this paper we present a method for action classification in videos using trajectory features. The novelty of our approach is in formulating the problem of simultaneous detection and localization as a probabilistic chains model. In our formulation, chains are sets of regions in the video that are connected based on their joint probabilities. We describe our approach for connecting subvolumes in the video into chains, and using them as spatio-temporal detectors for actions. Our approach allows the detection and localization of multiple actions occurring simultaneously or at different locations in a single video. We test the performance of our method on two challenging action recognition datasets, and compare to state of the art methods.
Gilad Sharir, Tinne Tuytelaars
WACV2
2014 Boosting masked dominant orientation templates for efficient object detection
Reyes Rios-Cabrera, Tinne Tuytelaars
Comput. Vis. Image Underst.2
2014 Mining Mid-level Features for Image Classification
Basura Fernando, Élisa Fromont, Tinne Tuytelaars
Int. J. Comput. Vis.3
2014 There are plenty of places like home: Using relational representations in hierarchies for distance-based image understanding
Laura Antanas, Martijn van Otterlo, José Oramas M., Tinne Tuytelaars, Luc De Raedt
Neurocomputing4
2014 Sketch classification and classification-driven analysis using Fisher vectors
abstract
We introduce an approach for sketch classification based on Fisher vectors that significantly outperforms existing techniques. For the TU-Berlin sketch benchmark [Eitz et al. 2012a], our recognition rate is close to human performance on the same task. Motivated by these results, we propose a different benchmark for the evaluation of sketch classification algorithms. Our key idea is that the relevant aspect when recognizing a sketch is not the intention of the person who made the drawing, but the information that was effectively expressed. We modify the original benchmark to capture this concept more precisely and, as such, to provide a more adequate tool for the evaluation of sketch classification techniques. Finally, we perform a classification-driven analysis which is able to recover semantic aspects of the individual sketches, such as the quality of the drawing and the importance of each part of the sketch for the recognition.
Rosália G. Schneider, Tinne Tuytelaars
ACM Trans. Graph.2
2013 Seeking the Strongest Rigid Detector
abstract
The current state of the art solutions for object detection describe each class by a set of models trained on discovered sub-classes (so called "components"), with each model itself composed of collections of interrelated parts (deformable models). These detectors build upon the now classic Histogram of Oriented Gradients+linear SVM combo. In this paper we revisit some of the core assumptions in HOG+SVM and show that by properly designing the feature pooling, feature selection, preprocessing, and training methods, it is possible to reach top quality, at least for pedestrian detections, using a single rigid component. Abstract We provide experiments for a large design space, that give insights into the design of classifiers, as well as relevant information for practitioners. Our best detector is fully feed-forward, has a single unified architecture, uses only histograms of oriented gradients and colour information in monocular static images, and improves over 23 other methods on the INRIA, ETH and Caltech-USA datasets, reducing the average miss-rate over HOG+SVM by more than 30%.
Rodrigo Benenson, Markus Mathias, Tinne Tuytelaars, Luc Van Gool
CVPR3
2013 Allocentric Pose Estimation
abstract
The task of object pose estimation has been a challenge since the early days of computer vision. To estimate the pose (or viewpoint) of an object, people have mostly looked at object intrinsic features, such as shape or appearance. Surprisingly, informative features provided by other, external elements in the scene, have so far mostly been ignored. At the same time, contextual cues have been shown to be of great benefit for related tasks such as object detection or action recognition. In this paper, we explore how information from other objects in the scene can be exploited for pose estimation. In particular, we look at object configurations. We show that, starting from noisy object detections and pose estimates, exploiting the estimated pose and location of other objects in the scene can help to estimate the objects' poses more accurately. We explore both a camera-centered as well as an object-centered representation for relations. Experiments on the challenging KITTI dataset show that object configurations can indeed be used as a complementary cue to appearance-based pose estimation. In addition, object-centered relational representations can also assist object detection.
José Oramas M., Luc De Raedt, Tinne Tuytelaars
ICCV3
2013 Unsupervised Visual Domain Adaptation Using Subspace Alignment
abstract
In this paper, we introduce a new domain adaptation (DA) algorithm where the source and target domains are represented by subspaces described by eigenvectors. In this context, our method seeks a domain adaptation solution by learning a mapping function which aligns the source subspace with the target one. We show that the solution of the corresponding optimization problem can be obtained in a simple closed form, leading to an extremely fast algorithm. We use a theoretical result to tune the unique hyper parameter corresponding to the size of the subspaces. We run our method on various datasets and show that, despite its intrinsic simplicity, it outperforms state of the art DA methods.
Basura Fernando, Amaury Habrard, Marc Sebban, Tinne Tuytelaars
ICCV4
2013 Mining Multiple Queries for Image Retrieval: On-the-Fly Learning of an Object-Specific Mid-level Representation
abstract
In this paper we present a new method for object retrieval starting from multiple query images. The use of multiple queries allows for a more expressive formulation of the query object including, e.g., different viewpoints and/or viewing conditions. This, in turn, leads to more diverse and more accurate retrieval results. When no query images are available to the user, they can easily be retrieved from the internet using a standard image search engine. In particular, we propose a new method based on pattern mining. Using the minimal description length principle, we derive the most suitable set of patterns to describe the query object, with patterns corresponding to local feature configurations. This results in a powerful object-specific mid-level image representation. The archive can then be searched efficiently for similar images based on this representation, using a combination of two inverted file systems. Since the patterns already encode local spatial information, good results on several standard image retrieval datasets are obtained even without costly re-ranking based on geometric verification.
Basura Fernando, Tinne Tuytelaars
ICCV2
2013 Fine-Grained Categorization by Alignments
abstract
The aim of this paper is fine-grained categorization without human interaction. Different from prior work, which relies on detectors for specific object parts, we propose to localize distinctive details by roughly aligning the objects using just the overall shape, since implicit to fine-grained categorization is the existence of a super-class shape shared among all classes. The alignments are then used to transfer part annotations from training images to test images (supervised alignment), or to blindly yet consistently segment the object in a number of regions (unsupervised alignment). We furthermore argue that in the distinction of fine grained sub-categories, classification-oriented encodings like Fisher vectors are better suited for describing localized information than popular matching oriented features like HOG. We evaluate the method on the CU-2011 Birds and Stanford Dogs fine-grained datasets, outperforming the state-of-the-art.
Efstratios Gavves, Basura Fernando, Cees Snoek, Arnold W. M. Smeulders, Tinne Tuytelaars
ICCV5
2013 Discriminatively Trained Templates for 3D Object Detection: A Real Time Scalable Approach
abstract
In this paper we propose a new method for detecting multiple specific 3D objects in real time. We start from the template-based approach based on the LINE2D/LINEMOD representation introduced recently by Hinterstoisser et al., yet extend it in two ways. First, we propose to learn the templates in a discriminative fashion. We show that this can be done online during the collection of the example images, in just a few milliseconds, and has a big impact on the accuracy of the detector. Second, we propose a scheme based on cascades that speeds up detection. Since detection of an object is fast, new objects can be added with very low cost, making our approach scale well. In our experiments, we easily handle 10-30 3D objects at frame rates above 10fps using a single CPU core. We outperform the state-of-the-art both in terms of speed as well as in terms of accuracy, as validated on 3 different datasets. This holds both when using monocular color images (with LINE2D) and when using RGBD images (with LINEMOD). Moreover, we propose a challenging new dataset made of 12 objects, for future competing methods on monocular color images.
Reyes Rios-Cabrera, Tinne Tuytelaars
ICCV2
2013 Multi RGB-D camera setup for generating large 3D point clouds
abstract
The advent of inexpensive RGB-D cameras brings new opportunities to capture a 3D environment. This paper presents a method to create a modular setup for generating a large 3D point cloud, with attention to the study of interference, the influence of a USB extension cable, and the calibration procedure. The study of interference includes the influence of the distance between the cameras, the orientation of the cameras, and the illumination. Furthermore, this paper proposes a number of evaluation metrics for similar setups.
Wim Lemkens, Koen Buys, Peter Slaets, Tinne Tuytelaars, Joris De Schutter
IROS5
2013 A relational kernel-based approach to scene classification
abstract
Real-world scenes involve many objects that interact with each other in complex semantic patterns. For example, a bar scene can be naturally described as having a variable number of chairs of similar size, close to each other and aligned horizontally. This high-level interpretation of a scene relies on semantically meaningful entities and is most generally described using relational representations or (hyper-) graphs. Popular in early work on syntactic and structural pattern recognition, relational representations are rarely used in computer vision due to their pure symbolic nature. Yet, today recent successes in combining them with statistical learning principles motivates us to reinvestigate their use. In this paper we show that relational techniques can also improve scene classification. More specifically, we employ a new relational language for learning with kernels, called kLog. With this language we define higher-order spatial relations among semantic objects. When applied to a particular image, they characterize a particular object arrangement and provide discriminative cues for the scene category. The kernel allows us to tractably learn from such complex features. Thus, our contribution is a principled and interpretable approach to learn from symbolic relations how to classify scenes in a statistical framework. We obtain results comparable to state-of-the-art methods on 15 Scenes and a subset of the MIT indoor dataset.
Laura Antanas, McElory Hoffmann, Paolo Frasconi, Tinne Tuytelaars, Luc De Raedt
WACV4
2013 Naming persons in video: Using the weak supervision of textual stories
Phi The Pham, Koen Deschacht, Tinne Tuytelaars, Marie-Francine Moens
J. Vis. Commun. Image Represent.3
2013 Finding a needle in a haystack: an interactive video archive explorer for professional video searchers
Mieke Haesen, Jan Meskens, Kris Luyten, Karin Coninx, Jan Hendrik Becker, Tinne Tuytelaars, Gert-Jan Poulisse, Phi The Pham, Marie-Francine Moens
Multim. Tools Appl.6
2012 The Pooled NBNN Kernel: Beyond Image-to-Class and Image-to-Image
Konstantinos Rematas, Mario Fritz, Tinne Tuytelaars
ACCV (1)3
2012 Naive Bayes Image Classification: Beyond Nearest Neighbors
Radu Timofte, Tinne Tuytelaars, Luc Van Gool
ACCV (1)2
2012 Effective Use of Frequent Itemset Mining for Image Classification
Basura Fernando, Élisa Fromont, Tinne Tuytelaars
ECCV (1)3
2012 A Warping Window Approach to Real-time Vision-based Pedestrian Detection in a Truck's Blind Spot Zone
Kristof Van Beeck, Toon Goedemé, Tinne Tuytelaars
ICINCO (2)3
2012 Integrating video and accelerometer signals for nocturnal epileptic seizure detection
abstract
Epileptic seizure detection is traditionally done using video/electroencephalogram (EEG) monitoring, which is not applicable in a home situation. In recent years, attempts have been made to detect the seizures using other modalities. In this paper we investigate if a combined usage of accelerometers attached to the limbs and video data would increase the performance compared to a single modality approach. Therefore, we used two existing approaches for seizure detection in accelerometers and video and combined them using a linear discriminant analysis (LDA) classifier. The results for a combined detection have a better positive predictive value (PPV) of 95.00% compared to the single modality detection and reached a sensitivity of 83.33%.
Kris Cuppens, Chih-Wei Chen, Kevin Bing-Yung Wong, Anouk Van de Vel, Lieven Lagae, Berten Ceulemans, Tinne Tuytelaars, Sabine Van Huffel, Bart Vanrumste, Hamid K. Aghajan
ICMI7
2012 A Relational Distance-based Framework for Hierarchical Image Understanding
Laura Antanas, Martijn van Otterlo, José Oramas M., Tinne Tuytelaars, Luc De Raedt
ICPRAM (2)4
2012 Efficient multi-camera vehicle detection, tracking, and identification in a tunnel surveillance application
Reyes Rios-Cabrera, Tinne Tuytelaars, Luc Van Gool
Comput. Vis. Image Underst.2
2011 Automatic Occlusion Removal from Facades for 3D Urban Reconstruction
Chris Engels, David Tingdahl, Mathias Vercruysse, Tinne Tuytelaars, Hichem Sahli, Luc Van Gool
ACIVS4
2011 Efficient multi-camera detection, tracking, and identification using a shared set of haar-features
abstract
This paper presents an integrated solution for the problem of detecting, tracking and identifying vehicles in a tunnel surveillance application, taking into account practical constraints including realtime operation, poor imaging conditions, and a decentralized architecture. Vehicles are followed through the tunnel by a network of non-overlapping cameras. They are detected and tracked in each camera and then identified, i.e. matched to any of the vehicles detected in the previous camera(s). To limit the computational load, we propose to reuse the same set of Haar-features for each of these steps. For the detection, we use an Adaboost cascade. Here we introduce a composite confidence score, integrating information from all stage of the cascades. A subset of the features used for detection is then selected, optimizing for the identification problem. This results in a compact binary `vehicle fingerprint', requiring very limited bandwidth. Finally, we show that the same set of features can also be used for tracking. This haar features based `tracking-by-identification' yields surprisingly good results on standard datasets, without the need to update the model online.
Reyes Rios-Cabrera, Tinne Tuytelaars, Luc Van Gool
CVPR2
2011 The NBNN kernel
abstract
Naive Bayes Nearest Neighbor (NBNN) has recently been proposed as a powerful, non-parametric approach for object classification, that manages to achieve remarkably good results thanks to the avoidance of a vector quantization step and the use of image-to-class comparisons, yielding good generalization. In this paper, we introduce a kernelized version of NBNN. This way, we can learn the classifier in a discriminative setting. Moreover, it then becomes straightforward to combine it with other kernels. In particular, we show that our NBNN kernel is complementary to standard bag-of-features based kernels, focussing on local generalization as opposed to global image composition. By combining them, we achieve state-of-the-art results on Caltech101 and 15 Scenes datasets. As a side contribution, we also investigate how to speed up the NBNN computations.
Tinne Tuytelaars, Mario Fritz, Kate Saenko, Trevor Darrell
ICCV1
2011 Towards a more discriminative and semantic visual vocabulary
Roberto Javier López-Sastre, Tinne Tuytelaars, Francisco Javier Acevedo-Rodríguez, Saturnino Maldonado-Bascón
Comput. Vis. Image Underst.2
2010 Automatic annotation of unique locations from video and text
abstract
Given a video and associated text, we propose an automatic annotation scheme in which we employ a latent topic model to generate topic distributions from weighted text and then modify these distributions based on visual similarity. We apply this scheme to location annotation of a television series for which transcripts are available. The topic distributions allow us to avoid explicit classification, which is useful in cases where the exact number of locations is unknown. Moreover, many locations are unique to a single episode, making it impossible to obtain representative training data for a supervised approach. Our method first segments the episode into scenes by fusing cues from both images and text. We then assign location-oriented weights to the text and generate topic distributions for each scene using Latent Dirichlet Allocation. Finally, we update the topic distributions using the distributions of visually similar scenes. We formulate our visual similarity between scenes as an Earth Mover's Distance problem. We quantitatively validate our multi-modal approach to segmentation and qualitatively evaluate the resulting location annotations. Our results demonstrate that we are able to generate accurate annotations, even for locations only seen in a single episode. © 2010. The copyright of this document resides with its authors.
Chris Engels, Koen Deschacht, Jan Hendrik Becker, Tinne Tuytelaars, Marie-Francine Moens, Luc Van Gool
BMVC4
2010 Dense interest points
abstract
Local features or image patches have become a standard tool in computer vision, with numerous application domains. Roughly speaking, two different types of patch-based image representations can be distinguished: interest points, such as corners or blobs, whose position, scale and shape are computed by a feature detector algorithm, and dense sampling, where patches of fixed size and shape are placed on a regular grid (possibly repeated over multiple scales). Interest points focus on `interesting' locations in the image and include various degrees of viewpoint and illumination invariance, resulting in better repeatability scores. Dense sampling, on the other hand, gives a better coverage of the image, a constant amount of features per image area, and simple spatial relations between features. In this paper, we propose a hybrid scheme, which we call dense interest points, where we start from densely sampled patches yet optimize their position and scale parameters locally. We investigate whether doing so it is possible to get the best of both worlds.
Tinne Tuytelaars
CVPR1
2010 Naming persons in news video with label propagation
abstract
Labeling persons appearing in video frames with names detected from the video transcript helps improving the video content identification and search task. We develop a face naming method that learns from labeled and unlabeled examples using iterative label propagation in a graph of connected faces or name-face pairs. The advantage of this method is that it can use very few labeled data points and incorporate the unlabeled data points during the learning process. Anchor detection and metric learning for face classification techniques are incorporated into the label propagation process to help boosting the face naming performance. On BBC News videos, the label propagation algorithm yields better results than a Support Vector Machine classifier trained on the same labeled data.
Phi The Pham, Marie-Francine Moens, Tinne Tuytelaars
ICME3
2010 Not Far Away from Home: A Relational Distance-Based Approach to Understanding Images of Houses
Laura Antanas, Martijn van Otterlo, José Oramas M., Tinne Tuytelaars, Luc De Raedt
ILP4
2010 Unsupervised Object Discovery: A Comparison
abstract
The goal of this paper is to evaluate and compare models and methods for learning to recognize basic entities in images in an unsupervised setting. In other words, we want to discover the objects present in the images by analyzing unlabeled data and searching for re-occurring patterns. We experiment with various baseline methods, methods based on latent variable models, as well as spectral clustering methods. The results are presented and compared both on subsets of Caltech256 and MSRC2, data sets that are larger and more challenging and that include more object classes than what has previously been reported in the literature. A rigorous framework for evaluating unsupervised object discovery methods is proposed.
Tinne Tuytelaars, Christoph H. Lampert, Matthew B. Blaschko, Wray L. Buntine
Int. J. Comput. Vis.1
2010 Kernelized Sorting
abstract
Object matching is a fundamental operation in data analysis. It typically requires the definition of a similarity measure between the classes of objects to be matched. Instead, we develop an approach which is able to perform matching by requiring a similarity measure only within each of the classes. This is achieved by maximizing the dependency between matched pairs of observations by means of the Hilbert-Schmidt Independence Criterion. This problem can be cast as one of maximizing a quadratic assignment problem with special structure and we present a simple algorithm for finding a locally optimal solution.
Novi Quadrianto, Alexander J. Smola, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.4
2010 Cross-Media Alignment of Names and Faces
abstract
In this paper we report on our experiments on aligning names and faces as found in images and captions of online news Websites. Developing accurate technologies for linking names and faces is valuable when retrieving or mining information from multimedia collections. We perform exhaustive and systematic experiments exploiting the (a)symmetry between the visual and textual modalities. This leads to different schemes for assigning names to the faces, assigning faces to the names, and establishing name-face link pairs. On top of that, we investigate generic approaches to the use of textual and visual structural information to predict the presence of the corresponding entity in the other modality. The proposed methods are completely unsupervised and are inspired by methods for aligning phrases and words in texts of different languages developed for constructing dictionaries for machine translation. The results are competitive with state-of-the-art performance on the ¿Labeled Faces in the Wild¿ dataset in terms of recall values, now reported on the complete dataset, include excellent precision values, and show the value of text and image analysis for identifying the probability of being pictured or named in the alignment process.
Phi The Pham, Marie-Francine Moens, Tinne Tuytelaars
IEEE Trans. Multim.3
2009 Exemplar-based Action Recognition in Video
abstract
In this work, we present a method for action localization and recognition using an exemplar-based approach. It starts from local dense yet scale-invariant spatio-temporal features. The most discriminative visual words are selected and used to cast bounding box hypotheses, which are then verified and further grouped into the final detections. To the best of our knowledge, we are the first to extend the exemplar-based approach using local features into the spatio-temporal domain. This allows us to avoid the problems that typically plague sliding window-based approaches - in particular the exhaustive search over spatial coordinates, time, and spatial as well as temporal scales. We report state-of-the-art results on challenging datasets, extracted from real movies, for both classification and localization. © 2009. The copyright of this document resides with its authors.
Geert Willems, Jan Hendrik Becker, Tinne Tuytelaars, Luc Van Gool
BMVC3
2009 Exploring Scale-Induced Feature Hierarchies in Natural Images
abstract
Recently there has been considerable interest in topic models based on the bag-of-features representation of images. The strong independence assumption inherent in the bag-of-features representation is not realistic however: patches often overlap and share underlying image structures. Moreover, important information with respect to relative scales of the features is completely ignored, for the sake of scale invariance. Considering both spatial and scale-based constraints one can derive spatially constrained natural feature hierarchies within images. We explore the use of topic models that build such spatially constrained scale-induced hierarchies of the features in an unsupervised fashion. Our model uses standard topic models as a starting point. We then incorporate information about the hierarchical and spatial relations of the features into the model. We illustrate the hierarchical nature of the resulting models using datasets of natural images, including the MSRC2 dataset as well as a challenging set of images of trees collected from the Internet.
Jukka Perkiö, Tinne Tuytelaars, Wray L. Buntine
ICMLA2
2009 Special issue on 3D representation for object and scene recognition
Silvio Savarese, Tinne Tuytelaars, Luc Van Gool
Comput. Vis. Image Underst.2
2009 Shape-from-recognition: Recognition enables meta-data transfer
Alexander Thomas, Vittorio Ferrari, Bastian Leibe, Tinne Tuytelaars, Luc Van Gool
Comput. Vis. Image Underst.4
2008 An Efficient Dense and Scale-Invariant Spatio-Temporal Interest Point Detector
Geert Willems, Tinne Tuytelaars, Luc Van Gool
ECCV (2)2
2008 Speeded-Up Robust Features (SURF)
Herbert Bay, Andreas Ess, Tinne Tuytelaars, Luc Van Gool
Comput. Vis. Image Underst.3
2007 Depth-From-Recognition: Inferring Meta-data by Cognitive Feedback
abstract
Thanks to recent progress in category-level object recognition, we have now come to a point where these techniques have gained sufficient maturity and accuracy to succesfully feed back their output to other processes. This is what we refer to as cognitive feedback. In this paper, we study one particular form of cognitive feedback, where the ability to recognize objects of a given category is exploited to infer meta-data such as depth cues, 3D points, or object decomposition in images of previously unseen object instances. Our approach builds on the implicit shape model of Leibe and Schiele, and extends it to transfer annotations from training images to test images. Experimental results validate the viability of our approach.
Alexander Thomas, Vittorio Ferrari, Bastian Leibe, Tinne Tuytelaars, Luc Van Gool
ICCV4
2007 Vector Quantizing Feature Space with a Regular Lattice
abstract
Most recent class-level object recognition systems work with visual words, i.e., vector quantized local descriptors. In this paper we examine the feasibility of a data- independent approach to construct such a visual vocabulary, where the feature space is discretized using a regular lattice. Using hashing techniques, only non-empty bins are stored, and fine-grained grids become possible in spite of the high dimensionality of typical feature spaces. Based on this representation, we can explore the structure of the feature space, and obtain state-of-the-art pixelwise classification results. In the case of image classification, we introduce a class-specific feature selection step, which takes the spatial structure of SIFT-like descriptors into account. Results are reported on the Graz02 dataset.
Tinne Tuytelaars, Cordelia Schmid
ICCV1
2007 Omnidirectional Vision Based Topological Navigation
Toon Goedemé, Marnix Nuttin, Tinne Tuytelaars, Luc Van Gool
Int. J. Comput. Vis.3
2007 A Thousand Words in a Scene
abstract
This paper presents a novel approach for visual scene modeling and classification, investigating the combined use of text modeling methods and local invariant features. Our work attempts to elucidate (1) whether a text-like bag-of-visterms representation (histogram of quantized local visual features) is suitable for scene (rather than object) classification, (2) whether some analogies between discrete scene representations and text documents exist, and (3) whether unsupervised, latent space models can be used both as feature extractors for the classification task and to discover patterns of visual co-occurrence. Using several data sets, we validate our approach, presenting and discussing experiments on each of these issues. We first show, with extensive experiments on binary and multi-class scene classification tasks using a 9,500-image data set, that the bag-of-visterms representation consistently outperforms classical scene classification approaches. In other data sets we show that our approach competes with or outperforms other recent, more complex, methods. We also show that Probabilistic Latent Semantic Analysis (PLSA) generates a compact scene representation, discriminative for accurate classification, and more robust than the bag-of-visterms representation when less labeled training data is available. Finally, through aspect-based image ranking experiments, we show the ability of PLSA to automatically extract visually meaningful scene patterns, making such representation useful for browsing image collections.
Pedro Quelhas, Florent Monay, Jean-Marc Odobez, Daniel Gatica-Perez, Tinne Tuytelaars
IEEE Trans. Pattern Anal. Mach. Intell.5
2006 Towards Multi-View Object Class Detection
abstract
We present a novel system for generic object class detection. In contrast to most existing systems which focus on a single viewpoint or aspect, our approach can detect object instances from arbitrary viewpoints. This is achieved by combining the Implicit Shape Model for object class detection proposed by Leibe and Schiele with the multi-view specific object recognition system of Ferrari et al. After learning single-view codebooks, these are interconnected by so-called activation links, obtained through multi-view region tracks across different training views of individual object instances. During recognition, these integrated codebooks work together to determine the location and pose of the object. Experimental results demonstrate the viability of the approach and compare it to a bank of independent single-view detectors
Alexander Thomas, Vittorio Ferrari, Bastian Leibe, Tinne Tuytelaars, Bernt Schiele, Luc Van Gool
CVPR (2)4
2006 SURF: Speeded Up Robust Features
Herbert Bay, Tinne Tuytelaars, Luc Van Gool
ECCV (1)2
2006 Object Detection by Contour Segment Networks
Vittorio Ferrari, Tinne Tuytelaars, Luc Van Gool
ECCV (3)2
2006 Localization with Omnidirectional Images using the Radial Trifocal Tensor
abstract
In this paper we present a technique to linearly recover 2D structure and motion in man made environments from three uncalibrated omnidirectional views. We use vertical lines from the scene which are projected as radial lines in the images and are automatically matched. The algorithm is based on a 1D radial trifocal tensor which encodes the relations of the three views and the projected lines. We include experiments with real images, which demonstrate the good performance of the method and its application to robotic tasks, such as robot localization based in a database of reference images or to obtain the initial values of robot and landmarks localization in SLAM algorithms
Carlos Sagüés, Ana Cristina Murillo, Josechu J. Guerrero, Toon Goedemé, Tinne Tuytelaars, Luc Van Gool
ICRA5
2006 Simultaneous Object Recognition and Segmentation from Single or Multiple Model Views
Vittorio Ferrari, Tinne Tuytelaars, Luc Van Gool
Int. J. Comput. Vis.2
2005 Modeling Scenes with Local Descriptors and Latent Aspects
abstract
We present a new approach to model visual scenes in image collections, based on local invariant features and probabilistic latent space models. Our formulation provides answers to three open questions:(l) whether the invariant local features are suitable for scene (rather than object) classification; (2) whether unsupennsed latent space models can be used for feature extraction in the classification task; and (3) whether the latent space formulation can discover visual co-occurrence patterns, motivating novel approaches for image organization and segmentation. Using a 9500-image dataset, our approach is validated on each of these issues. First, we show with extensive experiments on binary and multi-class scene classification tasks, that a bag-of-visterm representation, derived from local invariant descriptors, consistently outperforms state-of-the-art approaches. Second, we show that probabilistic latent semantic analysis (PLSA) generates a compact scene representation, discriminative for accurate classification, and significantly more robust when less training data are available. Third, we have exploited the ability of PLSA to automatically extract visually meaningful aspects, to propose new algorithms for aspect-based image ranking and context-sensitive image segmentation.
Pedro Quelhas, Florent Monay, Jean-Marc Odobez, Daniel Gatica-Perez, Tinne Tuytelaars, Luc Van Gool
ICCV5
2005 Feature based omnidirectional sparse visual path following
abstract
Vision sensors are attractive for autonomous robots because they are a rich source of environment information. The main challenge in using images for mobile robots is managing this wealth of information. A relatively recent approach is the use of fast wide baseline local features, which we developed and used in the novel approach to sparse visual path following described in this paper. These local features have the great advantage that they can be recognized even if the viewpoint differs significantly. This opens the door to a memory efficient description of a path by descriptors of sparse images. We propose a method for re-execution of these paths by a series of visual homing operations which yield a navigation method with unique properties: it is accurate, robust, fast, and without odometry error build-up.
Toon Goedemé, Tinne Tuytelaars, Luc Van Gool, Gerolf Vanacker, Marnix Nuttin
IROS2
2005 A Comparison of Affine Region Detectors
Krystian Mikolajczyk, Tinne Tuytelaars, Cordelia Schmid, Andrew Zisserman, Jiri Matas, Frederik Schaffalitzky, Timor Kadir, Luc Van Gool
Int. J. Comput. Vis.2
2004 Integrating Multiple Model Views for Object Recognition
Vittorio Ferrari, Tinne Tuytelaars, Luc Van Gool
CVPR (2)2
2004 Fast Wide Baseline Matching for Visual Navigation
Toon Goedemé, Tinne Tuytelaars, Luc Van Gool
CVPR (1)2
2004 Synchronizing Video Sequences
Tinne Tuytelaars, Luc Van Gool
CVPR (1)1
2004 Simultaneous Object Recognition and Segmentation by Image Exploration
Vittorio Ferrari, Tinne Tuytelaars, Luc Van Gool
ECCV (1)2
2004 Moment invariants for recognition under changing viewpoint and illumination
Florica Mindru, Tinne Tuytelaars, Luc Van Gool, Theo Moons
Comput. Vis. Image Underst.2
2004 Matching Widely Separated Views Based on Affine Invariant Regions
Tinne Tuytelaars, Luc Van Gool
Int. J. Comput. Vis.1
2003 Wide-baseline Multiple-view Correspondences
abstract
We present a novel approach for establishing multiple-view feature correspondences along an unordered set of images taken from substantially different viewpoints. Several wide-baseline stereo (WBS) algorithms have appeared, the N-view case is largely unexplored. In this paper, an established WBS algorithm is used to extract and match features in pairs of views. The pairwise matches are first integrated into disjoint feature tracks, each representing a single physical surface patch in several views. By exploiting the interplay between the tracks, they are extended over more views, while unrelated image features are removed. Similarity and spatial relationships between the features are simultaneously used. The output consists of many reliable and accurate feature tracks, strongly connecting the input views. Applications include 3D reconstruction and object recognition. The proposed approach is not restricted to the particular choice of features and matching criteria. It can extend any method that provides feature correspondences between pairs of images.
Vittorio Ferrari, Tinne Tuytelaars, Luc Van Gool
CVPR (1)2
2003 Dense Matching of Multiple Wide-baseline Views
abstract
This paper describes a PDE-based method for dense depth extraction from multiple wide-baseline images. Emphasis lies on the usage of only a small amount of images. The integration of these multiple wide-baseline views is guided by the relative confidence that the system has in the matching to different views. This weighting is fine-grained in that it is determined for every pixel at every iteration. Reliable information spreads fast at the expense of less reliable data, both in terms of spatial communications within a view and in terms of information exchange between the views. Changes in intensity between images can be handled in a similar fine grained fashion.
Christoph Strecha, Tinne Tuytelaars, Luc Van Gool
ICCV2
2003 Fast indexing for image retrieval based on local appearance with re-ranking
abstract
This paper describes an approach to retrieve images containing specific objects, scenes or buildings. The image content is captured by a set of local features. More precisely, we use so-called invariant regions. These are features with shapes that self-adapt to the viewpoint. The physical parts on the object surface that they carve out are the same in all views, even though the extraction proceeds from a single view only. The surface patterns within the regions are then characterized by a feature vector of moment invariants. Invariance is under affine geometric deformations and scaled color bands with an offset added. This allows regions from different views to be matched efficiently. An indexing technique based on vantage point tree organizes the feature vectors in such a way that a naive sequential search can be avoided. This results in sublinear computation times to retrieve images from a database. In order to get sufficient certainty about the correctness of the retrieved images, a method to increase the number of matched regions is introduced. This way, the system is both efficient and discriminant. It is demonstrated how scenes or buildings are recognized, even in case of partial visibility and under a large variety of viewing condition changes.
Hao Shao, Tomás Svoboda, Vittorio Ferrari, Tinne Tuytelaars, Luc Van Gool
ICIP (3)4
2003 Noncombinatorial Detection of Regular Repetitions under Perspective Skew
abstract
We present a geometric framework for the efficient detection of regular repetitions of planar (but not necessarily coplanar) patterns. At the heart of our system, lie the fixed structures of the transformations that describe these regular configurations. The approach detects a number of symmetric configurations that have traditionally been dealt with separately, in that all configurations corresponding to planar homologies are detected. These include important cases such as periodicities, mirror symmetries, and reflections about a point. The approach can handle perspective distortions. It avoids to get trapped in combinatorics; through invariant-based hashing for pattern matching and through Hough transforms for the detection of fixed structures. Additional efficiency and robustness are obtained from the system's ability to "reason" about the consistency of multiple homologies. The performance of the system is demonstrated with several examples.
Tinne Tuytelaars, Andreas Turina, Luc Van Gool
IEEE Trans. Pattern Anal. Mach. Intell.1
2001 Real-time affine region tracking and coplanar grouping
abstract
We present a novel approach for tracking locally planar regions in an image sequence and their grouping into larger planar surfaces. The tracker recovers the affine transformation of the region and therefore yields reliable point correspondences between frames. Both edges and texture information are exploited in an integrated way, while not requiring the complete region's contour. The tracker withstands zoom, out-of-plane rotations, discontinuous motion and changes in illumination conditions while achieving real-time performance for a region. Multiple tracked regions are grouped into disjoint coplanarity classes. We first define a coplanarity score between each pair of regions, based on motion and texture cues. The scores are then analyzed by a clique-partitioning algorithm yielding the coplanarity classes that best fit the data. The method works in the presence of perspective distortions, discontinuous planar surfaces and considerable amounts of measurement noise.
Vittorio Ferrari, Tinne Tuytelaars, Luc Van Gool
CVPR (2)2
2001 Efficient Grouping under Perspective Skew
abstract
We present an efficient grouping strategy for the detection of regular repetitions of planar (but not necessarily coplanar) patterns. At the heart of our system lie the fixed structures that typify the geometric transformations of the regularities. The approach unifies a number of grouping types that have traditionally been dealt with separately. It avoids the use of combinatorics in the search for pattern repetitions, through the combined use of invariants for hashing-based pattern matching on the one hand, and Hough transforms for the detection of the fixed structures on the other hand. In this paper we concentrate on planar homologies and elations in particular Results on real-world scenes demonstrate the performance of the approach.
Andreas Turina, Tinne Tuytelaars, Luc Van Gool
CVPR (1)2
2000 Wide Baseline Stereo Matching based on Local, Affinely Invariant Regions
abstract
`Invariant regions' are image patches that automatically deform with changing viewpoint as to keep on covering identical physical parts of a scene. Such regions are then described by a set of invariant features, which makes it relatively easy to match them between views and under changing illumination. In previous work, we have presented invariant regions that are based on a combination of corners and edges. The application discussed then was image database retrieval. Here, an alternative method for extracting (affinely) invariant regions is given, that does not depend on the presence of edges or corners in the image but is purely intensity-based. Also, we demonstrate the use of such regions for another application, which is wide baseline stereo matching. As a matter of fact, the goal is to build an opportunistic system that exploits several types of invariant regions as it sees fit. This yields more correspondences and a system that can deal with a wider range of images. To increase t...
Tinne Tuytelaars, Luc Van Gool
BMVC1
2000 Automatic Object Recognition as Part of an Integrated Supervisory Control System
abstract
The paper consists of two main contributions. First, a generic object recognition based on affinely invariant regions is proposed. Next, this algorithm is used in an integrated supervisory control system (ISCS), with special attention to the human machine interaction. Experiments on an industrial robotic system are included, performing simple tasks such as "Go to object O/sub 1/" or "Put object O/sub 2/ on top of object O/sub 3/", where an object is simply modeled by one or more images thereof.
Tinne Tuytelaars, A. Zaatri, Luc Van Gool, Hendrik Van Brussel
ICRA1
2000 On Satellite Vision-Aided Robotics Experiment
abstract
This paper describes the vision-based robotic control (VBRC) experiments executed on the Japanese research satellite ETS-VII. The VBRC experiments were designed to enhance image quality, refine calibration of different system components, facilitate robot-operation by automatically refining the robot-pose and provide data for robot-calibration.
Maarten Vergauwen, Marc Pollefeys, Tinne Tuytelaars, Luc Van Gool
ICRA3
1999 Adventurous Tourism for Couch Potatoes
Luc Van Gool, Tinne Tuytelaars, Marc Pollefeys
CAIP2
1999 Matching of Affinely Invariant Regions for Visual Servoing
abstract
This paper develops new image matching techniques for visual servoing based on affine invariants which allow one to deal with large viewpoint changes and that do not rely on specific markers. The only assumption is that there are some locally planar and unoccluded scene regions that have enough structure to be detected in the image. Those regions are classified by a set of illumination and viewpoint invariant features. The features represent the image in a very compact way and allow fast comparison and feature matching between quite different viewpoints. The matching procedure is embedded in a visual servoing system for a mobile robot. Experiments show its potential for navigation with large camera rotations and view point changes in a cluttered environment without the need for artificial landmarks.
Tinne Tuytelaars, Luc Van Gool, L. D'haene, Reinhard Koch
ICRA1
1998 A Cascaded Hough Transform as an Aid in Aerial Image Interpretation
abstract
Cartography and other applications of remote sensing have led to an increased interest in the (semi-)automatic interpretation of structures in aerial images of urban and suburban areas. Although these areas are particularly challenging because of their complexity, the degree of regularity in such man-made structures also helps to tackle the problems. The paper presents the iterated application of the Hough transform as a means to exploit such regularities. It shows how such 'Cascaded Hough Transform' (or CHT for short) yields straight lines, vanishing points, and vanishing lines. It also illustrates how the latter assist in improving the precision of the former. The examples are based on real aerial photographs.
Tinne Tuytelaars, Luc Van Gool, Marc Proesmans, Theo Moons
ICCV1
1997 The Cascaded Hough Transforms
abstract
When using the original slope-intercept parameterisation for the Hough transform, the resulting parameter space actually corresponds to the dual space. Indeed, lines are transformed into points, and for every point there is also a corresponding line. This paper presents a way of exploiting this special property, by the introduction of the cascaded Hough transform (CHT). This allows to look for the overall structure in an image, such as lines intersecting in a point or intersection points lying on a line. An interesting example is the detection of vanishing points and vanishing lines.
Tinne Tuytelaars, Marc Proesmans, Luc Van Gool
ICIP (2)1