VLDB 2026 Research / reviewers in the wild / expert
Esa Rahtu
dblp:34/3846
· DBLP profile ↗
95ranked-venue papers
8as first author
48since 2021 · last 2026
0000-0001-8767-0864ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 83 · 5 first-author · 44 since 2021Artificial intelligence and machine learning · 48 · 8 first-author · 17 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CorkHSI: Hyperspectral Anomaly Detection in Corks Using an Autoencoder with a Novel Spectral-Spatial Loss Optimization
Afshin Dini, Farnaz Delirie, Esa Rahtu |
ICPR (1) | 3 |
| 2026 | Pose-Guided Geometric Refinement for Feed-Forward 3D Gaussian Splatting
Yejun Zhang 0001, Esa Rahtu, Juho Kannala |
ICPR (12) | 4 |
| 2026 | SceneShine: Illumination-aware Human Scene Gaussian Re-Splatting from Mobile Device VideoabstractStandard 3DGS falls short in precise relighting and shadowing needed to realistically integrate humans into novel environments. We bridge this gap with SceneShine, an illumination-aware framework designed for seamless composition through physically-based avatar relighting and shadow casting. Relighting human surfaces in in-the-wild videos is inherently ill-posed, often making the simultaneous disentanglement of scene lighting and BRDF properties difficult. We overcome this ambiguity by utilizing a pseudo-global light map prior to guide BRDF parameter decomposition, significantly reducing relighting artifacts. Additionally, we implement point-based ray tracing to manage human-scene occlusions and dynamically update scene colors for accurate shadow casting. We also introduce a new synthetic dataset for evaluation. Extensive experiments show that our method surpasses existing approaches in reconstruction fidelity and identity preservation while achieving highly convincing illumination-aware integration1. Xuqian Ren, Wenjia Wang 0009, Mai Ngoc Nguyen, Juho Kannala, Esa Rahtu |
WACV | 5 |
| 2026 | Video Object Segmentation-Aware Audio GenerationabstractAbstract Existing multimodal audio generation models often lack precise user control, which limits their applicability in professional Foley workflows. In particular, these models focus on the entire video and do not provide precise methods for prioritizing a specific object within a scene, generating unnecessary background sounds, or focusing on the wrong objects. To address this gap, we introduce the novel task of video object segmentation-aware audio generation, which explicitly conditions sound synthesis on object-level segmentation maps. We present SAGANet, a new multimodal generative model that enables controllable audio generation for musical instruments by leveraging visual segmentation masks along with video and textual cues. Our model provides users with fine-grained and visually localized control over audio generation. To support this task and further research on segmentation-aware Foley, we propose Segmented Music Solos, a benchmark dataset of musical instrument performance videos with segmentation information. Our method demonstrates substantial improvements over current state-of-the-art methods and sets a new standard for controllable, high-fidelity Foley synthesis for musical audio. Code, samples, and Segmented Music Solos are available at http://saganet.notion.site . Ilpo Viertola, Vladimir Iashin, Esa Rahtu |
Int. J. Comput. Vis. | 3 |
| 2025 | GS-Pose: Generalizable Segmentation-Based 6D Object Pose Estimation with 3D Gaussian SplattingabstractThis paper introduces GS-Pose, a unified framework for localizing and estimating the 6D pose of novel objects. GS-Pose begins with a set of posed RGB images of a previously unseen object and builds three distinct representations stored in a database. At inference, GS-Pose operates sequentially by locating the object in the input image, estimating its initial 6D pose using a retrieval approach, and refining the pose with a render-and-compare method. The key insight is the application of the appropriate object representation at each stage of the process. In particular, for the refinement step, we leverage 3D Gaussian splatting, a novel differentiable rendering technique that offers high rendering speed and relatively low optimization time. Off-the-shelf toolchains and commodity hard-ware, such as mobile phones, can be used to capture new objects to be added to the database. Extensive evaluations on the LINEMOD and OnePose-LowTexture datasets demonstrate excellent performance, establishing the new state-of-the-art. The source code is publicly available at https://github.com/dingdingcai/GSPose. Dingding Cai, Janne Heikkilä, Esa Rahtu |
3DV | 3 |
| 2025 | AGS-Mesh: Adaptive Gaussian Splatting and Meshing with Geometric Priors for Indoor Room Reconstruction Using SmartphonesabstractGeometric priors are often used to enhance 3D reconstruction. With many smartphones featuring low-resolution depth sensors and the prevalence of off-the-shelf monocular geometry estimators, incorporating geometric priors as regularization signals has become common in 3D vision tasks. However, the accuracy of depth estimates from mobile devices is typically poor for highly detailed geometry, and monocular estimators often suffer from poor multi-view consistency and precision. In this work, we propose an approach for joint surface depth and normal refinement of Gaussian Splatting methods for accurate 3D reconstruction of indoor scenes. We develop supervision strategies that adaptively filters low-quality depth and normal estimates by comparing the consistency of the priors during optimization. We mitigate regularization in regions where prior estimates have high uncertainty or ambiguities. Our filtering strategy and optimization design demonstrate significant improvements in both mesh estimation and novel-view synthesis for both 3D and 2D Gaussian Splatting-based methods on challenging indoor room datasets. Furthermore, we explore the use of alternative meshing strategies for finer geometry extraction. We develop a scale-aware meshing strategy inspired by TSDF and octree-based isosurface extraction, which recovers finer details from Gaussian models compared to other commonly used open-source meshing tools. Our code is released in https://xuqianren.github.io/ags_mesh_website/. Xuqian Ren, Matias Turkulainen, Jiepeng Wang 0001, Otto Seiskari, Iaroslav Melekhov, Juho Kannala, Esa Rahtu |
3DV | 7 |
| 2025 | Temporally Aligned Audio for Video with AutoregressionabstractWe introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extractor and a cross-modal audio-visual feature fusion strategy to capture fine-grained visual motion events and ensure precise temporal alignment. Additionally, we propose VisualSound, a benchmark dataset with high audio-visual relevance. VisualSound is based on VGGSound, a video dataset consisting of in-the-wild samples extracted from YouTube. During the curation, we remove samples where auditory events are not aligned with the visual ones. V-AURA outperforms current state-of-the-art models in temporal alignment and semantic relevance while maintaining comparable audio quality. Code, samples, VisualSound and models are available at v-aura.notion.site. Ilpo Viertola, Vladimir Iashin, Esa Rahtu |
ICASSP | 3 |
| 2025 | A Hybrid Framework Integrating End-to-End Learned Image Codec with Conventional Codec
Nannan Zou, Antti Hallapuro, Francesco Cricri, Honglei Zhang 0001, A. Burakhan Koyuncu, Jukka I. Ahonen, Miska M. Hannuksela, Esa Rahtu |
PCS | 8 |
| 2025 | Heterogeneous Datasets for Unsupervised Image Anomaly DetectionabstractUnsupervised anomaly detection (AD) is a critical task in various domains, from manufacturing to infrastructure monitoring. To advance this field, we introduce two novel datasets: CARS-AD and ROADS-AD, designed to challenge existing unsupervised AD methods with their diverse and heterogeneous image content. CARS-AD comprises real images of cars with various defects, while ROADS-AD contains images of roads from multiple countries, each presenting unique challenges in anomaly detection. These datasets provide ground truth pixel-wise masks and image-level ground truth labels, enabling detailed evaluation and benchmarking of AD algorithms. We evaluate state-of-the-art unsupervised AD methods on both datasets, using the AUROC metric to assess detection and localization performance. Our results reveal significant room for improvement, underscoring the complexity of the datasets and the need for robust AD techniques. Notably, Csflow and U-Flow demonstrate superior performance on the CARS-AD Dataset, leveraging their ability to process multi-scale features effectively. Conversely, Reverse Distillation excels in anomaly localization on the ROADS-AD Dataset, showcasing the importance of nuanced approaches for diverse anomaly types. Our findings underscore the importance of addressing the challenges posed by heterogeneous datasets in unsu-pervised AD. We hope that the introduction of CARS-AD and ROADS-AD will inspire further research in this field, driving the development of innovative AD methods capa-ble of handling real-world anomalies with greater accuracy and reliability. CARS-AD Dataset and ROADS-AD Dataset are publicly available at https://github.com/juanb09111/heterogeneousAD. Juan Pablo Lagos, Adnan Faroque, Esa Rahtu |
WACV | 4 |
| 2025 | DN-Splatter: Depth and Normal Priors for Gaussian Splatting and MeshingabstractHigh-fidelity 3D reconstruction of common indoor scenes is crucial for VR and AR applications. 3D Gaussian splat-ting, a novel differentiable rendering technique, has achieved state-of-the-art novel view synthesis results with high ren-dering speeds and relatively low training times. However, its performance on scenes commonly seen in indoor datasets is poor due to the lack of geometric constraints during op-timization. In this work, we explore the use of readily accessible geometric cues to enhance Gaussian splatting op-timization in challenging, ill-posed, and textureless scenes. We extend 3D Gaussian splatting with depth and normal cues to tackle challenging indoor datasets and showcase techniques for efficient mesh extraction. Specifically, we regularize the optimization procedure with depth information, enforce local smoothness of nearby Gaussians, and use off-the-shelf monocular networks to achieve better align-ment with the true scene geometry. We propose an adaptive depth loss based on the gradient of color images, improving depth estimation and novel view synthesis results over various baselines. Our simple yet effective regularization technique enables direct mesh extraction from the Gaus-sian representation, yielding more physically accurate re-constructions of indoor scenes. Our code will be released in https://github.com/maturk/dn-splatter. Matias Turkulainen, Xuqian Ren, Iaroslav Melekhov, Otto Seiskari, Esa Rahtu, Juho Kannala |
WACV | 5 |
| 2024 | Gaussian Splatting on the Move: Blur and Rolling Shutter Compensation for Natural Camera Motion
Otto Seiskari, Jerry Ylilammi, Valtteri Kaatrasalo, Pekka Rantalankila, Matias Turkulainen, Juho Kannala, Esa Rahtu, Arno Solin |
ECCV (71) | 7 |
| 2024 | Synchformer: Efficient Synchronization From Sparse CuesabstractOur objective is audio-visual synchronization with a focus on ‘in-the-wild’ videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model, and training that decouples feature extraction from synchronization modelling through multi-modal segment-level contrastive pre-training. This approach achieves state-of-the-art performance in both dense and sparse settings. We also extend synchronization model training to AudioSet a million-scale ‘in-the-wild’ dataset, investigate evidence attribution techniques for interpretability, and explore a new capability for synchronization models: audio-visual synchronizability. robots.ox.ac.uk/~vgg/research/synchformer Vladimir Iashin, Weidi Xie, Esa Rahtu, Andrew Zisserman |
ICASSP | 3 |
| 2024 | MuSHRoom: Multi-Sensor Hybrid Room Dataset for Joint 3D Reconstruction and Novel View SynthesisabstractMetaverse technologies demand accurate, real-time, and immersive modeling on consumer-grade hardware for both non-human perception (e.g., drone/robot/autonomous car navigation) and immersive technologies like AR/VR, requiring both structural accuracy and photorealism. However, there exists a knowledge gap in how to apply geometric reconstruction and photorealism modeling (novel view synthesis) in a unified framework. To address this gap and promote the development of robust and immersive modeling and rendering with consumer-grade devices, first, we propose a real-world Multi-Sensor Hybrid Room Dataset (MuSHRoom). Our dataset presents exciting challenges and requires state-of-the-art methods to be cost-effective, robust to noisy data and devices, and can jointly learn 3D reconstruction and novel view synthesis instead of treating them as separate tasks, making them ideal for realworld applications. Second, we benchmark several famous pipelines on our dataset for joint 3D mesh reconstruction and novel view synthesis. Finally, in order to further improve the overall performance, we propose a new method that achieves a good trade-off between the two tasks. Our dataset and benchmark show great potential in promoting the improvements for fusing 3D reconstruction and highquality rendering in a robust and computationally efficient end-to-end fashion. The dataset and code are available at the project website: https://xuqianren.github.io/publications/MuSHRoom/. Xuqian Ren, Wenjia Wang 0009, Dingding Cai, Tuuli Tuominen, Juho Kannala, Esa Rahtu |
WACV | 6 |
| 2024 | Cascaded and Generalizable Neural Radiance Fields for Fast View SynthesisabstractWe present CG-NeRF, a cascade and generalizable neural radiance fields method for view synthesis. Recent generalizing view synthesis methods can render high-quality novel views using a set of nearby input views. However, the rendering speed is still slow due to the nature of uniformly-point sampling of neural radiance fields. Existing scene-specific methods can train and render novel views efficiently but can not generalize to unseen data. Our approach addresses the problems of fast and generalizing view synthesis by proposing two novel modules: a coarse radiance fields predictor and a convolutional-based neural renderer. This architecture infers consistent scene geometry based on the implicit neural fields and renders new views efficiently using a single GPU. We first train CG-NeRF on multiple 3D scenes of the DTU dataset, and the network can produce high-quality and accurate novel views on unseen real and synthetic data using only photometric losses. Moreover, our method can leverage a denser set of reference images of a single scene to produce accurate novel views without relying on additional explicit representations and still maintains the high-speed rendering of the pre-trained model. Experimental results show that CG-NeRF outperforms state-of-the-art generalizable neural rendering methods on various synthetic and real datasets. Phong Nguyen 0001, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Momentum Adapt: Robust Unsupervised Adaptation for Improving Temporal Consistency in Video Semantic Segmentation During Test-Time
Amirhossein Hassankhani, Hamed Rezazadegan Tavakoli, Esa Rahtu |
BMVC | 3 |
| 2023 | Toward Verifiable and Reproducible Human Evaluation for Text-to-Image GenerationabstractHuman evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works rely solely on automatic measures (e.g., FID) or perform poorly described human evaluations that are not reliable or repeatable. This paper proposes a standardized and well-defined human evaluation protocol to facilitate verifiable and reproducible human evaluation in future works. In our pilot data collection, we experimentally show that the current automatic measures are incompatible with human perception in evaluating the performance of the text-to-image generation results. Furthermore, we provide insights for designing human evaluation experiments reliably and conclusively. Finally, we make several resources publicly available to the community to facilitate easy and fast implementations. Mayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Shin'ichi Satoh 0001 |
CVPR | 6 |
| 2023 | NN-VVC: Versatile Video Coding boosted by self-supervisedly learned image coding for machinesabstractThe recent progress in artificial intelligence has led to an ever-increasing usage of images and videos by machine analysis algorithms, mainly neural networks. Nonetheless, compression, storage and transmission of media have traditionally been designed considering human beings as the viewers of the content. Recent research on image and video coding for machine analysis has progressed mainly in two almost orthogonal directions. The first is represented by end-to-end (E2E) learned codecs which, while offering high performance on image coding, are not yet on par with state-of-the-art conventional video codecs and lack interoperability. The second direction considers using the Versatile Video Coding (VVC) standard or any other conventional video codec (CVC) together with pre- and post-processing operations targeting machine analysis. While the CVC-based methods benefit from interoperability and broad hardware and software support, the machine task performance is often lower than the desired level, particularly in low bitrates. This paper proposes a hybrid codec for machines called NN-VVC, which combines the advantages of an E2E-learned image codec and a CVC to achieve high performance in both image and video coding for machines. Our experiments show that the proposed system achieved up to -43.20% and -26.8% Bjøntegaard Delta rate reduction over VVC for image and video data, respectively, when evaluated on multiple different datasets and machine vision tasks. To the best of our knowledge, this is the first research paper showing a hybrid video codec that outperforms VVC on multiple datasets and multiple machine vision tasks. Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Antti Hallapuro, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 8 |
| 2023 | Region of Interest Enabled Learned Image Coding for MachinesabstractImage and video coding for machines has been recently gaining more and more interest from both the industry and the research community. One successful approach is based on end-to-end (E2E) learned compression and has shown significant gains over the state-of-the-art conventional image coding methods. However, one of the remaining challenges for such E2E-learned image codecs for machines is to adaptively allocate the bits over different regions of the image, while retaining the machine vision performance. In this paper, we propose a method that leverages Regions-Of-Interest (ROIs) for bitrate allocation within a Learned Image Codec (LIC) for machines. In particular, the proposed method reduces the bits allocated for the background regions of the image by reducing the variance of the elements corresponding to the background regions in the latent representation. This results in more heavily quantized background areas, while keeping the quality of the ROI areas suitable for machine tasks. The proposed method achieves significant gains, -15.80% and -22.43% Pareto BD-rate reduction, over the baseline LIC on object detection and instance segmentation tasks, respectively. To the best of our knowledge, this is the first research paper proposing an ROI-based inference-time technology for Learned Image Coding for machines. Jukka I. Ahonen, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Esa Rahtu |
MMSP | 5 |
| 2022 | SC6D: Symmetry-agnostic and Correspondence-free 6D Object Pose EstimationabstractThis paper presents an efficient symmetry-agnostic and correspondence-free framework, referred to as SC6D, for 6D object pose estimation from a single monocular RGB image. SC6D requires neither the 3D CAD model of the object nor any prior knowledge of the symmetries. The pose estimation is decomposed into three sub-tasks: a) object 3D rotation representation learning and matching; b) estimation of the 2D location of the object center; and c) scale-invariant distance estimation (the translation along the z-axis) via classification. SC6D is evaluated on three benchmark datasets, T-LESS, YCB-V, and ITODD, and results in state-of-the-art performance on the T-LESS dataset. More-over, SC6D is computationally much more efficient than the previous state-of-the-art method SurfEmb. The implementation and pre-trained models are publicly available at https://github.com/dingdingcai/SC6D-pose. Dingding Cai, Janne Heikkilä, Esa Rahtu |
3DV | 3 |
| 2022 | Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors
Vladimir Iashin, Weidi Xie, Esa Rahtu, Andrew Zisserman |
BMVC | 3 |
| 2022 | OVE6D: Object Viewpoint Encoding for Depth-based 6D Object Pose EstimationabstractThis paper proposes a universal framework, called OVE6D, for model-based 6D object pose estimation from a single depth image and a target object mask. Our model is trained using purely synthetic data rendered from ShapeNet, and, unlike most of the existing methods, it generalizes well on new real-world objects without any fine-tuning. We achieve this by decomposing the 6D pose into viewpoint, in-plane rotation around the camera optical axis and translation, and introducing novel lightweight modules for estimating each component in a cascaded manner. The resulting network contains less than 4M parameters while demon-strating excellent performance on the challenging T-LESS and Occluded LINEMOD datasets without any dataset-specific training. We show that OVE6D outperforms some contemporary deep learning-based pose estimation methods specifically trained for individual objects or datasets with real-world training data. The implementation is available at https://github.com/dingdingcai/OVE6D-pose. Dingding Cai, Janne Heikkilä, Esa Rahtu |
CVPR | 3 |
| 2022 | Optimal Correction Cost for Object Detection EvaluationabstractMean Average Precision (mAP) is the primary evaluation measure for object detection. Although object detection has a broad range of applications, mAP evaluates detectors in terms of the performance of ranked instance retrieval. Such the assumption for the evaluation task does not suit some downstream tasks. To alleviate the gap between downstream tasks and the evaluation scenario, we propose Optimal Correction Cost (OC-cost), which assesses detection accuracy at image level. OC-cost computes the cost of correcting detections to ground truths as a measure of accuracy. The cost is obtained by solving an optimal transportation problem between the detections and the ground truths. Unlike mAp, OC-cost is designed to penalize false positive and false negative detections properly, and every image in a dataset is treated equally. Our experimental result validates that OCscost has better agreement with human preference than a ranking-based measure, i.e., mAP for a single image. We also show that detectors' rankings by OC-cost are more consistent on different data splits than mAP. Our goal is not to replace mAP with OC-cost but provide an additional tool to evaluate detectors from another aspect. To help future researchers and developers choose a target measure, we provide a series of experiments to clarify how mAP and OC-cost differ. Mayu Otani, Riku Togashi, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Shin'ichi Satoh 0001 |
CVPR | 4 |
| 2022 | AxIoU: An Axiomatically Justified Measure for Video Moment RetrievalabstractEvaluation measures have a crucial impact on the direction of research. Therefore, it is of utmost importance to develop appropriate and reliable evaluation measures for new applications where conventional measures are not well suited. Video Moment Retrieval (VMR) is one such application, and the current practice is to use R@K,$\theta$for evaluating VMR systems. However, this measure has two disadvantages. First, it is rank-insensitive: It ignores the rank positions of successfully localised moments in the top-K ranked list by treating the list as a set. Second, it binarizes the Intersection over Union (IoU) of each retrieved video moment using the threshold$\theta$and thereby ignoring fine-grained localisation quality of ranked moments. We propose an alternative measure for evaluating VMR, called Average Max IoU (AxIoU), which is free from the above two problems. We show that AxIoU satisfies two important axioms for VMR evaluation, namely, Invariance against Redundant Moments and Monotonicity with respect to the Best Moment, and also that R@ K,$\theta$satisfies the first axiom only. We also empirically examine how Ax-IoU agrees with R@K,$\theta$, as well as its stability with respect to change in the test data and human-annotated temporal boundaries. Riku Togashi, Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Tetsuya Sakai |
CVPR | 4 |
| 2022 | Bridging the Gap Between Image Coding for Machines and HumansabstractImage coding for machines (ICM) aims at reducing the bitrate required to represent an image while minimizing the drop in machine vision analysis accuracy. In many use cases, such as surveillance, it is also important that the visual quality is not drastically deteriorated by the compression process. Recent works on using neural network (NN) based ICM codecs have shown significant coding gains against traditional methods; however, the decompressed images, especially at low bitrates, often contain checkerboard artifacts. We propose an effective decoder finetuning scheme based on adversarial training to significantly enhance the visual quality of ICM codecs, while preserving the machine analysis accuracy, without adding extra bitcost or parameters at the inference phase. The results show complete removal of the checkerboard artifacts at the negligible cost of −1.6% relative change in task performance score. In the cases where some amount of artifacts is tolerable, such as when machine consumption is the primary target, this technique can enhance both pixel-fidelity and feature-fidelity scores without losing task performance. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ICIP | 8 |
| 2022 | Long-term Visual Place RecognitionabstractIn this work, we study the long-term performance of visual place recognition in urban outdoor environment. A long-term benchmark is constructed from the Oxford RobotCar dataset. It contains sequences of the same route traversed over a period of approx. 500 days. We carefully selected three gallery sequences, one training sequence and 15 query sequences that cover different seasons, times of day and weather. The RobotCar sequences from the first half year have several problems, for example, only partial routes and inaccurate location data. We circumvent these problems by reversing the time. In the benchmark dataset the gallery and training images are the latest and the query sequences go gradually back in time. Our experiments provide the following findings. 1) the selected gallery sequence has strong impact on performance, and 2) additional training sequences help to mitigate differences between the gallery sequences. In addition, results indicate that 3) there is a long-term trend of performance degradation over time. The degradation can be quantified as about 6 percentage points per 100 days and, therefore, the initial performance of 40% eventually drops below 20% at the end. Farid Alijani, Jukka Peltomäki, Jussi Puura, Heikki Huttunen, Joni-Kristian Kämäräinen, Esa Rahtu |
ICPR | 6 |
| 2022 | TPSAD: Learning to Detect and Localize Anomalies With Thin Plate Spline TransformationabstractWe present a self-supervised learning approach with a novel proxy task, based on thin-plate spline transformation, for detecting and localizing anomalies in images. The self-supervised model, referred as TPSAD, is firstly optimized to classify normal examples from the artificially anomalous ones which are created by a new data augmentation technique that applies random thin-plate spline transformation to a patch of an image, selected by the Canny edge detector. Then, the last layer representations of the model are utilized for detecting anomalies with the Gaussian density estimator technique, while the middle layer representations are used for localizing anomalies. By assessing the proposed method on the MVTec dataset, we discover that not only can it detect anomalous images and localize irregularities properly, but also it is computationally efficient in both training and testing stages, compared to previous methods. Moreover, the method is robust to images containing unaligned objects due to the usage of the Canny edge algorithm in proxy task learning. Lastly, high performance in addition to low computational cost makes our method a good candidate for image anomaly detection in industrial applications. Afshin Dini, Esa Rahtu |
ICPR | 2 |
| 2022 | The Lottery Ticket Adaptation for Neural Video CodingabstractRecently, learning based video compression methods have attracted increasing attention. However, most learning based video codecs are not adaptive to different video contents. Though adaptation at inference time is a solution to tackle this issue, adapting all the codec’s parameters is computationally expensive and brings heavy bitrate overhead. The recently proposed Lottery Ticket Hypothesis (LTH) states that an over-parameterized neural network contains smaller subnetworks (winning tickets) that can match the performance of the original network. In this paper, we present a novel lottery-ticket adaptation technique on decoder-side multiplicative parameters of a neural network, transferring the concept of winning lottery tickets to video compression tasks. At inference time, the winning multiplicative parameters are overfitted, compressed, and signaled together with encoded frames for decoding. We show that our approach outperforms the Versatile Video Coding (VVC) standard in the Multiscale Structural Similarity (MS-SSIM) at a low bitrate on both the UVG and JVET sequences. To the best of our knowledge, this is the first attempt to apply LTH in the video compression domain. Also, this is the first published end-to-end learned video codec working directly on YUV format, which outperforms VVC on UVG and JVET datasets in MS-SSIM. Nannan Zou, Francesco Cricri, Honglei Zhang 0001, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 6 |
| 2022 | Visually Guided Sound Source Separation and Localization using Self-Supervised Motion RepresentationsabstractIn this paper, we perform audio-visual sound source separation, i.e. to separate component audios from a mixture based on the videos of sound sources. Moreover, we aim to pinpoint the source location in the input video sequence. Recent works have shown impressive audio-visual separation results when using prior knowledge of the source type (e.g. human playing instrument) and pre-trained motion detectors (e.g. keypoints or optical flows). However, at the same time, the models are limited to a certain application domain. In this paper, we address these limitations and make the following contributions: i) we propose a two-stage architecture, called Appearance and Motion network (AM-net), where the stages specialise to appearance and motion cues, respectively. The entire system is trained in a self-supervised manner; ii) we introduce an Audio-Motion Embedding (AME) framework to explicitly represent the motions that related to sound; iii) we propose an audio-motion transformer architecture for audio and motion feature fusion; iv) we demonstrate state-of-the-art performance on two challenging datasets (MUSIC-21 and AVE) despite the fact that we do not use any pre-trained keypoint detectors or optical flow estimators. Project page: https://lyzhu.github.io/self-supervised-motion-representations Lingyu Zhu 0001, Esa Rahtu |
WACV | 2 |
| 2022 | V-SlowFast Network for Efficient Visual Sound SeparationabstractThe objective of this paper is to perform visual sound separation: i) we study visual sound separation on spectrograms of different temporal resolutions; ii) we propose a new light yet efficient three-stream framework V-SlowFast that operates on Visual frame, Slow spectrogram, and Fast spectrogram. The Slow spectrogram captures the coarse temporal resolution while the Fast spectrogram contains the fine-grained temporal resolution; iii) we introduce two contrastive objectives to encourage the network to learn discriminative visual features for separating sounds; iv) we propose an audio-visual global attention module for audio and visual feature fusion; v) the introduced V-SlowFast model outperforms previous state-of-the-art in single-frame based visual sound separation on small- and large-scale datasets: MUSIC-21, AVE, and VGG-Sound. We also propose a small V-SlowFast architecture variant, which achieves 74.2% reduction in the number of model parameters and 81.4% reduction in GMACs compared to the previous multi-stage models. Project page: https://ly-zhu.github.io/V-SlowFast Lingyu Zhu 0001, Esa Rahtu |
WACV | 2 |
| 2022 | Lightweight Monocular Depth with a Novel Neural Architecture Search MethodabstractThis paper presents a novel neural architecture search method, called LiDNAS, for generating lightweight monocular depth estimation models. Unlike previous neural architecture search (NAS) approaches, where finding optimized networks is computationally demanding, the introduced novel Assisted Tabu Search leads to efficient architecture exploration. Moreover, we construct the search space on a pre-defined backbone network to balance layer diversity and search space size. The LiDNAS method outperforms the state-of-the-art NAS approach, proposed for disparity and depth estimation, in terms of search efficiency and output model performance. The LiDNAS optimized models achieve result superior to compact depth estimation state-of-the-art on NYU-Depth-v2, KITTI, and ScanNet, while being 7%-500% more compact in size, i.e the number of model parameters. Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
WACV | 4 |
| 2022 | HybVIO: Pushing the Limits of Real-time Visual-inertial OdometryabstractWe present HybVIO, a novel hybrid approach for combining filtering-based visual-inertial odometry (VIO) with optimization-based SLAM. The core of our method is highly robust, independent VIO with improved IMU bias modeling, outlier rejection, stationarity detection, and feature track selection, which is adjustable to run on embedded hardware. Long-term consistency is achieved with a loosely-coupled SLAM module. In academic benchmarks, our solution yields excellent performance in all categories, especially in the real-time use case, where we outperform the current state-of-the-art. We also demonstrate the feasibility of VIO for vehicular tracking on consumer-grade hardware using a custom dataset, and show good performance in comparison to current commercial VISLAM alternatives. Otto Seiskari, Pekka Rantalankila, Juho Kannala, Jerry Ylilammi, Esa Rahtu, Arno Solin |
WACV | 5 |
| 2022 | Single Source One Shot Reenactment using Weighted Motion from Paired Feature PointsabstractImage reenactment is a task where the target object in the source image imitates the motion represented in the driving image. One of the most common reenactment tasks is face image animation. The major challenge in the current face reenactment approaches is to distinguish between facial motion and identity. For this reason, the previous models struggle to produce high-quality animations if the driving and source identities are different (cross-person reenactment). We propose a new (face) reenactment model that learns shape-independent motion features in a self-supervised setup. The motion is represented using a set of paired feature points extracted from the source and driving images simultaneously. The model is generalised to multiple reenactment tasks including faces and non-face objects using only a single source image. The extensive experiments show that the model faithfully transfers the driving motion to the source while retaining the source identity intact. Soumya Tripathy, Juho Kannala, Esa Rahtu |
WACV | 3 |
| 2022 | FATALRead - Fooling visual speech recognition models
Anup Kumar Gupta 0001, Puneet Gupta 0002, Esa Rahtu |
Appl. Intell. | 3 |
| 2021 | Monocular Depth Estimation Primed by Salient Point Detection and Normalized Hessian LossabstractDeep neural networks have recently thrived on single image depth estimation. That being said, current developments on this topic highlight an apparent compromise between accuracy and network size. This work proposes an accurate and lightweight framework for monocular depth estimation based on a self-attention mechanism stemming from salient point detection. Specifically, we utilize a sparse set of keypoints to train a FuSaNet model that consists of two major components: Fusion-Net and Saliency-Net. In addition, we introduce a normalized Hessian loss term invariant to scaling and shear along the depth direction, which is shown to substantially improve the accuracy. The proposed method achieves state-of-the-art results on NYU-Depth-v2 and KITTI while using 3.1-38.4 times smaller model in terms of the number of parameters than baseline approaches. Experiments on the SUN-RGBD further demonstrate the generalizability of the proposed method. Lam Huynh, Matteo Pedone, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
3DV | 5 |
| 2021 | RGBD-Net: Predicting Color and Depth Images for Novel Views SynthesisabstractWe propose a new cascaded architecture for novel view synthesis, called RGBD-Net, which consists of two core components: a hierarchical depth regression network and a depth-aware generator network. The former one predicts depth maps of the target views by using adaptive depth scaling, while the latter one leverages the predicted depths and renders spatially and temporally consistent target images. In the experimental evaluation on standard datasets, RGBD-Net not only outperforms the state-of-the-art by a clear margin, but it also generalizes well to new scenes without per-scene optimization. Moreover, we show that RGBD-Net can be optionally trained without depth supervision while still retaining high-quality rendering. Thanks to the depth regression network, RGBD-Net can be also used for creating dense 3D point clouds that are more accurate than those produced by some state-of-the-art multi-view stereo methods. Phong Nguyen 0001, Animesh Karnewar, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä |
3DV | 4 |
| 2021 | Taming Visually Guided Sound Generation
Vladimir Iashin, Esa Rahtu |
BMVC | 2 |
| 2021 | Image Coding For Machines: an End-To-End Learned ApproachabstractOver recent years, deep learning-based computer vision systems have been applied to images at an ever-increasing pace, oftentimes representing the only type of consumption for those images. Given the dramatic explosion in the number of images generated per day, a question arises: how much better would an image codec targeting machine-consumption perform against state-of-the-art codecs targeting human-consumption? In this paper, we propose an image codec for machines which is neural network (NN) based and end-to-end learned. In particular, we propose a set of training strategies that address the delicate problem of balancing competing loss functions, such as computer vision task losses, image distortion losses, and rate loss. Our experimental results show that our NN-based codec outperforms the state-of-the-art Versa-tile Video Coding (VVC) standard on the object detection and instance segmentation tasks, achieving -37.87% and -32.90% of BD-rate gain, respectively, while being fast thanks to its compact size. To the best of our knowledge, this is the first end-to-end learned machine-targeted image codec. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Esa Rahtu |
ICASSP | 5 |
| 2021 | Boosting Monocular Depth Estimation with Lightweight 3D Point FusionabstractIn this paper, we propose enhancing monocular depth estimation by adding 3D points as depth guidance. Unlike existing depth completion methods, our approach performs well on extremely sparse and unevenly distributed point clouds, which makes it agnostic to the source of the 3D points. We achieve this by introducing a novel multi-scale 3D point fusion network that is both lightweight and efficient. We demonstrate its versatility on two different depth estimation problems where the 3D points have been acquired with conventional structure-from-motion and Li-DAR. In both cases, our network performs on par with state-of-the-art depth completion methods and achieves significantly higher accuracy when only a small number of points is used while being more compact in terms of the number of parameters. We show that our method outperforms some contemporary deep learning based multi-view stereo and structure-from-motion methods both in accuracy and in compactness. Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
ICCV | 4 |
| 2021 | Learned Image Coding for Machines: A Content-Adaptive ApproachabstractToday, according to the Cisco Annual Internet Report (2018-2023), the fastest-growing category of Internet traffic is machine-to-machine communication. In particular, machine-to-machine communication of images and videos represents a new challenge and opens up new perspectives in the context of data compression. One possible solution approach consists of adapting current human-targeted image and video coding standards to the use case of machine consumption. Another approach consists of developing completely new compression paradigms and architectures for machine-to-machine communications. In this paper, we focus on image compression and present an inference-time content-adaptive fine-tuning scheme that optimizes the latent representation of an end-to-end learned image codec, aimed at improving the compression efficiency for machine-consumption. The conducted experiments targeting instance segmentation task network show that our online finetuning brings an average bitrate saving (BD-rate) of -3.66% with respect to our pretrained image codec. In particular, at low bitrate points, our proposed method results in a significant bitrate saving of -9.85%. Overall, our pretrained-and-then-finetuned system achieves - 30.54% BD-rate over the state-of-the-art image/video codec Versatile Video Coding (VVC) on instance segmentation. Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Esa Rahtu |
ICME | 6 |
| 2021 | Evaluation of Long-term LiDAR Place RecognitionabstractWe compare a state-of-the-art deep image retrieval and a deep place recognition method for place recognition using LiDAR data. Place recognition aims to detect previously visited locations and thus provides an important tool for navigation, mapping, and localisation. Experimental comparisons are conducted using challenging outdoor and indoor datasets, Oxford Radar RobotCar and COLD, in the "long-term" setting where the test conditions differ substantially from the training and gallery data. Based on our results the image retrieval methods using LiDAR depth images can achieve accurate localization (the single best match recall 80%) within 5.00 m in urban outdoors. In office indoors the comparable accuracy is 50 cm but is more sensitive to changes in the environment. Jukka Peltomäki, Farid Alijani, Jussi Puura, Heikki Huttunen, Esa Rahtu, Joni-Kristian Kämäräinen |
IROS | 5 |
| 2021 | Learned Enhancement Filters for Image Coding for MachinesabstractMachine-To-Machine (M2M) communication applications and use cases, such as object detection and instance segmentation, are becoming mainstream nowadays. As a consequence, majority of multimedia content is likely to be consumed by machines in the coming years. This opens up new challenges on efficient compression of this type of data. Two main directions are being explored in the literature, one being based on existing traditional codecs, such as the Versatile Video Coding (VVC) standard, that are optimized for human-targeted use cases, and another based on end-to-end trained neural networks. However, traditional codecs have significant benefits in terms of interoperability, real-time decoding, and availability of hardware implementations over end-to-end learned codecs. Therefore, in this paper, we propose learned post-processing filters that are targeted for enhancing the performance of machine vision tasks for images reconstructed by the VVC codec. The proposed enhancement filters provide significant improvements on the target tasks compared to VVC coded images. The conducted experiments show that the proposed post-processing filters provide about 45% and 49% Bjøntegaard Delta Rate gains over VVC in instance segmentation and object detection tasks, respectively. Jukka I. Ahonen, Ramin Ghaznavi Youvalari, Nam Le 0003, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Miska M. Hannuksela, Esa Rahtu |
ISM | 8 |
| 2021 | Content-adaptive convolutional neural network post-processing filterabstractNeural Network (NN)-based coding techniques are being developed for hybrid video coding schemes, such as the Versatile Video Coding (VVC) standard. In-loop filters and postprocessing filters are two types of coding tools that aim to improve the visual quality of the reconstructed content. These tools are usually trained on large video or image datasets with varying content, but they are rarely adaptive to different content types. This problem is addressed with the proposed content-adaptive Convolutional Neural Network (CNN) post-processing filter. The proposed approach is content-adaptive in two ways. Firstly, a relatively simple CNN is pre-trained on a general video dataset and then fine-tuned on the video to be coded. Since only the bias terms of the CNN are fine-tuned, the signalling overhead is reduced. Secondly, a scaling factor indicates the influence of the CNN post-processing filter on the final reconstruction. The CNN post-processing filter is evaluated on top of VVC Test Model (VTM) 11.0 with NN-based Video Coding (NNVC) 1.0 and, overall, it can save 2.37% (Y), 3.63% (U), 2.24% (V) Bjøntegaard Delta rate (BD-rate) in the Random Access (RA) configuration. María Santamaría 0001, Yat-Hong Lam, Francesco Cricri, Jani Lainema, Ramin Ghaznavi Youvalari, Honglei Zhang 0001, Miska M. Hannuksela, Esa Rahtu, Moncef Gabbouj |
ISM | 8 |
| 2021 | Enhancing Image Coding for Machines with Compressed Feature ResidualsabstractAs computer vision technologies have tremendously improved over the last decade, videos and images are often consumed by machines instead of humans which are the main target for traditional video codecs. In many use cases, although machines are the main consumers, human involvement is also required, or even mandatory. In this paper, we propose a novel image coding technique targeted for machines, while maintaining the capability for human consumption. Our proposed codec generates two bitstreams: one bitstream from a traditional codec, referred to as human bitstream, optimized for human consumption; the other bitstream, referred to as machine bitstream, generated from an end-to-end learned neural network-based codec and optimized for machine tasks. Instead of working on the image domain, the proposed machine bitstream is derived from feature residuals – the difference between the features extracted from the input image and the features extracted from the reconstructed image generated by the traditional codec. With the help of the machine bitstream, we can significantly improve machine task performance in the low bitrate range. Our system beats the state-of-the-art traditional codec, the Versatile Video Coding (VVC/H.266), achieving −40.5% in Bjontegaard delta bitrate reduction on average for bitrates up to 0.07 BPP. Joni Seppälä, Honglei Zhang 0001, Nam Le 0003, Ramin Ghaznavi Youvalari, Francesco Cricri, Hamed Rezazadegan Tavakoli, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 9 |
| 2021 | Adaptation and Attention for Neural Video CodingabstractNeural image coding represents now the state-of-the-art image compression approach. However, a lot of work is still to be done in the video domain. In this work, we propose an end-to-end learned video codec that introduces several architectural novelties as well as training novelties, revolving around the concepts of adaptation and attention. Our codec is organized as an intra-frame codec paired with an inter-frame codec. As one architectural novelty, we propose to train the inter-frame codec model to adapt the motion estimation process based on the resolution of the input video. A second architectural novelty is a new neural block that combines concepts from split-attention based neural networks and from DenseNets. Finally, we propose to overfit a set of decoder-side multiplicative parameters at inference time. Through ablation studies and comparisons to prior art, we show the benefits of our proposed techniques in terms of coding gains. We compare our codec to VVC/H.266 and RLVC, which represent the state-of-the-art traditional and end-to-end learned codecs, respectively, and to the top performing end-to-end learned approach in 2021 CLIC competition, E2E_T_OL. Our codec clearly outperforms E2E_T_OL, and compare favorably to VVC and RLVC in some settings. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Ramin Ghaznavi Youvalari, Hamed Rezazadegan Tavakoli, Jani Lainema, Emre Aksu, Miska M. Hannuksela, Esa Rahtu |
ISM | 9 |
| 2021 | Towards a Real-Time Facial Analysis SystemabstractFacial analysis is an active research area in computer vision, with many practical applications. Most of the existing studies focus on addressing one specific task and maximizing its performance. For a complete facial analysis system, one needs to solve these tasks efficiently to ensure a smooth experience. In this work, we present a system-level design of a real-time facial analysis system. With a collection of deep neural networks for object detection, classification, and regression, the system recognizes age, gender, facial expression, and facial similarity for each person that appears in the camera view. We investigate the parallelization and interplay of individual tasks. Results on common off-the-shelf architecture show that the system’s accuracy is comparable to the state-of-the-art methods, and the recognition speed satisfies real-time requirements. Moreover, we propose a multitask network for jointly predicting the first three attributes, i.e., age, gender, and facial expression. Source code and trained models are available at https://github.com/mahehu/TUT-live-age-estimator. Bishwo Adhikari, Xingyang Ni, Esa Rahtu, Heikki Huttunen |
MMSP | 3 |
| 2021 | Supervised Fine-tuning Evaluation for Long-term Visual Place RecognitionabstractIn this paper, we present a comprehensive study on the utility of deep convolutional neural networks with two state-of-the-art pooling layers which are placed after convolutional layers and fine-tuned in an end-to-end manner for visual place recognition task in challenging conditions, including seasonal and illumination variations. We compared extensively the performance of deep learned global features with three different loss functions, e.g. triplet, contrastive and ArcFace, for learning the parameters of the architectures in terms of fraction of the correct matches during deployment. To verify effectiveness of our results, we utilized two real world datasets in place recognition, both indoor and outdoor. Our investigation demonstrates that fine tuning architectures with ArcFace loss in an end-to-end manner outperforms other two losses by approximately 1 ~ 4 % in outdoor and 1 ~ 2 % in indoor datasets, given certain thresholds, for the visual place recognition tasks. Farid Alijani, Esa Rahtu |
MMSP | 2 |
| 2021 | Supervised Fine-tuning Evaluation for Long-term Visual Place RecognitionabstractIn this paper, we present a comprehensive study on the utility of deep convolutional neural networks with two state-of-the-art pooling layers which are placed after convolutional layers and fine-tuned in an end-to-end manner for visual place recognition task in challenging conditions, including seasonal and illumination variations. We compared extensively the performance of deep learned global features with three different loss functions, e.g. triplet, contrastive and ArcFace, for learning the parameters of the architectures in terms of fraction of the correct matches during deployment. To verify effectiveness of our results, we utilized two real world datasets in place recognition, both indoor and outdoor. Our investigation demonstrates that fine tuning architectures with ArcFace loss in an end-to-end manner outperforms other two losses by approximately 1 ∼ 4 % in outdoor and 1 ∼ 2 % in indoor datasets, given certain thresholds, for the visual place recognition tasks. Farid Alijani, Esa Rahtu |
MMSP | 2 |
| 2021 | FACEGAN: Facial Attribute Controllable rEenactment GANabstractThe face reenactment is a popular facial animation method where the person's identity is taken from the source image and the facial motion from the driving image. Recent works have demonstrated high quality results by combining the facial landmark based motion representations with the generative adversarial networks. These models perform best if the source and driving images depict the same person or if the facial structures are otherwise very similar. However, if the identity differs, the driving facial structures leak to the output distorting the reenactment result. We propose a novel Facial Attribute Controllable rEenactment GAN (FACEGAN), which transfers the facial motion from the driving face via the Action Unit (AU) representation. Unlike facial landmarks, the AUs are independent of the facial structure preventing the identity leak. Moreover, AUs provide a human interpretable way to control the reenactment. FACEGAN processes background and face regions separately for optimized output quality. The extensive quantitative and qualitative comparisons show a clear improvement over the state-of-the-art in a single source reenactment task. The results are best illustrated in the reenactment video provided in the supplementary material. The source code will be made available upon publication of the paper. Soumya Tripathy, Juho Kannala, Esa Rahtu |
WACV | 3 |
| 2020 | Visually Guided Sound Source Separation Using Cascaded Opponent Filter Network
Lingyu Zhu 0001, Esa Rahtu |
ACCV (6) | 2 |
| 2020 | Sequential View Synthesis with Transformer
Phong Nguyen 0001, Lam Huynh, Esa Rahtu, Janne Heikkilä |
ACCV (4) | 3 |
| 2020 | A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer
Vladimir Iashin, Esa Rahtu |
BMVC | 2 |
| 2020 | Uncovering Hidden Challenges in Query-Based Video Moment Retrieval
Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä |
BMVC | 3 |
| 2020 | Guiding Monocular Depth Estimation Using Depth-Attention Volume
Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
ECCV (26) | 4 |
| 2020 | Deep Learning Off-the-shelf Holistic Feature Descriptors for Visual Place Recognition in Challenging ConditionsabstractIn this paper, we present a comprehensive study on the utility of deep learning feature extraction methods for visual place recognition task in three challenging conditions, appearance variation, viewpoint variation and combination of both appearance and viewpoint variation. We extensively compared the performance of convolutional neural network architectures with batch normalization layers in terms of fraction of the correct matches. These architectures are primarily trained for image classification and object detection problems and used as holistic feature descriptors for visual place recognition task. To verify effectiveness of our results, we utilized four real world datasets in place recognition. Our investigation demonstrates that convolutional neural network architectures coupled with batch normalization and trained for other tasks in computer vision outperform architectures which are specifically designed for place recognition tasks. Farid Alijani, Esa Rahtu |
MMSP | 2 |
| 2020 | L2C - Learning to Learn to CompressabstractIn this paper we present an end-to-end meta-learned system for image compression. Traditional machine learning based approaches to image compression train one or more neural network for generalization performance. However, at inference time, the encoder or the latent tensor output by the encoder can be optimized for each test image. This optimization can be regarded as a form of adaptation or benevolent overfitting to the input content. In order to reduce the gap between training and inference conditions, we propose a new training paradigm for learned image compression, which is based on meta-learning. In a first phase, the neural networks are trained normally. In a second phase, the Model-Agnostic Meta-learning approach is adapted to the specific case of image compression, where the inner-loop performs latent tensor overfitting, and the outer loop updates both encoder and decoder neural networks based on the overfitting performance. Furthermore, after meta-learning, we propose to overfit and cluster the bias terms of the decoder on training image patches, so that at inference time the optimal content-specific bias terms can be selected at encoder-side. Finally, we propose a new probability model for lossless compression, which combines concepts from both multi-scale and super-resolution probability model approaches. We show the benefits of all our proposed ideas via carefully designed experiments. Nannan Zou, Honglei Zhang 0001, Francesco Cricri, Hamed Rezazadegan Tavakoli, Jani Lainema, Miska M. Hannuksela, Emre Aksu, Esa Rahtu |
MMSP | 8 |
| 2020 | ICface: Interpretable and Controllable Face Reenactment Using GANsabstractThis paper presents a generic face animator that is able to control the pose and expressions of a given face image. The animation is driven by human interpretable control signals consisting of head pose angles and the Action Unit (AU) values. The control information can be obtained from multiple sources including external driving videos and manual controls. Due to the interpretable nature of the driving signal, one can easily mix the information between multiple sources (e.g. pose from one image and expression from another) and apply selective postproduction editing. The proposed face animator is implemented as a two stage neural network model that is learned in self-supervised manner using a large video collection. The proposed Interpretable and Controllable face reenactment network (ICface) is compared to the state-of-the-art neural network based face animation techniques in multiple tasks. The results indicate that ICface produces better visual quality, while being more versatile than most of the comparison methods. The introduced model could provide a lightweight and easy to use tool for multitude of advanced image and video editing tasks. The program code will be publicly available upon the acceptance of the paper. Soumya Tripathy, Juho Kannala, Esa Rahtu |
WACV | 3 |
| 2020 | Automated Video Face Labelling for Films and TV MaterialabstractThe objective of this work is automatic labelling of characters in TV video and movies, given weak supervisory information provided by an aligned transcript. We make five contributions: (i) a new strategy for obtaining stronger supervisory information from aligned transcripts; (ii) an explicit model for classifying background characters, based on their face-tracks; (iii) employing new ConvNet based face features, and (iv) a novel approach for labelling all face tracks jointly using linear programming. Each of these contributions delivers a boost in performance, and we demonstrate this on standard benchmarks using tracks provided by authors of prior work. As a fifth contribution, we also investigate the generalisation and strength of the features and classifiers by applying them "in the raw" on new video material where no supervisory information is used. In particular, to provide high quality tracks on those material, we propose efficient track classifiers to remove false positive tracks by the face tracker. Overall we achieve a dramatic improvement over the state of the art on both TV series and film datasets, and almost saturate performance on some benchmarks. Omkar M. Parkhi, Esa Rahtu, Qiong Cao, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Rethinking the Evaluation of Video SummariesabstractVideo summarization is a technique to create a short skim of the original video while preserving the main stories/content. There exists a substantial interest in automatizing this process due to the rapid growth of the available material. The recent progress has been facilitated by public benchmark datasets, which enable easy and fair comparison of methods. Currently the established evaluation protocol is to compare the generated summary with respect to a set of reference summaries provided by the dataset. In this paper, we will provide in-depth assessment of this pipeline using two popular benchmark datasets. Surprisingly, we observe that randomly generated summaries achieve comparable or better performance to the state-of-the-art. In some cases, the random summaries outperform even the human generated summaries in leave-one-out experiments. Moreover, it turns out that the video segmentation, which is often considered as a fixed pre-processing method, has the most significant impact on the performance measure. Based on our observations, we propose alternative approaches for assessing the importance scores as well as an intuitive visualization of correlation between the estimated scoring and human annotations. Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä |
CVPR | 3 |
| 2019 | CIIDefence: Defeating Adversarial Attacks by Fusing Class-Specific Image Inpainting and Image DenoisingabstractThis paper presents a novel approach for protecting deep neural networks from adversarial attacks, i.e., methods that add well-crafted imperceptible modifications to the original inputs such that they are incorrectly classified with high confidence. The proposed defence mechanism is inspired by the recent works mitigating the adversarial disturbances by the means of image reconstruction and denoising. However, unlike the previous works, we apply the reconstruction only for small and carefully selected image areas that are most influential to the current classification outcome. The selection process is guided by the class activation map responses obtained for multiple top-ranking class labels. The same regions are also the most prominent for the adversarial perturbations and hence most important to purify. The resulting inpainting task is substantially more tractable than the full image reconstruction, while still being able to prevent the adversarial attacks. Furthermore, we combine the selective image inpainting with wavelet based image denoising to produce a non differentiable layer that prevents attacker from using gradient backpropagation. Moreover, the proposed nonlinearity cannot be easily approximated with simple differentiable alternative as demonstrated in the experiments with Backward Pass Differentiable Approximation (BPDA) attack. Finally, we experimentally show that the proposed Class-specific Image Inpainting Defence (CIIDefence) is able to withstand several powerful adversarial attacks including the BPDA. The obtained results are consistently better compared to the other recent defence approaches. Esa Rahtu |
ICCV | 2 |
| 2019 | DGC-Net: Dense Geometric Correspondence NetworkabstractThis paper addresses the challenge of dense pixel correspondence estimation between two images. This problem is closely related to optical flow estimation task where ConvNets (CNNs) have recently achieved significant progress. While optical flow methods produce very accurate results for the small pixel translation and limited appearance variation scenarios, they hardly deal with the strong geometric transformations that we consider in this work. In this paper, we propose a coarse-to-fine CNN-based framework that can leverage the advantages of optical flow approaches and extend them to the case of large transformations providing dense and subpixel accurate estimates. It is trained on synthetic transformations and demonstrates very good performance to unseen, realistic, data. Further, we apply our method to the problem of relative camera pose estimation and demonstrate that the model outperforms existing dense approaches. Iaroslav Melekhov, Aleksei Tiulpin, Torsten Sattler, Marc Pollefeys, Esa Rahtu, Juho Kannala |
WACV | 5 |
| 2019 | Digging Deeper Into Egocentric Gaze PredictionabstractThis paper digs deeper into factors that influence egocentric gaze. Instead of training deep models for this purpose in a blind manner, we propose to inspect factors that contribute to gaze guidance during daily tasks. Bottom-up saliency and optical flow are assessed versus strong spatial prior baselines. Task-specific cues such as vanishing point, manipulation point, and hand regions are analyzed as representatives of top-down information. We also look into the contribution of these factors by investigating a simple recurrent neural model for ego-centric gaze prediction. First, deep features are extracted for all input video frames. Then, a gated recurrent unit is employed to integrate information over time and to predict the next fixation. We also propose an integrated model that combines the recurrent model with several top-down and bottom-up cues. Extensive experiments over multiple datasets reveal that (1) spatial biases are strong in egocentric videos, (2) bottom-up saliency models perform poorly in predicting gaze and underperform spatial biases, (3) deep features perform better compared to traditional features, (4) as opposed to hand regions, the manipulation point is a strong influential cue for gaze prediction, (5) combining the proposed recurrent model with bottom-up cues, vanishing points and, in particular, manipulation point results in the best gaze prediction accuracy over egocentric videos, (6) the knowledge transfer works best for cases where the tasks or sequences are similar, and (7) task and activity recognition can benefit from gaze prediction. Our findings suggest that (1) there should be more emphasis on hand-object interaction and (2) the egocentric vision community should consider larger datasets including diverse stimuli and more subjects. Hamed Rezazadegan Tavakoli, Esa Rahtu, Juho Kannala, Ali Borji |
WACV | 2 |
| 2018 | Learning Image-to-Image Translation Using Paired and Unpaired Training Samples
Soumya Tripathy, Juho Kannala, Esa Rahtu |
ACCV (2) | 3 |
| 2018 | ADVIO: An Authentic Dataset for Visual-Inertial Odometry
Santiago Cortés Reina, Arno Solin, Esa Rahtu, Juho Kannala |
ECCV (10) | 3 |
| 2018 | Inertial Odometry on Handheld SmartphonesabstractBuilding a complete inertial navigation system using the limited quality data provided by current smartphones has been regarded challenging, if not impossible. This paper shows that by careful crafting and accounting for the weak information in the sensor samples, smartphones are capable of pure inertial navigation. We present a probabilistic approach for orientation and use-case free inertial odometry, which is based on double-integrating rotated accelerations. The strength of the model is in learning additive and multiplicative IMU biases online. We are able to track the phone position, velocity, and pose in realtime and in a computationally lightweight fashion by solving the inference with an extended Kalman filter. The information fusion is completed with zero-velocity updates (if the phone remains stationary), altitude correction from barometric pressure readings (if available), and pseudo-updates constraining the momentary speed. We demonstrate our approach using an iPad and iPhone in several indoor dead-reckoning applications and in a measurement tool setup. Arno Solin, Santiago Cortés Reina, Esa Rahtu, Juho Kannala |
FUSION | 3 |
| 2018 | Bottom-Up Attention Guidance for Recurrent Image RecognitionabstractThis paper presents a recurrent neural network architecture, guided by the bottom-up attention, for the recognition task. The proposed architecture processes an input image as a sequence of selectively chosen patches. The patches are chosen from the salient regions of the input image. Using human driven saliency maps from gaze, the benefit of such a selection process is first shown. Next, the performance of computational models of bottom-up attention are assessed as alternative to human attention. Hamed Rezazadegan Tavakoli, Ali Borji, Rao Muhammad Anwer, Esa Rahtu, Juho Kannala |
ICIP | 4 |
| 2018 | PIVO: Probabilistic Inertial-Visual Odometry for Occlusion-Robust NavigationabstractThis paper presents a novel method for visual-inertial odometry. The method is based on an information fusion framework employing low-cost IMU sensors and the monocular camera in a standard smartphone. We formulate a sequential inference scheme, where the IMU drives the dynamical model and the camera frames are used in coupling trailing sequences of augmented poses. The novelty in the model is in taking into account all the cross-terms in the updates, thus propagating the inter-connected uncertainties throughout the model. Stronger coupling between the inertial and visual data sources leads to robustness against occlusion and feature-poor environments. We demonstrate results on data collected with an iPhone and provide comparisons against the Tango device and using the EuRoC data set. Arno Solin, Santiago Cortés Reina, Esa Rahtu, Juho Kannala |
WACV | 3 |
| 2018 | Summarization of User-Generated Sports Video by Using Deep Action Recognition FeaturesabstractAutomatically generating a summary of a sports video poses the challenge of detecting interesting moments, or highlights, of a game. Traditional sports video summarization methods leverage editing conventions of broadcast sports video that facilitate the extraction of high-level semantics. However, user-generated videos are not edited and, thus, traditional methods are not suitable to generate a summary. In order to solve this problem, this paper proposes a novel video summarization method that uses players' actions as a cue to determine the highlights of the original video. A deep neural-network-based approach is used to extract two types of action-related features and to classify video segments into interesting or uninteresting parts. The proposed method can be applied to any sports in which games consist of a succession of actions. Especially, this paper considers the case of Kendo (Japanese fencing) as an example of a sport to evaluate the proposed method. The method is trained using Kendo videos with ground truth labels that indicate the video highlights. The labels are provided by annotators possessing a different experience with respect to Kendo to demonstrate how the proposed method adapts to different needs. The performance of the proposed method is compared with several combinations of different features, and the results show that it outperforms previous summarization methods. Antonio Tejero-de-Pablos, Yuta Nakashima, Tomokazu Sato, Naokazu Yokoya, Marko Linna, Esa Rahtu |
IEEE Trans. Multim. | 6 |
| 2017 | Relative Camera Pose Estimation Using Convolutional Neural Networks
Iaroslav Melekhov, Juha Ylioinas, Juho Kannala, Esa Rahtu |
ACIVS | 4 |
| 2017 | Exploiting inter-image similarity and ensemble of extreme learners for fixation prediction using deep features
Hamed Rezazadegan Tavakoli, Ali Borji, Jorma Laaksonen, Esa Rahtu |
Neurocomputing | 4 |
| 2016 | Video Summarization Using Deep Semantic Features
Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Naokazu Yokoya |
ACCV (5) | 3 |
| 2016 | Robust loop closures for scene reconstruction by combining odometry and visual correspondencesabstractGiven an image sequence and odometry from a moving camera, we propose a batch-based approach for robust reconstruction of scene structure and camera motion. A key part of our method is robust loop closure disambiguation. First, a structure-from-motion pipeline is used to get a set of candidate feature correspondences and the respective triangulated 3D landmarks. Thereafter, the compatibility of each correspondence constraint and the odometry is evaluated in a bundle-adjustment optimization, where only compatible constraints affect. Our approach is evaluated using data from a Google Tango device. The results show that it produces better reconstructions than the device's built-in software or a state-of-the-art pose-graph formulation. Zakaria Laskar, Sami Huttunen, Daniel Herrera C., Esa Rahtu, Juho Kannala |
ICIP | 4 |
| 2016 | Siamese network features for image matchingabstractFinding matching images across large datasets plays a key role in many computer vision applications such as structure-from-motion (SfM), multi-view 3D reconstruction, image retrieval, and image-based localisation. In this paper, we propose finding matching and non-matching pairs of images by representing them with neural network based feature vectors, whose similarity is measured by Euclidean distance. The feature vectors are obtained with convolutional neural networks which are learnt from labeled examples of matching and non-matching image pairs by using a contrastive loss function in a Siamese network architecture. Previously Siamese architecture has been utilised in facial image verification and in matching local image patches, but not yet in generic image retrieval or whole-image matching. Our experimental results show that the proposed features improve matching performance compared to baseline features obtained with networks which are trained for image classification task. The features generalize well and improve matching of images of new landmarks which are not seen at training time. This is despite the fact that the labeling of matching and non-matching pairs is imperfect in our training data. The results are promising considering image retrieval applications, and there is potential for further improvement by utilising more training image pairs with more accurate ground truth labels. Iaroslav Melekhov, Juho Kannala, Esa Rahtu |
ICPR | 3 |
| 2015 | Online Face Recognition System Based on Local Binary Patterns and Facial Landmark Tracking
Marko Linna, Juho Kannala, Esa Rahtu |
ACIVS | 3 |
| 2015 | Adaptive Kalman filtering and smoothing for gravitation tracking in mobile systemsabstractThis paper is concerned with inertial-sensor-based tracking of the gravitation direction in mobile devices such as smartphones. Although this tracking problem is a classical one, choosing a good state-space for this problem is not entirely trivial. Even though for many other orientation related tasks a quaternion-based representation tends to work well, for gravitation tracking their use is not always advisable. In this paper we present a convenient linear quaternion-free state-space model for gravitation tracking. We also discuss the efficient implementation of the Kalman filter and smoother for the model. Furthermore, we propose an adaption mechanism for the Kalman filter which is able to filter out shot-noises similarly as has been proposed in context of adaptive and robust Kalman filtering. We compare the proposed approach to other approaches using measurement data collected with a smartphone. Simo Särkkä, Ville Tolvanen, Juho Kannala, Esa Rahtu |
IPIN | 4 |
| 2014 | Generating Object Segmentation Proposals Using Global and Local SearchabstractWe present a method for generating object segmentation proposals from groups of superpixels. The goal is to propose accurate segmentations for all objects of an image. The proposed object hypotheses can be used as input to object detection systems and thereby improve efficiency by replacing exhaustive search. The segmentations are generated in a class-independent manner and therefore the computational cost of the approach is independent of the number of object classes. Our approach combines both global and local search in the space of sets of superpixels. The local search is implemented by greedily merging adjacent pairs of superpixels to build a bottom-up segmentation hierarchy. The regions from such a hierarchy directly provide a part of our region proposals. The global search provides the other part by performing a set of graph cut segmentations on a superpixel graph obtained from an intermediate level of the hierarchy. The parameters of the graph cut problems are learnt in such a manner that they provide complementary sets of regions. Experiments with Pascal VOC images show that we reach state-of-the-art with greatly reduced computational cost. Pekka Rantalankila, Juho Kannala, Esa Rahtu |
CVPR | 3 |
| 2014 | Understanding Objects in Detail with Fine-Grained AttributesabstractWe study the problem of understanding objects in detail, intended as recognizing a wide array of fine-grained object attributes. To this end, we introduce a dataset of 7, 413 airplanes annotated in detail with parts and their attributes, leveraging images donated by airplane spotters and crowd-sourcing both the design and collection of the detailed annotations. We provide a number of insights that should help researchers interested in designing fine-grained datasets for other basic level categories. We show that the collected data can be used to study the relation between part detection and attribute prediction by diagnosing the performance of classifiers that pool information from different parts of an object. We note that the prediction of certain attributes can benefit substantially from accurate part detection. We also show that, differently from previous results in object detection, employing a large number of part templates can improve detection accuracy at the expenses of detection speed. We finally propose a coarse-to-fine approach to speed up detection through a hierarchical cascade algorithm. Andrea Vedaldi, Siddharth Mahendran, Stavros Tsogkas, Subhransu Maji, Ross B. Girshick, Juho Kannala, Esa Rahtu, Iasonas Kokkinos, Matthew B. Blaschko, David J. Weiss, Ben Taskar, Karen Simonyan, Naomi Saphra, Sammy Mohamed |
CVPR | 7 |
| 2014 | Emotional Valence Recognition, Analysis of Salience and Eye MovementsabstractThis paper studies the performance of recorded eye movements and computational visual attention models (i.e. saliency models) in the recognition of emotional valence of an image. In the first part of this study, it employs eye movement data (fixation & saccade) to build image content descriptors and use them with support vector machines to classify the emotional valence. In the second part, it examines if the human saliency map can be substituted with the state-of-the-art computational visual attention models in the task of valence recognition. The results indicate that the eye movement based descriptors provide significantly better performance compared to the baselines, which apply low-level visual cues (e.g. color, texture and shape). Furthermore, it will be shown that the current computational models for visual attention are not able to capture the emotional information in similar extent as the real eye movements. Hamed Rezazadegan Tavakoli, Victoria Yanulevskaya, Esa Rahtu, Janne Heikkilä, Nicu Sebe |
ICPR | 3 |
| 2013 | Spherical Center-Surround for Video Saliency Detection Using Sparse Sampling
Hamed Rezazadegan Tavakoli, Esa Rahtu, Janne Heikkilä |
ACIVS | 2 |
| 2013 | Stochastic bottom-up fixation prediction and saccade generation
Hamed Rezazadegan Tavakoli, Esa Rahtu, Janne Heikkilä |
Image Vis. Comput. | 2 |
| 2013 | Automatic Dynamic Texture Segmentation Using Local Descriptors and Optical FlowabstractA dynamic texture (DT) is an extension of the texture to the temporal domain. How to segment a DT is a challenging problem. In this paper, we address the problem of segmenting a DT into disjoint regions. A DT might be different from its spatial mode (i.e., appearance) and/or temporal mode (i.e., motion field). To this end, we develop a framework based on the appearance and motion modes. For the appearance mode, we use a new local spatial texture descriptor to describe the spatial mode of the DT; for the motion mode, we use the optical flow and the local temporal texture descriptor to represent the temporal variations of the DT. In addition, for the optical flow, we use the histogram of oriented optical flow (HOOF) to organize them. To compute the distance between two HOOFs, we develop a simple effective and efficient distance measure based on Weber's law. Furthermore, we also address the problem of threshold selection by proposing a method for determining thresholds for the segmentation method by an offline supervised statistical learning. The experimental results show that our method provides very good segmentation results compared to the state-of-the-art methods in segmenting regions that differ in their dynamics. Jie Chen 0001, Guoying Zhao 0001, Mikko Salo, Esa Rahtu, Matti Pietikäinen |
IEEE Trans. Image Process. | 4 |
| 2012 | TriCoS: A Tri-level Class-Discriminative Co-segmentation Method for Image Classification
Yuning Chai, Esa Rahtu, Victor S. Lempitsky, Luc Van Gool, Andrew Zisserman |
ECCV (1) | 2 |
| 2012 | BSIF: Binarized statistical image features
Juho Kannala, Esa Rahtu |
ICPR | 2 |
| 2012 | Local phase quantization for blur-insensitive image analysis
Esa Rahtu, Janne Heikkilä, Ville Ojansivu, Timo Ahonen |
Image Vis. Comput. | 1 |
| 2011 | Learning a category independent object detection cascadeabstractCascades are a popular framework to speed up object detection systems. Here we focus on the first layers of a category independent object detection cascade in which we sample a large number of windows from an objectness prior, and then discriminatively learn to filter these candidate windows by an order of magnitude. We make a number of contributions to cascade design that substantially improve over the state of the art: (i) our novel objectness prior gives much higher recall than competing methods, (ii) we propose objectness features that give high performance with very low computational cost, and (iii) we make use of a structured output ranking approach to learn highly effective, but inexpensive linear feature combinations by directly optimizing cascade performance. Thorough evaluation on the PASCAL VOC data set shows consistent improvement over the current state of the art, and over alternative discriminative learning strategies. Esa Rahtu, Juho Kannala, Matthew B. Blaschko |
ICCV | 1 |
| 2010 | Segmenting Salient Objects from Images and Videos
Esa Rahtu, Juho Kannala, Mikko Salo, Janne Heikkilä |
ECCV (5) | 1 |
| 2010 | Improved Blur Insensitivity for Decorrelated Local Phase QuantizationabstractThis paper presents a novel blur tolerant decor relation scheme for local phase quantization (LPQ) texture descriptor. As opposed to previous methods, the introduced model can be applied with virtually any kind of blur regardless of the point spread function. The new technique takes also into account the changes in the image characteristics originating from the blur itself. The implementation does not suffer from multiple solutions like the decor relation in original LPQ, but still retains the same run-time computational complexity. The texture classification experiments illustrate considerable improvements in the performance of LPQ descriptors in the case of blurred images and show only negligible loss of accuracy with sharp images. Janne Heikkilä, Ville Ojansivu, Esa Rahtu |
ICPR | 3 |
| 2010 | Compressing Sparse Feature Vectors Using Random Ortho-ProjectionsabstractIn this paper we investigate the usage of random ortho-projections in the compression of sparse feature vectors. The study is carried out by evaluating the compressed features in classification tasks instead of concentrating on reconstruction accuracy. In the random ortho-projection method, the mapping for the compression can be obtained without any further knowledge of the original features. This makes the approach favorable if training data is costly or impossible to obtain. The independence from the data also enables one to embed the compression scheme directly into the computation of the original features. Our study is inspired by the results in compressive sensing, which state that up to a certain compression ratio and with high probability, such projections result in no loss of information. In comparison to learning based compression, namely principal component analysis (PCA), the random projections resulted in comparable performance already at high compression ratios depending on the sparsity of the original features. Esa Rahtu, Mikko Salo, Janne Heikkilä |
ICPR | 1 |
| 2008 | Object recognition and segmentation by non-rigid quasi-dense matchingabstractIn this paper, we present a non-rigid quasi-dense matching method and its application to object recognition and segmentation. The matching method is based on the match propagation algorithm which is here extended by using local image gradients for adapting the propagation to smooth non-rigid deformations of the imaged surfaces. The adaptation is based entirely on the local properties of the images and the method can be hence used in non-rigid image registration where global geometric constraints are not available. Our approach for object recognition and segmentation is directly built on the quasi-dense matching. The quasi-dense pixel matches between the model and test images are grouped into geometrically consistent groups using a method which utilizes the local affine transformation estimates obtained during the propagation. The number and quality of geometrically consistent matches is used as a recognition criterion and the location of the matching pixels directly provides the segmentation. The experiments demonstrate that our approach is able to deal with extensive background clutter, partial occlusion, large scale and viewpoint changes, and notable geometric deformations. Juho Kannala, Esa Rahtu, Sami S. Brandt, Janne Heikkilä |
CVPR | 2 |
| 2008 | Recognition of blurred faces using Local Phase QuantizationabstractIn this paper, recognition of blurred faces using the recently introduced Local Phase Quantization (LPQ) operator is proposed. LPQ is based on quantizing the Fourier transform phase in local neighborhoods. The phase can be shown to be a blur invariant property under certain commonly fulfilled conditions. In face image analysis, histograms of LPQ labels computed within local regions are used as a face descriptor similarly to the widely used Local Binary Pattern (LBP) methodology for face image description. The experimental results on CMU PIE and FRGC 1.0.4 datasets show that the LPQ descriptor is highly tolerant to blur but still very descriptive outperforming LBP both with blurred and sharp images. Timo Ahonen, Esa Rahtu, Ville Ojansivu, Janne Heikkilä |
ICPR | 2 |
| 2008 | Rotation invariant local phase quantization for blur insensitive texture analysisabstractThis paper introduces a rotation invariant extension to the blur insensitive local phase quantization texture descriptor. The new method consists of two stages, the first of which estimates the local characteristic orientation, and the second one extracts a binary descriptor vector. Both steps of the algorithm apply the phase of the locally computed Fourier transform coefficients, which can be shown to be insensitive to centrally symmetric image blurring. The new descriptors are assessed in comparison with the well known texture descriptors, local binary patterns (LBP) and Gabor filtering. The results illustrate that the proposed method has superior performance in those cases where the image contains blur and is slightly better even with sharp images. Ville Ojansivu, Esa Rahtu, Janne Heikkilä |
ICPR | 2 |
| 2006 | A New Affine Invariant Image Transform Based on RidgeletsabstractIn this paper we present a new affine invariant image transform, based on ridgelets. The proposed transform is directly applicable to segmented image patches. The new method has some similarities with the previously proposed Multiscale Autoconvolution, but it will offer a more general framework and possibilities for variations. The obtained transform coefficients can be used in affine invariant pattern classification, and as shown in the experiments, already a small subset of them is enough for reliable recognition of complex patterns. The new method is assessed in several experiments and it is observed to perform well under many nonaffine distortions. 1 Esa Rahtu, Janne Heikkilä, Mikko Salo |
BMVC | 1 |
| 2006 | Multiscale Autoconvolution Histograms for Affine Invariant Pattern RecognitionabstractIn this paper we present a new way of producing affine invariant histograms from images. The approach is based on a probabilistic interpretation of the image function as in the multiscale autoconvolution (MSA) transform, but the histograms extract much more information of the image than traditional MSA. The new histograms can be considered as generalizations of the image gray scale histogram, encoding also the spatial information. It turns out that the proposed method can be efficiently computed using the Fast Fourier Transform, and it will be shown to have essentially the same computational load as MSA. The experiments performed indicate that the new invariants are capable of reliable classification of complex patterns, outperforming MSA and many other methods. 1 Esa Rahtu, Mikko Salo, Janne Heikkilä |
BMVC | 1 |
| 2006 | A New Convexity Measure Based on a Probabilistic Interpretation of ImagesabstractIn this paper, we present a novel convexity measure for object shape analysis. The proposed method is based on the idea of generating pairs of points from a set and measuring the probability that a point dividing the corresponding line segments belongs to the same set. The measure is directly applicable to image functions representing shapes and also to gray-scale images which approximate image binarizations. The approach introduced gives rise to a variety of convexity measures which make it possible to obtain more information about the object shape. The proposed measure turns out to be easy to implement using the Fast Fourier Transform and we will consider this in detail. Finally, we illustrate the behavior of our measure in different situations and compare it to other similar ones. Esa Rahtu, Mikko Salo, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Affine registration with multi-scale autoconvolutionabstractIn this paper we propose a novel method for the recovery of affine transformation parameters between two images. Registration is achieved without separate feature extraction by directly utilizing the intensity distribution of the images. The method can also be used for matching point sets under affine transformations. Our approach is based on the same probabilistic interpretation of the image function as the recently introduced multi-scale autoconvolution (MSA) transform. Here we describe how the framework may be used in image registration and present two variants of the method for practical implementation. The proposed method is experimented with binary and grayscale images and compared with other non-feature-based registration methods. The experiments show that the new method can efficiently align images of isolated objects and is relatively robust. Juho Kannala, Esa Rahtu, Janne Heikkilä |
ICIP (3) | 2 |
| 2005 | Affine Invariant Pattern Recognition Using Multiscale AutoconvolutionabstractThis paper presents a new affine invariant image transform called Multiscale Autoconvolution (MSA). The proposed transform is based on a probabilistic interpretation of the image function. The method is directly applicable to isolated objects and does not require extraction of boundaries or interest points, and the computational load is significantly reduced using the Fast Fourier Transform. The transform values can be used as descriptors for affine invariant pattern classification and, in this article, we illustrate their performance in various object classification tasks. As shown by a comparison with other affine invariant techniques, the new method appears to be suitable for problems where image distortions can be approximated with affine transformations. Esa Rahtu, Mikko Salo, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |