VLDB 2026 Research / reviewers in the wild / expert
Juho Kannala
dblp:47/4656
· DBLP profile ↗
112ranked-venue papers
7as first author
47since 2021 · last 2026
0000-0001-5088-4041ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 84 · 5 first-author · 34 since 2021Artificial intelligence and machine learning · 65 · 6 first-author · 29 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Predicting Video Slot Attention Queries from Random Slot-Feature PairsabstractUnsupervised video Object-Centric Learning (OCL) is promising as it enables object-level scene representation and understanding as we humans do. Mainstream video OCL methods adopt a recurrent architecture: An aggregator aggregates current video frame into object features, termed slots, under some queries; A transitioner transits current slots to queries for the next frame. This is an effective architecture but all existing implementations both (i1) neglect to incorporate next frame features, the most informative source for query prediction, and (i2) fail to learn transition dynamics, the knowledge essential for query prediction. To address these issues, we propose Random Slot-Feature pair for learning Query prediction (RandSF.Q): (t1) We design a new transitioner to incorporate both slots and features, which provides more information for query prediction; (t2) We train the transitioner to predict queries from slot-feature pairs randomly sampled from available recurrences, which drives it to learn transition dynamics. Experiments on scene representation demonstrate that our method surpass existing video OCL methods significantly, e.g., up to 10 points on object discovery, setting new state-of-the-art. Such superiority also benefits downstream tasks like scene understanding. Rongzhen Zhao, Juho Kannala, Joni Pajarinen |
AAAI | 3 |
| 2026 | Pose-Guided Geometric Refinement for Feed-Forward 3D Gaussian Splatting
Yejun Zhang 0001, Esa Rahtu, Juho Kannala |
ICPR (12) | 5 |
| 2026 | SceneShine: Illumination-aware Human Scene Gaussian Re-Splatting from Mobile Device VideoabstractStandard 3DGS falls short in precise relighting and shadowing needed to realistically integrate humans into novel environments. We bridge this gap with SceneShine, an illumination-aware framework designed for seamless composition through physically-based avatar relighting and shadow casting. Relighting human surfaces in in-the-wild videos is inherently ill-posed, often making the simultaneous disentanglement of scene lighting and BRDF properties difficult. We overcome this ambiguity by utilizing a pseudo-global light map prior to guide BRDF parameter decomposition, significantly reducing relighting artifacts. Additionally, we implement point-based ray tracing to manage human-scene occlusions and dynamically update scene colors for accurate shadow casting. We also introduce a new synthetic dataset for evaluation. Extensive experiments show that our method surpasses existing approaches in reconstruction fidelity and identity preservation while achieving highly convincing illumination-aware integration1. Xuqian Ren, Wenjia Wang 0009, Mai Ngoc Nguyen, Juho Kannala, Esa Rahtu |
WACV | 4 |
| 2026 | Tuning Qwen2.5-VL to Improve Its Web Interaction SkillsabstractRecent advances in vision–language models (VLMs) have sparked growing interest in using them to automate web tasks, yet their feasibility as independent agents that reason and act purely from visual input remains underexplored. We investigate this setting using Qwen2.5-VL-32B, one of the strongest open-source VLMs available, and focus on improving its reliability in web-based control. Through initial experimentation, we observe three key challenges: (i)~inaccurate localization of target elements, the cursor, and their relative positions, (ii)~sensitivity to instruction phrasing, and (iii)~an overoptimistic bias toward its own actions, often assuming they succeed rather than analyzing their actual outcomes. To address these issues, we fine-tune Qwen2.5-VL-32B for a basic web interaction task: moving the mouse and clicking on a page element described in natural language. Our training pipeline consists of two stages: (1)~teaching the model to determine whether the cursor already hovers over the target element or whether movement is required, and (2)~training it to execute a single command (a mouse move or a mouse click) at a time, verifying the resulting state of the environment before planning the next action. Evaluated on a custom benchmark of single-click web tasks, our approach increases success rates from 86% to 94% under the most challenging setting. Alexandra Yakovleva, Henrik Pärssinen, Harri Valpola, Juho Kannala, Alexander Ilin |
WWW | 4 |
| 2026 | Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation ModelsabstractDomain adaptation and generalization are crucial for real-world applications, such as autonomous driving and medical imaging where the model must operate reliably across environments with distinct data distributions. However, these tasks are challenging because the model needs to overcome various domain gaps caused by variations in, for example, lighting, weather, sensor configurations, and so on. Addressing domain gaps simultaneously in different modalities, known as multimodal domain adaptation and generalization, is even more challenging due to unique challenges in different modalities. Over the past few years, significant progress has been made in these areas, with applications ranging from action recognition to semantic segmentation, and more. Recently, the emergence of large-scale pre-trained multimodal foundation models, such as CLIP, has inspired numerous research studies, which leverage these models to enhance downstream adaptation and generalization. This survey summarizes recent advances in multimodal adaptation and generalization, particularly how these areas evolve from traditional approaches to foundation models. Specifically, this survey covers (1) multimodal domain adaptation, (2) multimodal test-time adaptation, (3) multimodal domain generalization, (4) domain adaptation and generalization with the help of multimodal foundation models, and (5) adaptation of multimodal foundation models. For each topic, we formally define the problem and give a thorough review of existing methods. Additionally, we analyze relevant datasets and applications, highlighting open challenges and potential future research directions. Hao Dong 0011, Moru Liu, Kaiyang Zhou, Eleni N. Chatzi, Juho Kannala, Cyrill Stachniss, Olga Fink |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | AGS-Mesh: Adaptive Gaussian Splatting and Meshing with Geometric Priors for Indoor Room Reconstruction Using SmartphonesabstractGeometric priors are often used to enhance 3D reconstruction. With many smartphones featuring low-resolution depth sensors and the prevalence of off-the-shelf monocular geometry estimators, incorporating geometric priors as regularization signals has become common in 3D vision tasks. However, the accuracy of depth estimates from mobile devices is typically poor for highly detailed geometry, and monocular estimators often suffer from poor multi-view consistency and precision. In this work, we propose an approach for joint surface depth and normal refinement of Gaussian Splatting methods for accurate 3D reconstruction of indoor scenes. We develop supervision strategies that adaptively filters low-quality depth and normal estimates by comparing the consistency of the priors during optimization. We mitigate regularization in regions where prior estimates have high uncertainty or ambiguities. Our filtering strategy and optimization design demonstrate significant improvements in both mesh estimation and novel-view synthesis for both 3D and 2D Gaussian Splatting-based methods on challenging indoor room datasets. Furthermore, we explore the use of alternative meshing strategies for finer geometry extraction. We develop a scale-aware meshing strategy inspired by TSDF and octree-based isosurface extraction, which recovers finer details from Gaussian models compared to other commonly used open-source meshing tools. Our code is released in https://xuqianren.github.io/ags_mesh_website/. Xuqian Ren, Matias Turkulainen, Jiepeng Wang 0001, Otto Seiskari, Iaroslav Melekhov, Juho Kannala, Esa Rahtu |
3DV | 6 |
| 2025 | A2-GNN: Angle-Annular GNN for Visual Descriptor-Free Camera RelocalizationabstractVisual localization involves estimating the 6-degree-of-freedom (6-DoF) camera pose within a known scene. A critical step in this process is identifying pixel-to-point correspondences between 2D query images and 3D models. Most advanced approaches currently rely on extensive visual descriptors to establish these correspondences, facing challenges in storage, privacy issues and model maintenance. Direct 2D-3D keypoint matching without visual descriptors is becoming popular as it can overcome those challenges. However, existing descriptor-free methods suffer from low accuracy or heavy computation. Addressing this gap, this paper introduces the Angle-Annular Graph Neural Network (A2-GNN), a simple approach that efficiently learns robust geometric structural representations with annular feature extraction. Specifically, this approach clusters neighbors and embeds each group's distance information and angle as supplementary information to capture local structures. Evaluation on matching and visual localization datasets demonstrates that our approach achieves state-of-the-art accuracy with low computational overhead among visual description-free methods. Yejun Zhang 0001, Shuzhe Wang, Juho Kannala |
3DV | 3 |
| 2025 | Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual LocalizationabstractVisual localization aims to determine the camera pose of a query image relative to a database of posed images. In recent years, deep neural networks that directly regress camera poses have gained popularity due to their fast inference capabilities. However, existing methods struggle to either generalize well to new scenes or provide accurate camera pose estimates. To address these issues, we present Reloc3r, a simple yet effective visual localization framework. It consists of an elegantly designed relative pose regression network, and a minimalist motion averaging module for absolute pose estimation. Trained on approximately eight million posed image pairs, Reloc3r achieves surprisingly good performance and generalization ability. We conduct extensive experiments on six public datasets, consistently demonstrating the effectiveness and efficiency of the proposed method. It provides high-quality camera pose estimates in real time and generalizes to novel scenes. Code: https://github.com/ffrivera0/reloc3r. Siyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai, Qingnan Fan, Juho Kannala, Yanchao Yang 0001 |
CVPR | 6 |
| 2025 | A Dataset for Semantic Segmentation in the Presence of UnknownsabstractBefore deployment in the real-world deep neural networks require thorough evaluation of how they handle both knowns, inputs represented in the training data, and unknowns (anomalies). This is especially important for scene understanding tasks with safety critical applications, such as in autonomous driving. Existing datasets allow evaluation of only knowns or unknowns - but not both, which is required to establish "in the wild" suitability of deep neural network models. To bridge this gap, we propose a novel anomaly segmentation dataset, ISSU, that features a diverse set of anomaly inputs from cluttered real-world environments. The dataset is twice larger than existing anomaly segmentation datasets, and provides a training, validation and test set for controlled in-domain evaluation. The test set consists of a static and temporal part, with the latter comprised of videos. The dataset provides annotations for both closed-set (knowns) and anomalies, enabling closed-set and open-set evaluation. The dataset covers diverse conditions, such as domain and cross-sensor shift, illumination variation and allows ablation of anomaly detection methods with respect to these variations. Evaluation results of current state-of-the-art methods confirm the need for improvements especially in domain-generalization, small and large object segmentation. The code and the dataset are available at https://github.com/vojirt/benchmark_issu. Zakaria Laskar, Tomás Vojír, Matej Grcic, Iaroslav Melekhov, Shankar Gangisetty, Juho Kannala, Jiri Matas, Giorgos Tolias, C. V. Jawahar |
CVPR | 6 |
| 2025 | DeSplat: Decomposed Gaussian Splatting for Distractor-Free RenderingabstractGaussian splatting enables fast novel view synthesis in static 3D environments. However, reconstructing real-world environments remains challenging as distractors or occluders break the multi-view consistency assumption required for accurate 3D reconstruction. Most existing methods rely on external semantic information from pre-trained models, introducing additional computational overhead as pre-processing steps or during optimization. In this work, we propose a novel method, DeSplat, that directly separates distractors and static scene elements purely based on volume rendering of Gaussian primitives. We initialize Gaussians within each camera view for reconstructing the view-specific dis-tractors to separately model the static 3D scene and dis-tractors in the alpha compositing stages. DeSplat yields an explicit scene separation of static elements and distrac-tors, achieving comparable results to prior distractor-free approaches without sacrificing rendering speed. We demonstrate DeSplat’s effectiveness on three benchmark data sets for distractor-free novel view synthesis. See the project website at https://aaltoml.github.io/desplat/. Yihao Wang 0014, Marcus Klasson, Matias Turkulainen, Shuzhe Wang, Juho Kannala, Arno Solin |
CVPR | 5 |
| 2025 | Multi-Scale Fusion for Object RepresentationabstractRepresenting images or videos as object-level feature vectors, rather than pixel-level feature maps, facilitates advanced visual tasks.
Object-Centric Learning (OCL) primarily achieves this by reconstructing the input under the guidance of Variational Autoencoder (VAE) intermediate representation to drive so-called slots to aggregate as much object information as possible.
However, existing VAE guidance does not explicitly address that objects can vary in pixel sizes while models typically excel at specific pattern scales.
We propose Multi-Scale Fusion (MSF) to enhance VAE guidance for OCL training.
To ensure objects of all sizes fall within VAE's comfort zone, we adopt the image pyramid, which produces intermediate representations at multiple scales;
To foster scale-invariance/variance in object super-pixels, we devise inter/intra-scale fusion, which augments low-quality object super-pixels of one scale with corresponding high-quality super-pixels from another scale.
On standard OCL benchmarks, our technique improves mainstream methods, including state-of-the-art diffusion-based ones.
The source code is available on https://github.com/Genera1Z/MultiScaleFusion. Rongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni Pajarinen |
ICLR | 3 |
| 2025 | FingerVeinSyn-5M: A Million-Scale Dataset and Benchmark for Finger Vein RecognitionabstractA major challenge in finger vein recognition is the lack of large-scale public datasets. Existing datasets contain few identities and limited samples per finger, restricting the advancement of deep learning-based methods. To address this, we introduce FVeinSyn, a synthetic generator capable of producing diverse finger vein patterns with rich intra-class variations. Using FVeinSyn, we created FingerVeinSyn-5M -- the largest available finger vein dataset -- containing 5 million samples from 50,000 unique fingers, each with 100 variations including shift, rotation, scale, roll, varying exposure levels, skin scattering blur, optical blur, and motion blur. FingerVeinSyn-5M is also the first to offer fully annotated finger vein images, supporting deep learning applications in this field. Models pretrained on FingerVeinSyn-5M and fine-tuned with minimal real data achieve an average 53.91% performance gain across multiple benchmarks. The dataset is publicly available at: https://github.com/EvanWang98/FingerVeinSyn-5M. Yifan Wang 0036, Jie Gui, Baosheng Yu, Qi Li 0005, Zhenan Sun, Juho Kannala, Guoying Zhao 0001 |
ACM Multimedia | 6 |
| 2025 | Slot Attention with Re-Initialization and Self-DistillationabstractUnlike popular solutions based on dense feature maps, Object-Centric Learning (OCL) represents visual scenes as sub-symbolic object-level feature vectors, termed slots, which are highly versatile for tasks involving visual modalities. OCL typically aggregates object superpixels into slots by iteratively applying competitive cross attention, known as Slot Attention, with the slots as the query. However, once initialized, these slots are reused naively, causing redundant slots to compete with informative ones for representing objects. This often results in objects being erroneously segmented into parts. Additionally, mainstream methods derive supervision signals solely from decoding slots into the input's reconstruction, overlooking potential supervision based on internal information. To address these issues, we propose Slot Attention with re-Initialization and self-Distillation (DIAS): i) We reduce redundancy in the aggregated slots and re-initialize extra aggregation to update the remaining slots; ii) We drive the bad attention map at the first aggregation iteration to approximate the good at the last iteration to enable self-distillation. Experiments demonstrate that DIAS achieves state-of-the-art on OCL tasks like object discovery and recognition, while also improving advanced visual prediction and reasoning. Our source code and model checkpoints are available on https://github.com/Genera1Z/DIAS. Rongzhen Zhao, Yi Zhao 0014, Juho Kannala, Joni Pajarinen |
ACM Multimedia | 3 |
| 2025 | Vector-Quantized Vision Foundation Models for Object-Centric LearningabstractObject-Centric Learning (OCL) aggregates image or video feature maps into object-level feature vectors, termed slots. It's self-supervision of reconstructing the input from slots struggles with complex object textures, thus Vision Foundation Model (VFM) representations are used as the aggregation input and reconstruction target. Existing methods leverage VFM representations in diverse ways yet fail to fully exploit their potential. In response, we propose a unified architecture, Vector-Quantized VFMs for OCL (VQ-VFM-OCL, or VVO). The key to our unification is simply shared quantizing VFM representations in OCL aggregation and decoding. Experiments show that across different VFMs, aggregators and decoders, our VVO consistently outperforms baselines in object discovery and recognition, as well as downstream visual prediction and reasoning. We also mathematically analyze why VFM representations facilitate OCL aggregation and why their shared quantization as reconstruction targets strengthens OCL supervision. Our source code and model checkpoints are available on https://github.com/Genera1Z/VQ-VFM-OCL. Rongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni Pajarinen |
ACM Multimedia | 3 |
| 2025 | Grouped Discrete Representation for Object-Centric Learning
Rongzhen Zhao, Vivienne Huiling Wang, Juho Kannala, Joni Pajarinen |
ECML/PKDD (6) | 3 |
| 2025 | DN-Splatter: Depth and Normal Priors for Gaussian Splatting and MeshingabstractHigh-fidelity 3D reconstruction of common indoor scenes is crucial for VR and AR applications. 3D Gaussian splat-ting, a novel differentiable rendering technique, has achieved state-of-the-art novel view synthesis results with high ren-dering speeds and relatively low training times. However, its performance on scenes commonly seen in indoor datasets is poor due to the lack of geometric constraints during op-timization. In this work, we explore the use of readily accessible geometric cues to enhance Gaussian splatting op-timization in challenging, ill-posed, and textureless scenes. We extend 3D Gaussian splatting with depth and normal cues to tackle challenging indoor datasets and showcase techniques for efficient mesh extraction. Specifically, we regularize the optimization procedure with depth information, enforce local smoothness of nearby Gaussians, and use off-the-shelf monocular networks to achieve better align-ment with the true scene geometry. We propose an adaptive depth loss based on the gradient of color images, improving depth estimation and novel view synthesis results over various baselines. Our simple yet effective regularization technique enables direct mesh extraction from the Gaus-sian representation, yielding more physically accurate re-constructions of indoor scenes. Our code will be released in https://github.com/maturk/dn-splatter. Matias Turkulainen, Xuqian Ren, Iaroslav Melekhov, Otto Seiskari, Esa Rahtu, Juho Kannala |
WACV | 6 |
| 2024 | Projected Stochastic Gradient Descent with Quantum Annealed Binary Gradients
Maximilian Krahn, Michele Sasdelli, Frances Fengyi Yang, Vladislav Golyanik, Juho Kannala, Tat-Jun Chin, Tolga Birdal |
BMVC | 5 |
| 2024 | Gaussian Splatting in Mirrors: Reflection-aware Rendering via Virtual Camera Optimization
Zihan Wang 0011, Shuzhe Wang, Matias Turkulainen, Junyuan Fang, Juho Kannala |
BMVC | 5 |
| 2024 | DGC-GNN: Leveraging Geometry and Color Cues for Visual Descriptor-Free 2D-3D MatchingabstractMatching 2D keypoints in an image to a sparse 3D point cloud of the scene without requiring visual descriptors has garnered increased interest due to its low memory requirements, inherent privacy preservation, and reduced need for expensive 3D model maintenance compared to visual descriptor-based methods. However, existing algorithms of-ten compromise on performance, resulting in a significant de-terioration compared to their descriptor-based counterparts. In this paper, we introduce DGC-GNN, a novel algorithm that employs a global-to-local Graph Neural Network (GNN) that progressively exploits geometric and color cues to rep-resent keypoints, thereby improving matching accuracy. Our procedure encodes both Euclidean and angular relations at a coarse level, forming the geometric embedding to guide the point matching. We evaluate DGC-GNN on both indoor and outdoor datasets, demonstrating that it not only doubles the accuracy of the state-of-the-art visual descriptor-free algorithm but also substantially narrows the performance gap between descriptor-based and descriptor-free methods.11The code and trained models are available at: https://github.com/AaltoVision/DGC-GNN-release. Shuzhe Wang, Juho Kannala, Daniel Barath |
CVPR | 2 |
| 2024 | Efficient NeRF Optimization - Not All Samples Remain Equally Hard
Juuso Korhonen, Rangu Goutham, Hamed Rezazadegan Tavakoli, Juho Kannala |
ECCV (36) | 4 |
| 2024 | Differentiable Product Quantization for Memory Efficient Camera Relocalization
Zakaria Laskar, Iaroslav Melekhov, Assia Benbihi, Shuzhe Wang, Juho Kannala |
ECCV (85) | 5 |
| 2024 | Gaussian Splatting on the Move: Blur and Rolling Shutter Compensation for Natural Camera Motion
Otto Seiskari, Jerry Ylilammi, Valtteri Kaatrasalo, Pekka Rantalankila, Matias Turkulainen, Juho Kannala, Esa Rahtu, Arno Solin |
ECCV (71) | 6 |
| 2024 | Optimistic Multi-Agent Policy GradientabstractRelative overgeneralization (RO) occurs in cooperative multi-agent learning tasks when agents converge towards a suboptimal joint policy due to overfitting to suboptimal behaviors of other agents. No methods have been proposed for addressing RO in multi-agent policy gradient (MAPG) methods although these methods produce state-of-the-art results. To address this gap, we propose a general, yet simple, framework to enable optimistic updates in MAPG methods that alleviate the RO problem. Our approach involves clipping the advantage to eliminate negative values, thereby facilitating optimistic updates in MAPG. The optimism prevents individual agents from quickly converging to a local optimum. Additionally, we provide a formal analysis to show that the proposed method retains optimality at a fixed point. In extensive evaluations on a diverse set of tasks including the Multi-agent MuJoCo and Overcooked benchmarks, our method outperforms strong baselines on 13 out of 19 tested tasks and matches the performance on the rest. Wenshuai Zhao, Yi Zhao 0014, Juho Kannala, Joni Pajarinen |
ICML | 4 |
| 2024 | Dense Road Surface Grip Map Prediction from Multimodal Image DataabstractAbstract Slippery road weather conditions are prevalent in many regions and cause a regular risk for traffic. Still, there has been less research on how autonomous vehicles could detect slippery driving conditions on the road to drive safely. In this work, we propose a method to predict a dense grip map from the area in front of the car, based on postprocessed multimodal sensor data. We trained a convolutional neural network to predict pixelwise grip values from fused RGB camera, thermal camera, and LiDAR reflectance images, based on weakly supervised ground truth from an optical road weather sensor. The experiments show that it is possible to predict dense grip values with good accuracy from the used data modalities as the produced grip map follows both ground truth measurements and local weather conditions, such as snowy areas on the road. The model using only the RGB camera or LiDAR reflectance modality provided good baseline results for grip prediction accuracy while using models fusing the RGB camera, thermal camera, and LiDAR modalities improved the grip predictions significantly. Jyri Maanpää, Julius Pesonen, Heikki Hyyti, Iaroslav Melekhov, Juho Kannala, Petri Manninen, Antero Kukko, Juha Hyyppä |
ICPR (17) | 5 |
| 2024 | SuperFusion: Multilevel LiDAR-Camera Fusion for Long-Range HD Map GenerationabstractHigh-definition (HD) semantic map generation of the environment is an essential component of autonomous driving. Existing methods have achieved good performance in this task by fusing different sensor modalities, such as LiDAR and camera. However, current works are based on raw data or network feature-level fusion and only consider short-range HD map generation, limiting their deployment to realistic autonomous driving applications. In this paper, we focus on the task of building the HD maps in both short ranges, i.e., within 30m, and also predicting long-range HD maps up to 90m, which is required by downstream path planning and control tasks to improve the smoothness and safety of autonomous driving. To this end, we propose a novel network named SuperFusion, exploiting the fusion of LiDAR and camera data at multiple levels. We use LiDAR depth to improve image depth estimation and use image features to guide long-range LiDAR feature prediction. We benchmark our SuperFusion on the nuScenes dataset and a self-recorded dataset and show that it outperforms the state-of-the-art baseline methods with large margins on all intervals. Additionally, we apply the generated HD map to a downstream path planning task, demonstrating that the long-range HD maps predicted by our method can lead to better path planning for autonomous vehicles. Our code and self-recorded dataset have been released at https://github.com/haomo-ai/SuperFusion. Hao Dong 0011, Weihao Gu, Xianjing Zhang, Jintao Xu 0001, Rui Ai 0001, Huimin Lu 0002, Juho Kannala, Xieyuanli Chen |
ICRA | 7 |
| 2024 | MuSHRoom: Multi-Sensor Hybrid Room Dataset for Joint 3D Reconstruction and Novel View SynthesisabstractMetaverse technologies demand accurate, real-time, and immersive modeling on consumer-grade hardware for both non-human perception (e.g., drone/robot/autonomous car navigation) and immersive technologies like AR/VR, requiring both structural accuracy and photorealism. However, there exists a knowledge gap in how to apply geometric reconstruction and photorealism modeling (novel view synthesis) in a unified framework. To address this gap and promote the development of robust and immersive modeling and rendering with consumer-grade devices, first, we propose a real-world Multi-Sensor Hybrid Room Dataset (MuSHRoom). Our dataset presents exciting challenges and requires state-of-the-art methods to be cost-effective, robust to noisy data and devices, and can jointly learn 3D reconstruction and novel view synthesis instead of treating them as separate tasks, making them ideal for realworld applications. Second, we benchmark several famous pipelines on our dataset for joint 3D mesh reconstruction and novel view synthesis. Finally, in order to further improve the overall performance, we propose a new method that achieves a good trade-off between the two tasks. Our dataset and benchmark show great potential in promoting the improvements for fusing 3D reconstruction and highquality rendering in a robust and computationally efficient end-to-end fashion. The dataset and code are available at the project website: https://xuqianren.github.io/publications/MuSHRoom/. Xuqian Ren, Wenjia Wang 0009, Dingding Cai, Tuuli Tuominen, Juho Kannala, Esa Rahtu |
WACV | 5 |
| 2024 | HSCNet++: Hierarchical Scene Coordinate Classification and Regression for Visual Localization with TransformerabstractAbstract Visual localization is critical to many applications in computer vision and robotics. To address single-image RGB localization, state-of-the-art feature-based methods match local descriptors between a query image and a pre-built 3D model. Recently, deep neural networks have been exploited to regress the mapping between raw pixels and 3D coordinates in the scene, and thus the matching is implicitly performed by the forward pass through the network. However, in a large and ambiguous environment, learning such a regression task directly can be difficult for a single network. In this work, we present a new hierarchical scene coordinate network to predict pixel scene coordinates in a coarse-to-fine manner from a single RGB image. The proposed method, which is an extension of HSCNet, allows us to train compact models which scale robustly to large environments. It sets a new state-of-the-art for single-image localization on the 7-Scenes, 12-Scenes, Cambridge Landmarks datasets, and the combined indoor scenes. Shuzhe Wang, Zakaria Laskar, Iaroslav Melekhov, Yi Zhao 0014, Giorgos Tolias, Juho Kannala |
Int. J. Comput. Vis. | 7 |
| 2023 | Guiding Local Feature Matching with Surface CurvatureabstractWe propose a new method, called curvature similarity extractor (CSE), for improving local feature matching across images. CSE calculates the curvature of the local 3D surface patch for each detected feature point in a viewpoint-invariant manner via fitting quadrics to predicted monocular depth maps. This curvature is then leveraged as an additional signal in feature matching with off-the-shelf matchers like SuperGlue and LoFTR. Additionally, CSE enables end-to-end joint training by connecting the matcher and depth predictor networks. Our experiments demonstrate on large-scale real-world datasets that CSE consistently improves the accuracy of state-of-the-art methods. Fine-tuning the depth prediction network further enhances the accuracy. The proposed approach achieves state-of-the-art results on the ScanNet dataset, showcasing the effectiveness of incorporating 3D geometric information into feature matching.1 Shuzhe Wang, Juho Kannala, Marc Pollefeys, Daniel Barath |
ICCV | 2 |
| 2023 | Simplified Temporal Consistency Reinforcement LearningabstractReinforcement learning (RL) is able to solve complex sequential decision-making tasks but is currently limited by sample efficiency and required computation. To improve sample efficiency, recent work focuses on model-based RL which interleaves model learning with planning. Recent methods further utilize policy learning, value estimation, and, self-supervised learning as auxiliary objectives. In this paper we show that, surprisingly, a simple representation learning approach relying only on a latent dynamics model trained by latent temporal consistency is sufficient for high-performance RL. This applies when using pure planning with a dynamics model conditioned on the representation, but, also when utilizing the representation as policy and value function features in model-free RL. In experiments, our approach learns an accurate dynamics model to solve challenging high-dimensional locomotion tasks with online planners while being 4.1$\times$ faster to train compared to ensemble-based methods. With model-free RL without planning, especially on high-dimensional tasks, such as the Deepmind Control Suite Humanoid and Dog tasks, our approach outperforms model-free methods by a large margin and matches model-based methods’ sample efficiency while training 2.4$\times$ faster. Yi Zhao 0014, Wenshuai Zhao, Rinu Boney, Juho Kannala, Joni Pajarinen |
ICML | 4 |
| 2023 | MixupE: Understanding and improving Mixup from directional derivative perspectiveabstractMixup is a popular data augmentation technique for training deep neural networks where additional samples are generated by linearly interpolating pairs of inputs and their labels. This technique is known to improve the generalization performance in many learning paradigms and applications. In this work, we first analyze Mixup and show that it implicitly regularizes infinitely many directional derivatives of all orders. Based on this new insight, we propose an improved version of Mixup, theoretically justified to deliver better generalization performance than the vanilla Mixup. To demonstrate the effectiveness of the proposed method, we conduct experiments across various domains such as images, tabular data, speech, and graphs. Our results show that the proposed method improves Mixup across multiple datasets using a variety of architectures, for instance, exhibiting an improvement over Mixup by 0.8% in ImageNet top-1 accuracy. Yingtian Zou, Vikas Verma, Sarthak Mittal, Wai Hoh Tang, Hieu Pham 0001, Juho Kannala, Yoshua Bengio, Arno Solin, Kenji Kawaguchi |
UAI | 6 |
| 2023 | Expansion of Visual Hints for Improved Generalization in Stereo MatchingabstractWe introduce visual hints expansion for guiding stereo matching to improve generalization. Our work is motivated by the robustness of Visual Inertial Odometry (VIO) in computer vision and robotics, where a sparse and unevenly distributed set of feature points characterizes a scene. To improve stereo matching, we propose to elevate 2D hints to 3D points. These sparse and unevenly distributed 3D visual hints are expanded using a 3D random geometric graph, which enhances the learning and inference process. We evaluate our proposal on multiple widely adopted benchmarks and show improved performance without access to additional sensors other than the image sequence. To highlight practical applicability and symbiosis with visual odometry, we demonstrate how our methods run on embedded hardware. Andrea Pilzer, Yuxin Hou, Niki Andreas Lopi, Arno Solin, Juho Kannala |
WACV | 5 |
| 2022 | Visual Localization via Few-Shot Scene Region ClassificationabstractVisual (re)localization addresses the problem of estimating the 6-DoF (Degree of Freedom) camera pose of a query image captured in a known scene, which is a key building block of many computer vision and robotics applications. Recent advances in structure-based localization solve this problem by memorizing the mapping from image pixels to scene coordinates with neural networks to build 2D-3D correspondences for camera pose optimization. However, such memorization requires training by amounts of posed images in each scene, which is heavy and inefficient. On the contrary, few-shot images are usually sufficient to cover the main regions of a scene for a human operator to perform visual localization. In this paper, we propose a scene region classification approach to achieve fast and effective scene memorization with few-shot images. Our insight is leveraging a) pre-learned feature extractor, b) scene region classifier, and c) meta-learning strategy to accelerate training while mitigating overfitting. We evaluate our method on both indoor and outdoor benchmarks. The experiments validate the effectiveness of our method in the few-shot setting, and the training time is significantly reduced to only a few minutes.11Code available at: https://github.com/siyandong/SRC Siyan Dong, Shuzhe Wang, Yixin Zhuang, Juho Kannala, Marc Pollefeys, Baoquan Chen |
3DV | 4 |
| 2022 | Uncertainty-Guided Source-Free Domain Adaptation
Subhankar Roy, Martin Trapp 0001, Andrea Pilzer, Juho Kannala, Nicu Sebe, Elisa Ricci 0001, Arno Solin |
ECCV (25) | 4 |
| 2022 | Adaptive Behavior Cloning Regularization for Stable Offline-to-Online Reinforcement LearningabstractOffline reinforcement learning, by learning from a fixed dataset, makes it possible to learn agent behaviors without interacting with the environment.However, depending on the quality of the offline dataset, such pre-trained agents may have limited performance and would further need to be fine-tuned online by interacting with the environment.During online fine-tuning, the performance of the pre-trained agent may collapse quickly due to the sudden distribution shift from offline to online data.We propose to adaptively weigh the behavior cloning loss during online fine-tuning based on the agent's performance and training stability.Moreover, we use a randomized ensemble of Q functions to further increase the sample efficiency of online fine-tuning by performing a large number of learning updates.Experiments show that the proposed method yields state-of-the-art offline-to-online reinforcement learning performance on the popular D4RL benchmark. Yi Zhao 0014, Rinu Boney, Alexander Ilin, Juho Kannala, Joni Pajarinen |
ESANN | 4 |
| 2022 | Multiple Offsets Multilateration: A New Paradigm for Sensor Network Calibration with Unsynchronized Reference NodesabstractPositioning using wave signal measurements is used in several applications, such as GPS systems, structure from sound and Wifi based positioning. Mathematically, such problems require the computation of the positions of receivers and/or transmitters as well as time offsets if the devices are unsynchronized. In this paper, we expand the previous state-of-the-art on positioning formulations by introducing Multiple Offsets Multilateration (MOM), a new mathematical framework to compute the receivers positions with pseudoranges from unsynchronized reference transmitters at known positions. This could be applied in several scenarios, for example structure from sound and positioning with LEO satellites. We mathematically describe MOM, determining how many receivers and transmitters are needed for the network to be solvable, a study on the number of possible distinct solutions is presented and stable solvers based on homotopy continuation are derived. The solvers are shown to be efficient and robust to noise both for synthetic and real audio data. Luca Ferranti, Kalle Åström, Magnus Oskarsson, Jani Boutellier, Juho Kannala |
ICASSP | 5 |
| 2022 | HybVIO: Pushing the Limits of Real-time Visual-inertial OdometryabstractWe present HybVIO, a novel hybrid approach for combining filtering-based visual-inertial odometry (VIO) with optimization-based SLAM. The core of our method is highly robust, independent VIO with improved IMU bias modeling, outlier rejection, stationarity detection, and feature track selection, which is adjustable to run on embedded hardware. Long-term consistency is achieved with a loosely-coupled SLAM module. In academic benchmarks, our solution yields excellent performance in all categories, especially in the real-time use case, where we outperform the current state-of-the-art. We also demonstrate the feasibility of VIO for vehicular tracking on consumer-grade hardware using a custom dataset, and show good performance in comparison to current commercial VISLAM alternatives. Otto Seiskari, Pekka Rantalankila, Juho Kannala, Jerry Ylilammi, Esa Rahtu, Arno Solin |
WACV | 3 |
| 2022 | Single Source One Shot Reenactment using Weighted Motion from Paired Feature PointsabstractImage reenactment is a task where the target object in the source image imitates the motion represented in the driving image. One of the most common reenactment tasks is face image animation. The major challenge in the current face reenactment approaches is to distinguish between facial motion and identity. For this reason, the previous models struggle to produce high-quality animations if the driving and source identities are different (cross-person reenactment). We propose a new (face) reenactment model that learns shape-independent motion features in a self-supervised setup. The motion is represented using a set of paired feature points extracted from the source and driving images simultaneously. The model is generalised to multiple reenactment tasks including faces and non-face objects using only a single source image. The extensive experiments show that the model faithfully transfers the driving motion to the source while retaining the source identity intact. Soumya Tripathy, Juho Kannala, Esa Rahtu |
WACV | 2 |
| 2022 | Interpolated Adversarial Training: Achieving robust neural networks without sacrificing too much accuracyabstractAdversarial robustness has become a central goal in deep learning, both in the theory and the practice. However, successful methods to improve the adversarial robustness (such as adversarial training) greatly hurt generalization performance on the unperturbed data. This could have a major impact on how the adversarial robustness affects real world systems (i.e. many may opt to forego robustness if it can improve accuracy on the unperturbed data). We propose Interpolated Adversarial Training, which employs recently proposed interpolation based training methods in the framework of adversarial training. On CIFAR-10, adversarial training increases the standard test error ( when there is no adversary) from 4.43% to 12.32%, whereas with our Interpolated adversarial training we retain the adversarial robustness while achieving a standard test error of only 6.45%. With our technique, the relative increase in the standard error for the robust model is reduced from 178.1% to just 45.5%. Moreover, we provide mathematical analysis of Interpolated Adversarial Training to confirm its efficiencies and demonstrate its advantages in terms of robustness and generalization. Alex Lamb, Vikas Verma, Kenji Kawaguchi, Alexander Matyasko, Savya Khosla, Juho Kannala, Yoshua Bengio |
Neural Networks | 6 |
| 2022 | Interpolation consistency training for semi-supervised learningabstractWe introduce Interpolation Consistency Training (ICT), a simple and computation efficient algorithm for training Deep Neural Networks in the semi-supervised learning paradigm. ICT encourages the prediction at an interpolation of unlabeled points to be consistent with the interpolation of the predictions at those points. In classification problems, ICT moves the decision boundary to low-density regions of the data distribution. Our experiments show that ICT achieves state-of-the-art performance when applied to standard neural network architectures on the CIFAR-10 and SVHN benchmark datasets. Our theoretical analysis shows that ICT corresponds to a certain type of data-adaptive regularization with unlabeled points which reduces overfitting to labeled points under high confidence values. Vikas Verma, Kenji Kawaguchi, Alex Lamb, Juho Kannala, Arno Solin, Yoshua Bengio, David Lopez-Paz |
Neural Networks | 4 |
| 2022 | Learning to Play Imperfect-Information Games by Imitating an Oracle PlannerabstractWe consider learning to play multiplayer imperfect-information games with simultaneous moves and large state-action spaces. Previous attempts to tackle such challenging games have largely focused on model-free learning methods, often requiring hundreds of years of experience to produce competitive agents. Our approach is based on model-based planning. We tackle the problem of partial observability by first building an (oracle) planner that has access to the full state of the environment and then distilling the knowledge of the oracle to a (follower) agent which is trained to play the imperfect-information game by imitating the oracle’s choices. We experimentally show that planning with naive Monte Carlo tree search does not perform very well in large combinatorial action spaces. We, therefore, propose planning with a fixed-depth tree search and decoupled TS for action selection. We show that the planner is able to discover efficient playing strategies in the games ofClash RoyaleandPommermanand the follower policy successfully learns to implement them by training on a few hundred battles. Rinu Boney, Alexander Ilin, Juho Kannala, Jarno Seppänen |
IEEE Trans. Games | 3 |
| 2021 | Digging Into Self-Supervised Learning of Feature DescriptorsabstractFully-supervised CNN-based approaches for learning local image descriptors have shown remarkable results in a wide range of geometric tasks. However, most of them require per-pixel ground-truth keypoint correspondence data which is difficult to acquire at scale. To address this challenge, recent weakly-and self-supervised methods can learn feature descriptors from relative camera poses or using only synthetic rigid transformations such as homographies. In this work, we focus on understanding the limitations of existing self-supervised approaches and propose a set of improvements that combined lead to powerful feature descriptors. We show that increasing the search space from in-pair to in-batch for hard negative mining brings consistent improvement. To enhance the discriminativeness of feature descriptors, we propose a coarse-to-fine method for mining local hard negatives from a wider search space by using global visual image descriptors. We demonstrate that a combination of synthetic homography transformation, color augmentation, and photorealistic image stylization produces useful representations that are viewpoint and illumination invariant. The feature descriptors learned by the proposed approach perform competitively and surpass their fully- and weakly-supervised counterparts on various geometric benchmarks such as image-based localization, sparse feature matching, and image retrieval. Iaroslav Melekhov, Zakaria Laskar, Shuzhe Wang, Juho Kannala |
3DV | 5 |
| 2021 | GraphMix: Improved Training of GNNs for Semi-Supervised LearningabstractWe present GraphMix, a regularization method for Graph Neural Network based semi-supervised object classification, whereby we propose to train a fully-connected network jointly with the graph neural network via parameter sharing and interpolation-based regularization. Further, we provide a theoretical analysis of how GraphMix improves the generalization bounds of the underlying graph neural network, without making any assumptions about the "aggregation" layer or the depth of the graph neural networks. We experimentally validate this analysis by applying GraphMix to various architectures such as Graph Convolutional Networks, Graph Attention Networks and Graph-U-Net. Despite its simplicity, we demonstrate that GraphMix can consistently improve or closely match state-of-the-art performance using even simpler architectures such as Graph Convolutional Networks, across three established graph benchmarks: Cora, Citeseer and Pubmed citation network datasets, as well as three newly proposed datasets: Cora-Full, Co-author-CS and Co-author-Physics. Vikas Verma, Meng Qu, Kenji Kawaguchi, Alex Lamb, Yoshua Bengio, Juho Kannala, Jian Tang 0005 |
AAAI | 6 |
| 2021 | Interpolation-Based Semi-Supervised Learning for Object DetectionabstractDespite the data labeling cost for the object detection tasks being substantially more than that of the classification tasks, semi-supervised learning methods for object detection have not been studied much. In this paper, we propose an Interpolation-based Semi-supervised learning method for object Detection (ISD), which considers and solves the problems caused by applying conventional Interpolation Regularization (IR) directly to object detection. We divide the output of the model into two types according to the objectness scores of both original patches that are mixed in IR. Then, we apply a separate loss suitable for each type in an unsupervised manner. The proposed losses dramatically improve the performance of semi-supervised learning as well as supervised learning. In the supervised learning setting, our method improves the baseline methods by a significant margin. In the semi-supervised learning setting, our algorithm improves the performance on a benchmark dataset (PASCAL VOC and MSCOCO) in a benchmark architecture (SSD). Our code is available at https://github.com/soo89/ISD-SSD Jisoo Jeong, Vikas Verma, Minsung Hyun, Juho Kannala, Nojun Kwak |
CVPR | 4 |
| 2021 | Sensor Networks TDOA Self-Calibration: 2D Complexity Analysis and SolutionsabstractGiven a network of receivers and transmitters, the process of determining their positions from measured pseudoranges is known as network self-calibration. In this paper we consider 2D networks with synchronized receivers but unsynchronized transmitters and the corresponding calibration techniques, known as Time-Difference-Of-Arrival (TDOA) techniques. Despite previous work, TDOA self-calibration is computationally challenging. Iterative algorithms are very sensitive to the initialization, causing convergence issues. In this paper, we present a novel approach, which gives an algebraic solution to two previously unsolved scenarios. We also demonstrate that our solvers produce an excellent initial value for non-linear optimisation algorithms, leading to a full pipeline robust to noise. Luca Ferranti, Kalle Åström, Magnus Oskarsson, Jani Boutellier, Juho Kannala |
ICASSP | 5 |
| 2021 | Continual Learning for Image-Based Camera LocalizationabstractFor several emerging technologies such as augmented reality, autonomous driving and robotics, visual localization is a critical component. Directly regressing camera pose/3D scene coordinates from the input image using deep neural networks has shown great potential. However, such methods assume a stationary data distribution with all scenes simultaneously available during training. In this paper, we approach the problem of visual localization in a continual learning setup – whereby the model is trained on scenes in an incremental manner. Our results show that similar to the classification domain, non-stationary data induces catastrophic forgetting in deep networks for visual localization. To address this issue, a strong baseline based on storing and replaying images from a fixed buffer is proposed. Furthermore, we propose a new sampling method based on coverage score (Buff-CS) that adapts the existing sampling strategies in the buffering process to the problem of visual localization. Results demonstrate consistent improvements over standard buffering methods on two challenging datasets – 7Scenes, 12Scenes, and also 19Scenes by combining the former scenes1. Shuzhe Wang, Zakaria Laskar, Iaroslav Melekhov, Juho Kannala |
ICCV | 5 |
| 2021 | Novel View Synthesis via Depth-guided Skip ConnectionsabstractWe introduce a principled approach for synthesizing new views of a scene given a single source image. Previous methods for novel view synthesis can be divided into image-based rendering methods (e.g., flow prediction) or pixel generation methods. Flow predictions enable the target view to re-use pixels directly, but can easily lead to distorted results. Directly regressing pixels can produce structurally consistent results but generally suffer from the lack of low-level details. In this paper, we utilize an encoder-decoder architecture to regress pixels of a target view. In order to maintain details, we couple the decoder aligned feature maps with skip connections, where the alignment is guided by predicted depth map of the target view. Our experimental results show that our method does not suffer from distortions and successfully preserves texture details with aligned skip connections. Yuxin Hou, Arno Solin, Juho Kannala |
WACV | 3 |
| 2021 | FACEGAN: Facial Attribute Controllable rEenactment GANabstractThe face reenactment is a popular facial animation method where the person's identity is taken from the source image and the facial motion from the driving image. Recent works have demonstrated high quality results by combining the facial landmark based motion representations with the generative adversarial networks. These models perform best if the source and driving images depict the same person or if the facial structures are otherwise very similar. However, if the identity differs, the driving facial structures leak to the output distorting the reenactment result. We propose a novel Facial Attribute Controllable rEenactment GAN (FACEGAN), which transfers the facial motion from the driving face via the Action Unit (AU) representation. Unlike facial landmarks, the AUs are independent of the facial structure preventing the identity leak. Moreover, AUs provide a human interpretable way to control the reenactment. FACEGAN processes background and face regions separately for optimized output quality. The extensive quantitative and qualitative comparisons show a clear improvement over the state-of-the-art in a single source reenactment task. The results are best illustrated in the reenactment video provided in the supplementary material. The source code will be made available upon publication of the paper. Soumya Tripathy, Juho Kannala, Esa Rahtu |
WACV | 2 |
| 2020 | Data-Efficient Ranking Distillation for Image Retrieval
Zakaria Laskar, Juho Kannala |
ACCV (1) | 2 |
| 2020 | LSD_2 - Joint Denoising and Deblurring of Short and Long Exposure Images with CNNs
Janne Mustaniemi, Juho Kannala, Jiri Matas, Simo Särkkä, Janne Heikkilä |
BMVC | 2 |
| 2020 | Hierarchical Scene Coordinate Classification and Regression for Visual LocalizationabstractVisual localization is critical to many applications in computer vision and robotics. To address single-image RGB localization, state-of-the-art feature-based methods match local descriptors between a query image and a pre-built 3D model. Recently, deep neural networks have been exploited to regress the mapping between raw pixels and 3D coordinates in the scene, and thus the matching is implicitly performed by the forward pass through the network. However, in a large and ambiguous environment, learning such a regression task directly can be difficult for a single network. In this work, we present a new hierarchical scene coordinate network to predict pixel scene coordinates in a coarse-to-fine manner from a single RGB image. The network consists of a series of output layers, each of them conditioned on the previous ones. The final output layer predicts the 3D coordinates and the others produce progressively finer discrete location labels. The proposed method outperforms the baseline regression-only network and allows us to train compact models which scale robustly to large environments. It sets a new state-of-the-art for single-image RGB localization performance on the 7-Scenes, 12-Scenes, Cambridge Landmarks datasets, and three combined scenes. Moreover, for large-scale outdoor localization on the Aachen Day-Night dataset, we present a hybrid approach which outperforms existing scene coordinate regression methods, and reduces significantly the performance gap w.r.t. explicit feature matching methods. Shuzhe Wang, Yi Zhao 0014, Jakob Verbeek, Juho Kannala |
CVPR | 5 |
| 2020 | Can You Trust Your Pose? Confidence Estimation in Visual LocalizationabstractCamera pose estimation in large-scale environments is still an open question and, despite recent promising results, it may still fail in some situations. The research so far has focused on improving subcomponents of estimation pipelines, to achieve more accurate poses. However, there is no guarantee for the result to be correct, even though the correctness of pose estimation is critically important in several visual localization applications, such as in autonomous navigation. In this paper we bring to attention a novel research question, pose confidence estimation, where we aim at quantifying how reliable the visually estimated pose is. We develop a novel confidence measure to fulfill this task and show that it can be flexibly applied to different datasets, indoor or outdoor, and for various visual localization pipelines. We also show that the proposed techniques can be used to accomplish a secondary goal: improving the accuracy of existing pose estimation pipelines. Finally, the proposed approach is computationally light-weight and adds only a negligible increase to the computational effort of pose estimation. Luca Ferranti, Jani Boutellier, Juho Kannala |
ICPR | 4 |
| 2020 | Movement-induced Priors for Deep StereoabstractWe propose a method for fusing stereo disparity estimation with movement-induced prior information. Instead of independent inference frame-by-frame, we formulate the problem as a non-parametric learning task in terms of a temporal Gaussian process prior with a movement-driven kernel for inter-frame reasoning. We present a hierarchy of three Gaussian process kernels depending on the availability of motion information, where our main focus is on a new gyroscope-driven kernel for handheld devices with low-quality MEMS sensors, thus also relaxing the requirement of having full 6D camera poses available. We show how our method can be combined with two state-of-the-art deep stereo methods. The method either work in a plug-and-play fashion with pre-trained deep stereo networks, or further improved by jointly training the kernels together with encoder-decoder architectures, leading to consistent improvement. Yuxin Hou, Muhammad Kamran Janjua, Juho Kannala, Arno Solin |
ICPR | 3 |
| 2020 | Deep AutomodulatorsabstractWe introduce a new category of generative autoencoders called automodulators. These networks can faithfully reproduce individual real-world input images like regular autoencoders, but also generate a fused sample from an arbitrary combination of several such images, allowing instantaneous "style-mixing" and other new applications. An automodulator decouples the data flow of decoder operations from statistical properties thereof and uses the latent vector to modulate the former by the latter, with a principled approach for mutual disentanglement of decoder layers. Prior work has explored similar decoder architecture with GANs, but their focus has been on random sampling. A corresponding autoencoder could operate on real input images. For the first time, we show how to train such a general-purpose model with sharp outputs in high resolution, using novel training techniques, demonstrated on four image data sets. Besides style-mixing, we show state-of-the-art results in autoencoder comparison, and visual image quality nearly indistinguishable from state-of-the-art GANs. We expect the automodulator variants to become a useful building block for image applications and other data domains. Ari Heljakka, Yuxin Hou, Juho Kannala, Arno Solin |
NeurIPS | 3 |
| 2020 | Towards Photographic Image Manipulation with Balanced Growing of Generative AutoencodersabstractWe present a generative autoencoder that provides fast encoding, faithful reconstructions (e.g. retaining the identity of a face), sharp generated/reconstructed samples in high resolutions, and a well-structured latent space that supports semantic manipulation of the inputs. There are no current autoencoder or GAN models that satisfactorily achieve all of these. We build on the progressively growing autoencoder model PIONEER , for which we completely alter the training dynamics based on a careful analysis of recently introduced normalization schemes. We show significantly improved visual and quantitative results for face identity conservation in CELEBA-HQ. Our model achieves state-of-the-art disentanglement of latent space, both quantitatively and via realistic image attribute manipulations. On the LSUN Bedrooms dataset, we improve the disentanglement performance of the vanilla PIONEER, despite having a simpler model. Overall, our results indicate that the PIONEER networks provide a way towards photorealistic face manipulation. Ari Heljakka, Arno Solin, Juho Kannala |
WACV | 3 |
| 2020 | Devon: Deformable Volume Network for Learning Optical Flow
Jack Valmadre, Juho Kannala, Mehrtash Harandi, Philip Torr 0001 |
WACV | 4 |
| 2020 | ICface: Interpretable and Controllable Face Reenactment Using GANsabstractThis paper presents a generic face animator that is able to control the pose and expressions of a given face image. The animation is driven by human interpretable control signals consisting of head pose angles and the Action Unit (AU) values. The control information can be obtained from multiple sources including external driving videos and manual controls. Due to the interpretable nature of the driving signal, one can easily mix the information between multiple sources (e.g. pose from one image and expression from another) and apply selective postproduction editing. The proposed face animator is implemented as a two stage neural network model that is learned in self-supervised manner using a large video collection. The proposed Interpretable and Controllable face reenactment network (ICface) is compared to the state-of-the-art neural network based face animation techniques in multiple tasks. The results indicate that ICface produces better visual quality, while being more versatile than most of the comparison methods. The introduced model could provide a lightweight and easy to use tool for multitude of advanced image and video editing tasks. The program code will be publicly available upon the acceptance of the paper. Soumya Tripathy, Juho Kannala, Esa Rahtu |
WACV | 2 |
| 2020 | MAMBA: Adaptive and Bi-directional Data Transfer for Reliable Camera-display CommunicationabstractCamera-display communication leverages visible light to transfer data wirelessly by using a screen as a transmitter and a camera as a receiver. Such an approach faces several challenges to be employed in practice, including unreliable decoding due to imperfect synchronization and channel impairments. This article introduces MAMBA, a mobile application framework for adaptive camera-display communication with color barcodes. MAMBA employs efficient computer vision techniques and scales with the number of blocks in the barcode, rather than with the number of pixels in the captured image. As a consequence, it allows to take full advantage from the high-resolution cameras available on modern mobile devices. In addition, MAMBA realizes an adaptive and bi-directional protocol for camera-display communication with fast feedback. Specifically, MAMBA carries out dynamic adaptation of both frame rate and length based on environmental conditions and the processing capabilities of the devices. Experimental results show that MAMBA is effective, thereby allowing reliable realtime communication in a variety of operating conditions. Jacopo Bufalino, Maria L. Montoya Freire, Juho Kannala, Mario Di Francesco |
WoWMoM | 3 |
| 2019 | Iterative Path Reconstruction for Large-Scale Inertial Navigation on Smartphones
Santiago Cortés Reina, Yuxin Hou, Juho Kannala, Arno Solin |
FUSION | 3 |
| 2019 | Multi-View Stereo by Temporal Nonparametric FusionabstractWe propose a novel idea for depth estimation from multi-view image-pose pairs, where the model has capability to leverage information from previous latent-space encodings of the scene. This model uses pairs of images and poses, which are passed through an encoder-decoder model for disparity estimation. The novelty lies in soft-constraining the bottleneck layer by a nonparametric Gaussian process prior. We propose a pose-kernel structure that encourages similar poses to have resembling latent spaces. The flexibility of the Gaussian process (GP) prior provides adapting memory for fusing information from nearby views. We train the encoder-decoder and the GP hyperparameters jointly end-to-end. In addition to a batch method, we derive a lightweight estimation scheme that circumvents standard pitfalls in scaling Gaussian process inference, and demonstrate how our scheme can run in real-time on smart devices. Yuxin Hou, Juho Kannala, Arno Solin |
ICCV | 2 |
| 2019 | Interpolation Consistency Training for Semi-supervised Learning
Vikas Verma, Alex Lamb, Juho Kannala, Yoshua Bengio, David Lopez-Paz |
IJCAI | 3 |
| 2019 | Learning Image Relations with Contrast Association NetworksabstractInferring the relations between two images is an important class of tasks in computer vision. Examples of such tasks include computing optical flow and stereo disparity. We treat the relation inference tasks as a machine learning problem and tackle it with neural networks. A key to the problem is learning a representation of relations. We propose a new neural network module, contrast association unit (CAU), which explicitly models the relations between two sets of input variables. Due to the non-negativity of the weights in CAU, we adopt a multiplicative update algorithm for learning these weights. Experiments show that neural networks with CAUs are more effective in learning five fundamental image transformations than conventional neural networks. Yao Lu 0027, Zhirong Yang, Juho Kannala, Samuel Kaski |
IJCNN | 3 |
| 2019 | Regularizing Trajectory Optimization with Denoising AutoencodersabstractTrajectory optimization using a learned model of the environment is one of the core elements of model-based reinforcement learning. This procedure often suffers from exploiting inaccuracies of the learned model. We propose to regularize trajectory optimization by means of a denoising autoencoder that is trained on the same trajectories as the model of the environment. We show that the proposed regularization leads to improved planning with both gradient-based and gradient-free optimizers. We also demonstrate that using regularized trajectory optimization leads to rapid initial learning in a set of popular motor control tasks, which suggests that the proposed approach can be a useful tool for improving sample efficiency. Rinu Boney, Norman Di Palo, Mathias Berglund, Alexander Ilin, Juho Kannala, Antti Rasmus, Harri Valpola |
NeurIPS | 5 |
| 2019 | Semantic Matching by Weakly Supervised 2D Point Set RegistrationabstractIn this paper we address the problem of establishing correspondences between different instances of the same object. The problem is posed as finding the geometric transformation that aligns a given image pair. We use a convolutional neural network (CNN) to directly regress the parameters of the transformation model. The alignment problem is defined in the setting where an unordered set of semantic key-points per image are available, but, without the correspondence information. To this end we propose a novel loss function based on cyclic consistency that solves this 2D point set registration problem by inferring the optimal geometric transformation model parameters. We train and test our approach on a standard benchmark dataset Proposal-Flow (PF-PASCAL). The proposed approach achieves state-of-the-art results demonstrating the effectiveness of the method. In addition, we show our approach further benefits from additional training samples in PF-PASCAL generated by using category level information. Zakaria Laskar, Hamed Rezazadegan Tavakoli, Juho Kannala |
WACV | 3 |
| 2019 | DGC-Net: Dense Geometric Correspondence NetworkabstractThis paper addresses the challenge of dense pixel correspondence estimation between two images. This problem is closely related to optical flow estimation task where ConvNets (CNNs) have recently achieved significant progress. While optical flow methods produce very accurate results for the small pixel translation and limited appearance variation scenarios, they hardly deal with the strong geometric transformations that we consider in this work. In this paper, we propose a coarse-to-fine CNN-based framework that can leverage the advantages of optical flow approaches and extend them to the case of large transformations providing dense and subpixel accurate estimates. It is trained on synthetic transformations and demonstrates very good performance to unseen, realistic, data. Further, we apply our method to the problem of relative camera pose estimation and demonstrate that the model outperforms existing dense approaches. Iaroslav Melekhov, Aleksei Tiulpin, Torsten Sattler, Marc Pollefeys, Esa Rahtu, Juho Kannala |
WACV | 6 |
| 2019 | Gyroscope-Aided Motion Deblurring with Deep NetworksabstractWe propose a deblurring method that incorporates gyroscope measurements into a convolutional neural network (CNN). With the help of such measurements, it can handle extremely strong and spatially-variant motion blur. At the same time, the image data is used to overcome the limitations of gyro-based blur estimation. To train our network, we also introduce a novel way of generating realistic training data using the gyroscope. The evaluation shows a clear improvement in visual quality over the state-of-the-art while achieving real-time performance. Furthermore, the method is shown to improve the performance of existing feature detectors and descriptors against the motion blur. Janne Mustaniemi, Juho Kannala, Simo Särkkä, Jiri Matas, Janne Heikkilä |
WACV | 2 |
| 2019 | Digging Deeper Into Egocentric Gaze PredictionabstractThis paper digs deeper into factors that influence egocentric gaze. Instead of training deep models for this purpose in a blind manner, we propose to inspect factors that contribute to gaze guidance during daily tasks. Bottom-up saliency and optical flow are assessed versus strong spatial prior baselines. Task-specific cues such as vanishing point, manipulation point, and hand regions are analyzed as representatives of top-down information. We also look into the contribution of these factors by investigating a simple recurrent neural model for ego-centric gaze prediction. First, deep features are extracted for all input video frames. Then, a gated recurrent unit is employed to integrate information over time and to predict the next fixation. We also propose an integrated model that combines the recurrent model with several top-down and bottom-up cues. Extensive experiments over multiple datasets reveal that (1) spatial biases are strong in egocentric videos, (2) bottom-up saliency models perform poorly in predicting gaze and underperform spatial biases, (3) deep features perform better compared to traditional features, (4) as opposed to hand regions, the manipulation point is a strong influential cue for gaze prediction, (5) combining the proposed recurrent model with bottom-up cues, vanishing points and, in particular, manipulation point results in the best gaze prediction accuracy over egocentric videos, (6) the knowledge transfer works best for cases where the tasks or sequences are similar, and (7) task and activity recognition can benefit from gaze prediction. Our findings suggest that (1) there should be more emphasis on hand-object interaction and (2) the egocentric vision community should consider larger datasets including diverse stimuli and more subjects. Hamed Rezazadegan Tavakoli, Esa Rahtu, Juho Kannala, Ali Borji |
WACV | 3 |
| 2018 | Pioneer Networks: Progressively Growing Generative Autoencoder
Ari Heljakka, Arno Solin, Juho Kannala |
ACCV (1) | 3 |
| 2018 | Learning Image-to-Image Translation Using Paired and Unpaired Training Samples
Soumya Tripathy, Juho Kannala, Esa Rahtu |
ACCV (2) | 2 |
| 2018 | Recursive Chaining of Reversible Image-to-Image Translators for Face Aging
Ari Heljakka, Arno Solin, Juho Kannala |
ACIVS | 3 |
| 2018 | ADVIO: An Authentic Dataset for Visual-Inertial Odometry
Santiago Cortés Reina, Arno Solin, Esa Rahtu, Juho Kannala |
ECCV (10) | 4 |
| 2018 | Robust Gyroscope-Aided Camera Self-CalibrationabstractCamera calibration for estimating the intrinsic parameters and lens distortion is a prerequisite for various monocular vision applications including feature tracking and video stabilization. This application paper proposes a model for estimating the parameters on the fly by fusing gyroscope and camera data, both readily available in modern day smartphones. The model is based on joint estimation of visual feature positions, camera parameters, and the camera pose, the movement of which is assumed to follow the movement predicted by the gyroscope. Our model assumes the camera movement to be free, but continuous and differentiable, and individual features are assumed to stay stationary. The estimation is performed online using an extended Kalman filter, and it is shown to outperform existing methods in robustness and insensitivity to initialization. We demonstrate the method using simulated data and empirical data from an iPad. Santiago Cortés Reina, Arno Solin, Juho Kannala |
FUSION | 3 |
| 2018 | Inertial Odometry on Handheld SmartphonesabstractBuilding a complete inertial navigation system using the limited quality data provided by current smartphones has been regarded challenging, if not impossible. This paper shows that by careful crafting and accounting for the weak information in the sensor samples, smartphones are capable of pure inertial navigation. We present a probabilistic approach for orientation and use-case free inertial odometry, which is based on double-integrating rotated accelerations. The strength of the model is in learning additive and multiplicative IMU biases online. We are able to track the phone position, velocity, and pose in realtime and in a computationally lightweight fashion by solving the inference with an extended Kalman filter. The information fusion is completed with zero-velocity updates (if the phone remains stationary), altitude correction from barometric pressure readings (if available), and pseudo-updates constraining the momentary speed. We demonstrate our approach using an iPad and iPhone in several indoor dead-reckoning applications and in a measurement tool setup. Arno Solin, Santiago Cortés Reina, Esa Rahtu, Juho Kannala |
FUSION | 4 |
| 2018 | Bottom-Up Attention Guidance for Recurrent Image RecognitionabstractThis paper presents a recurrent neural network architecture, guided by the bottom-up attention, for the recognition task. The proposed architecture processes an input image as a sequence of selectively chosen patches. The patches are chosen from the salient regions of the input image. Using human driven saliency maps from gaze, the benefit of such a selection process is first shown. Next, the performance of computational models of bottom-up attention are assessed as alternative to human attention. Hamed Rezazadegan Tavakoli, Ali Borji, Rao Muhammad Anwer, Esa Rahtu, Juho Kannala |
ICIP | 5 |
| 2018 | Fast Motion Deblurring for Feature Detection and Matching Using Inertial MeasurementsabstractMany computer vision and image processing applications rely on local features. It is well-known that motion blur decreases the performance of traditional feature detectors and descriptors. We propose an inertial-based deblurring method for improving the robustness of existing feature detectors and descriptors against the motion blur. Unlike most deblurring algorithms, the method can handle spatially-variant blur and rolling shutter distortion. Furthermore, it is capable of running in real-time contrary to state-of-the-art algorithms. The limitations of inertial-based blur estimation are taken into account by validating the blur estimates using image data. The evaluation shows that when the method is used with traditional feature detector and descriptor, it increases the number of detected keypoints, provides higher repeatability and improves the localization accuracy. We also demonstrate that such features will lead to more accurate and complete reconstructions when used in the application of 3D visual reconstruction. Janne Mustaniemi, Juho Kannala, Simo Särkkä, Jiri Matas, Janne Heikkilä |
ICPR | 2 |
| 2018 | Accurate 3-D Reconstruction with RGB-D Cameras using Depth Map Fusion and Pose RefinementabstractDepth map fusion is an essential part in both stereo and RGB-D based 3- D reconstruction pipelines. Whether produced with a passive stereo reconstruction or using an active depth sensor, such as Microsoft Kinect, the depth maps have noise and may have poor initial registration. In this paper, we introduce a method which is capable of handling outliers, and especially, even significant registration errors. The proposed method first fuses a sequence of depth maps into a single non-redundant point cloud so that the redundant points are merged together by giving more weight to more certain measurements. Then, the original depth maps are re-registered to the fused point cloud to refine the original camera extrinsic parameters. The fusion is then performed again with the refined extrinsic parameters. This procedure is repeated until the result is satisfying or no significant changes happen between iterations. The method is robust to outliers and erroneous depth measurements as well as even significant depth map registration errors due to inaccurate initial camera poses. Markus Ylimäki, Janne Heikkilä, Juho Kannala |
ICPR | 3 |
| 2018 | PIVO: Probabilistic Inertial-Visual Odometry for Occlusion-Robust NavigationabstractThis paper presents a novel method for visual-inertial odometry. The method is based on an information fusion framework employing low-cost IMU sensors and the monocular camera in a standard smartphone. We formulate a sequential inference scheme, where the IMU drives the dynamical model and the camera frames are used in coupling trailing sequences of augmented poses. The novelty in the model is in taking into account all the cross-terms in the updates, thus propagating the inter-connected uncertainties throughout the model. Stronger coupling between the inertial and visual data sources leads to robustness against occlusion and feature-poor environments. We demonstrate results on data collected with an iPhone and provide comparisons against the Tango device and using the EuRoC data set. Arno Solin, Santiago Cortés Reina, Esa Rahtu, Juho Kannala |
WACV | 4 |
| 2017 | Relative Camera Pose Estimation Using Convolutional Neural Networks
Iaroslav Melekhov, Juha Ylioinas, Juho Kannala, Esa Rahtu |
ACIVS | 3 |
| 2017 | Inertial-based scale estimation for structure from motion on mobile devicesabstractStructure from motion algorithms have an inherent limitation that the reconstruction can only be determined up to the unknown scale factor. Modern mobile devices are equipped with an inertial measurement unit (IMU), which can be used for estimating the scale of the reconstruction. We propose a method that recovers the metric scale given inertial measurements and camera poses. In the process, we also perform a temporal and spatial alignment of the camera and the IMU. Therefore, our solution can be easily combined with any existing visual reconstruction software. The method can cope with noisy camera pose estimates, typically caused by motion blur or rolling shutter artifacts, via utilizing a Rauch-Tung-Striebel (RTS) smoother. Furthermore, the scale estimation is performed in the frequency domain, which provides more robustness to inaccurate sensor time stamps and noisy IMU samples than the previously used time domain representation. In contrast to previous methods, our approach has no parameters that need to be tuned for achieving a good performance. In the experiments, we show that the algorithm outperforms the state-of-the-art in both accuracy and convergence speed of the scale estimate. The accuracy of the scale is around 1% from the ground truth depending on the recording. We also demonstrate that our method can improve the scale accuracy of the Project Tango's build-in motion tracking. Janne Mustaniemi, Juho Kannala, Simo Särkkä, Jiri Matas, Janne Heikkilä |
IROS | 2 |
| 2016 | Cell proposal network for microscopy image analysisabstractRobust cell detection plays a key role in the development of reliable methods for automated analysis of microscopy images. It is a challenging problem due to low contrast, variable fluorescence, weak boundaries, conjoined and overlapping cells, causing most cell detection methods to fail in difficult situations. One approach for overcoming these challenges is to use cell proposals, which enable the use of more advanced features from ambiguous regions and/or information from adjacent frames to make better decisions. However, most current methods rely on simple proposal generation and scoring methods, which limits the performance they can reach. In this paper, we propose a convolutional neural network based method which generates cell proposals to facilitate cell detection, segmentation and tracking. We compare our method against commonly used proposal generation and scoring methods and show that our method generates significantly better proposals, and achieves higher final recall and average precision. Saad Ullah Akram, Juho Kannala, Lauri Eklund, Janne Heikkilä |
ICIP | 2 |
| 2016 | Robust loop closures for scene reconstruction by combining odometry and visual correspondencesabstractGiven an image sequence and odometry from a moving camera, we propose a batch-based approach for robust reconstruction of scene structure and camera motion. A key part of our method is robust loop closure disambiguation. First, a structure-from-motion pipeline is used to get a set of candidate feature correspondences and the respective triangulated 3D landmarks. Thereafter, the compatibility of each correspondence constraint and the odometry is evaluated in a bundle-adjustment optimization, where only compatible constraints affect. Our approach is evaluated using data from a Google Tango device. The results show that it produces better reconstructions than the device's built-in software or a state-of-the-art pose-graph formulation. Zakaria Laskar, Sami Huttunen, Daniel Herrera C., Esa Rahtu, Juho Kannala |
ICIP | 5 |
| 2016 | Deep learning for magnification independent breast cancer histopathology image classificationabstractMicroscopic analysis of breast tissues is necessary for a definitive diagnosis of breast cancer which is the most common cancer among women. Pathology examination requires time consuming scanning through tissue images under different magnification levels to find clinical assessment clues to produce correct diagnoses. Advances in digital imaging techniques offers assessment of pathology images using computer vision and machine learning methods which could automate some of the tasks in the diagnostic pathology workflow. Such automation could be beneficial to obtain fast and precise quantification, reduce observer variability, and increase objectivity. In this work, we propose to classify breast cancer histopathology images independent of their magnifications using convolutional neural networks (CNNs). We propose two different architectures; single task CNN is used to predict malignancy and multi-task CNN is used to predict both malignancy and image magnification level simultaneously. Evaluations and comparisons with previous results are carried out on BreaKHis dataset. Experimental results show that our magnification independent CNN approach improved the performance of magnification specific model. Our results in this limited set of training data are comparable with previous state-of-the-art results obtained by hand-crafted features. However, unlike previous methods, our approach has potential to directly benefit from additional training data, and such additional data could be captured with same or different magnification levels than previous data. Neslihan Bayramoglu, Juho Kannala, Janne Heikkilä |
ICPR | 2 |
| 2016 | Siamese network features for image matchingabstractFinding matching images across large datasets plays a key role in many computer vision applications such as structure-from-motion (SfM), multi-view 3D reconstruction, image retrieval, and image-based localisation. In this paper, we propose finding matching and non-matching pairs of images by representing them with neural network based feature vectors, whose similarity is measured by Euclidean distance. The feature vectors are obtained with convolutional neural networks which are learnt from labeled examples of matching and non-matching image pairs by using a contrastive loss function in a Siamese network architecture. Previously Siamese architecture has been utilised in facial image verification and in matching local image patches, but not yet in generic image retrieval or whole-image matching. Our experimental results show that the proposed features improve matching performance compared to baseline features obtained with networks which are trained for image classification task. The features generalize well and improve matching of images of new landmarks which are not seen at training time. This is despite the fact that the labeling of matching and non-matching pairs is imperfect in our training data. The results are promising considering image retrieval applications, and there is potential for further improvement by utilising more training image pairs with more accurate ground truth labels. Iaroslav Melekhov, Juho Kannala, Esa Rahtu |
ICPR | 2 |
| 2016 | Forget the checkerboard: Practical self-calibration using a planar sceneabstractWe introduce a camera self-calibration method using a planar scene of unknown texture. Planar surfaces are everywhere but checkerboards are not, thus the method can be more easily applied outside the lab. We demonstrate that the accuracy is equivalent to a checkerboard-based calibration, so there is no need for printing checkerboards any more. Moreover, the use of a planar scene provides improved robustness and stronger constraints than a self-calibration with an arbitrary scene. We utilize a closed-form initialization of the focal length with minimal and practical assumptions. The method recovers the intrinsic and extrinsic parameters of the camera and the metric structure of the planar scene. The method is implemented in a real-time application for non-expert users that provides an easy and practical process to obtain high accuracy calibrations. Daniel Herrera C., Juho Kannala, Janne Heikkilä |
WACV | 2 |
| 2016 | Parallax correction via disparity estimation in a multi-aperture camera
Janne Mustaniemi, Juho Kannala, Janne Heikkilä |
Mach. Vis. Appl. | 2 |
| 2015 | Online Face Recognition System Based on Local Binary Patterns and Facial Landmark Tracking
Marko Linna, Juho Kannala, Esa Rahtu |
ACIVS | 2 |
| 2015 | Human Epithelial Type 2 cell classification with convolutional neural networksabstractAutomated cell classification in Indirect Immunofluorescence (IIF) images has potential to be an important tool in clinical practice and research. This paper presents a framework for classification of Human Epithelial Type 2 cell IIF images using convolutional neural networks (CNNs). Previuos state-of-the-art methods show classification accuracy of 75.6% on a benchmark dataset. We conduct an exploration of different strategies for enhancing, augmenting and processing training data in a CNN framework for image classification. Our proposed strategy for training data and pre-training and fine-tuning the CNN network led to a significant increase in the performance over other approaches that have been used until now. Specifically, our method achieves a 80.25% classification accuracy. Source code and models to reproduce the experiments in the paper is made publicly available. Neslihan Bayramoglu, Juho Kannala, Janne Heikkilä |
BIBE | 2 |
| 2015 | Disparity Estimation for Image Fusion in a Multi-aperture Camera
Janne Mustaniemi, Juho Kannala, Janne Heikkilä |
CAIP (2) | 2 |
| 2015 | Optimizing the Accuracy and Compactness of Multi-view Reconstructions
Markus Ylimäki, Juho Kannala, Janne Heikkilä |
CAIP (2) | 2 |
| 2015 | A novel feature descriptor based on microscopy image statisticsabstractIn this paper, we propose a novel feature description algorithm based on image statistics. The pipeline first performs independent component analysis on training image patches to obtain basis vectors (filters) for a lower dimensional representation. Then for a given image, a set of filter responses at each pixel is computed. Finally, a histogram representation, which considers the signs and magnitudes of the responses as well as the number of filters, is applied on local image patches. We propose to apply this idea to a microscopy image pixel identification system based on a learning framework. Experimental results show that the proposed algorithm performs better than the state-of-the-art descriptors in biomedical images of different microscopy modalities. Neslihan Bayramoglu, Juho Kannala, Malin Akerfelt, Mika Kaakinen, Lauri Eklund, Matthias Nees, Janne Heikkilä |
ICIP | 2 |
| 2015 | Adaptive Kalman filtering and smoothing for gravitation tracking in mobile systemsabstractThis paper is concerned with inertial-sensor-based tracking of the gravitation direction in mobile devices such as smartphones. Although this tracking problem is a classical one, choosing a good state-space for this problem is not entirely trivial. Even though for many other orientation related tasks a quaternion-based representation tends to work well, for gravitation tracking their use is not always advisable. In this paper we present a convenient linear quaternion-free state-space model for gravitation tracking. We also discuss the efficient implementation of the Kalman filter and smoother for the model. Furthermore, we propose an adaption mechanism for the Kalman filter which is able to filter out shot-noises similarly as has been proposed in context of adaptive and robust Kalman filtering. We compare the proposed approach to other approaches using measurement data collected with a smartphone. Simo Särkkä, Ville Tolvanen, Juho Kannala, Esa Rahtu |
IPIN | 3 |
| 2015 | Fast and accurate multi-view reconstruction by multi-stage prioritised matchingabstractIn this study, the authors propose a multi‐view stereo reconstruction method which creates a three‐dimensional point cloud of a scene from multiple calibrated images captured from different viewpoints. The method is based on a prioritised match expansion technique, which starts from a sparse set of seed points, and iteratively expands them into neighbouring areas by using multiple expansion stages. Each seed point represents a surface patch and has a position and a surface normal vector. The location and surface normal of the seeds are optimised using a homography‐based local image alignment. The propagation of seeds is performed in a prioritised order in which the most promising seeds are expanded first and removed from the list of seeds. The first expansion stage proceeds until the list of seeds is empty. In the following expansion stages, the current reconstruction may be further expanded by finding new seeds near the boundaries of the current reconstruction. The prioritised expansion strategy allows efficient generation of accurate point clouds and their experiments show its benefits compared with non‐prioritised expansion. In addition, a comparison to the widely used patch‐based multi‐view stereo software shows that their method is significantly faster and produces more accurate and complete reconstructions. Markus Ylimäki, Juho Kannala, Jukka Holappa, Sami S. Brandt, Janne Heikkilä |
IET Comput. Vis. | 2 |
| 2014 | DT-SLAM: Deferred Triangulation for Robust SLAMabstractObtaining a good baseline between different video frames is one of the key elements in vision-based monocular SLAM systems. However, if the video frames contain only a few 2D feature correspondences with a good baseline, or the camera only rotates without sufficient translation in the beginning, tracking and mapping becomes unstable. We introduce a real-time visual SLAM system that incrementally tracks individual 2D features, and estimates camera pose by using matched 2D features, regardless of the length of the baseline. Triangulating 2D features into 3D points is deferred until key frames with sufficient baseline for the features are available. Our method can also deal with pure rotational motions, and fuse the two types of measurements in a bundle adjustment step. Adaptive criteria for key frame selection are also introduced for efficient optimization and dealing with multiple maps. We demonstrate that our SLAM system improves camera pose estimates and robustness, even with purely rotational motions. Daniel Herrera C., Juho Kannala, Kari Pulli, Janne Heikkilä |
3DV | 3 |
| 2014 | Segmentation of Cells from Spinning Disk Confocal Images Using a Multi-stage Approach
Saad Ullah Akram, Juho Kannala, Mika Kaakinen, Lauri Eklund, Janne Heikkilä |
ACCV (3) | 2 |
| 2014 | Generating Object Segmentation Proposals Using Global and Local SearchabstractWe present a method for generating object segmentation proposals from groups of superpixels. The goal is to propose accurate segmentations for all objects of an image. The proposed object hypotheses can be used as input to object detection systems and thereby improve efficiency by replacing exhaustive search. The segmentations are generated in a class-independent manner and therefore the computational cost of the approach is independent of the number of object classes. Our approach combines both global and local search in the space of sets of superpixels. The local search is implemented by greedily merging adjacent pairs of superpixels to build a bottom-up segmentation hierarchy. The regions from such a hierarchy directly provide a part of our region proposals. The global search provides the other part by performing a set of graph cut segmentations on a superpixel graph obtained from an intermediate level of the hierarchy. The parameters of the graph cut problems are learnt in such a manner that they provide complementary sets of regions. Experiments with Pascal VOC images show that we reach state-of-the-art with greatly reduced computational cost. Pekka Rantalankila, Juho Kannala, Esa Rahtu |
CVPR | 2 |
| 2014 | Understanding Objects in Detail with Fine-Grained AttributesabstractWe study the problem of understanding objects in detail, intended as recognizing a wide array of fine-grained object attributes. To this end, we introduce a dataset of 7, 413 airplanes annotated in detail with parts and their attributes, leveraging images donated by airplane spotters and crowd-sourcing both the design and collection of the detailed annotations. We provide a number of insights that should help researchers interested in designing fine-grained datasets for other basic level categories. We show that the collected data can be used to study the relation between part detection and attribute prediction by diagnosing the performance of classifiers that pool information from different parts of an object. We note that the prediction of certain attributes can benefit substantially from accurate part detection. We also show that, differently from previous results in object detection, employing a large number of part templates can improve detection accuracy at the expenses of detection speed. We finally propose a coarse-to-fine approach to speed up detection through a hierarchical cascade algorithm. Andrea Vedaldi, Siddharth Mahendran, Stavros Tsogkas, Subhransu Maji, Ross B. Girshick, Juho Kannala, Esa Rahtu, Iasonas Kokkinos, Matthew B. Blaschko, David J. Weiss, Ben Taskar, Karen Simonyan, Naomi Saphra, Sammy Mohamed |
CVPR | 6 |
| 2014 | Detection of Tumor Cell Spheroids from Co-cultures Using Phase Contrast Images and Machine Learning ApproachabstractAutomated image analysis is demanded in cell biology and drug development research. The type of microscopy is one of the considerations in the trade-offs between experimental setup, image acquisition speed, molecular labelling, resolution and quality of images. In many cases, phase contrast imaging gets higher weights in this optimization. And it comes at the price of reduced image quality in imaging 3D cell cultures. For such data, the existing state-of-the-art computer vision methods perform poorly in segmenting specific cell type. Low SNR, clutter and occlusions are basic challenges for blind segmentation approaches. In this study we propose an automated method, based on a learning framework, for detecting particular cell type in cluttered 2D phase contrast images of 3D cell cultures that overcomes those challenges. It depends on local features defined over super pixels. The method learns appearance based features, statistical features, textural features and their combinations. Also, the importance of each feature is measured by employing Random Forest classifier. Experiments show that our approach does not depend on training data and the parameters. Neslihan Bayramoglu, Mika Kaakinen, Lauri Eklund, Malin Akerfelt, Matthias Nees, Juho Kannala, Janne Heikkilä |
ICPR | 6 |
| 2014 | An In-depth Examination of Local Binary Descriptors in Unconstrained Face RecognitionabstractAutomatic face recognition in unconstrained conditions is a difficult task which has recently attained increasing attention. In this domain, face verification methods have significantly improved since the release of the Labeled Faces in the Wild database, but the related problem of face identification, is still lacking considerations, which is partly because of the shortage of representative databases. Only recently, two new datasets called Remote Face and Point-and-Shoot Challenge were published providing appropriate benchmarks for the research community to investigate the problem of face recognition in challenging imaging conditions, in both, verification and identification modes. In this paper we provide an in-depth examination of three local binary description methods in unconstrained face recognition evaluating them on these two recently published datasets. In detail, we investigate three well established methods separately and fusing them at rank- and score-levels. We are using a well-defined evaluation protocol allowing a fair comparison of our results for future examinations. Juha Ylioinas, Abdenour Hadid, Juho Kannala, Matti Pietikäinen |
ICPR | 3 |
| 2014 | Learning local image descriptors using binary decision treesabstractIn this paper we propose a unified framework for learning such local image descriptors that describe pixel neighborhoods using binary codes. The descriptors are constructed using binary decision trees which are learnt from a set of training image patches. Our framework generalizes several previously proposed binary descriptors, such as BRIEF, LBP and their variants, and provides a principled way to learn new constructions which have not been previously studied. Further, the proposed framework can utilize both labeled or unlabeled training data, and hence fits to both supervised and unsupervised learning scenarios. We evaluate our framework using varying levels of supervision in the learning phase. The experiments show that our descriptor constructions perform comparably to benchmark descriptors in two different applications, namely texture categorization and age group classification from facial images. Juha Ylioinas, Juho Kannala, Abdenour Hadid, Matti Pietikäinen |
WACV | 2 |
| 2013 | A Learned Joint Depth and Intensity Prior Using Markov Random FieldsabstractWe present a joint prior that takes intensity and depth information into account. The prior is defined using a flexible Field-of-Experts model and is learned from a dataBase of natural images. It is a generative model and has an efficient method for sampling. We use sampling from the model to perform in painting and up sampling of depth maps when intensity information is available. We show that including the intensity information in the prior improves the results obtained from the model. We also compare to another two-channel in painting approach and show superior results. Daniel Herrera C., Juho Kannala, Peter F. Sturm, Janne Heikkilä |
3DV | 2 |
| 2012 | BSIF: Binarized statistical image features
Juho Kannala, Esa Rahtu |
ICPR | 1 |
| 2012 | Robust and accurate multi-view reconstruction by prioritized matching
Markus Ylimäki, Juho Kannala, Jukka Holappa, Janne Heikkilä, Sami S. Brandt |
ICPR | 2 |
| 2012 | Joint Depth and Color Camera Calibration with Distortion CorrectionabstractWe present an algorithm that simultaneously calibrates two color cameras, a depth camera, and the relative pose between them. The method is designed to have three key features: accurate, practical, and applicable to a wide range of sensors. The method requires only a planar surface to be imaged from various poses. The calibration does not use depth discontinuities in the depth image, which makes it flexible and robust to noise. We apply this calibration to a Kinect device and present a new depth distortion model for the depth sensor. We perform experiments that show an improved accuracy with respect to the manufacturer's calibration. Daniel Herrera C., Juho Kannala, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Accurate and Practical Calibration of a Depth and Color Camera Pair
Daniel Herrera C., Juho Kannala, Janne Heikkilä |
CAIP (2) | 2 |
| 2011 | Learning a category independent object detection cascadeabstractCascades are a popular framework to speed up object detection systems. Here we focus on the first layers of a category independent object detection cascade in which we sample a large number of windows from an objectness prior, and then discriminatively learn to filter these candidate windows by an order of magnitude. We make a number of contributions to cascade design that substantially improve over the state of the art: (i) our novel objectness prior gives much higher recall than competing methods, (ii) we propose objectness features that give high performance with very low computational cost, and (iii) we make use of a structured output ranking approach to learn highly effective, but inexpensive linear feature combinations by directly optimizing cascade performance. Thorough evaluation on the PASCAL VOC data set shows consistent improvement over the current state of the art, and over alternative discriminative learning strategies. Esa Rahtu, Juho Kannala, Matthew B. Blaschko |
ICCV | 2 |
| 2010 | Segmenting Salient Objects from Images and Videos
Esa Rahtu, Juho Kannala, Mikko Salo, Janne Heikkilä |
ECCV (5) | 2 |
| 2010 | Quasi-dense Wide Baseline Matching for Three ViewsabstractThis paper proposes a method for computing a quasi-dense set of matching points between three views of a scene. The method takes a sparse set of seed matches between pairs of views as input and then propagates the seeds to neighboring regions. The proposed method is based on the best-first match propagation strategy, which is here extended from two-view matching to the case of three views. The results show that utilizing the three-view constraint during the correspondence growing improves the accuracy of matching and reduces the occurrence of outliers. In particular, compared with two-view stereo, our method is more robust for repeating texture. Since the proposed approach is able to produce high quality depth maps from only three images, it could be used in multi-view stereo systems that fuse depth maps from multiple views. Pekka Koskenkorva, Juho Kannala, Sami S. Brandt |
ICPR | 2 |
| 2008 | Object recognition and segmentation by non-rigid quasi-dense matchingabstractIn this paper, we present a non-rigid quasi-dense matching method and its application to object recognition and segmentation. The matching method is based on the match propagation algorithm which is here extended by using local image gradients for adapting the propagation to smooth non-rigid deformations of the imaged surfaces. The adaptation is based entirely on the local properties of the images and the method can be hence used in non-rigid image registration where global geometric constraints are not available. Our approach for object recognition and segmentation is directly built on the quasi-dense matching. The quasi-dense pixel matches between the model and test images are grouped into geometrically consistent groups using a method which utilizes the local affine transformation estimates obtained during the propagation. The number and quality of geometrically consistent matches is used as a recognition criterion and the location of the matching pixels directly provides the segmentation. The experiments demonstrate that our approach is able to deal with extensive background clutter, partial occlusion, large scale and viewpoint changes, and notable geometric deformations. Juho Kannala, Esa Rahtu, Sami S. Brandt, Janne Heikkilä |
CVPR | 1 |
| 2008 | Measuring and modelling sewer pipes from video
Juho Kannala, Sami S. Brandt, Janne Heikkilä |
Mach. Vis. Appl. | 1 |
| 2007 | Quasi-Dense Wide Baseline Matching Using Match PropagationabstractIn this paper we propose extensions to the match propagation algorithm which is a technique for computing quasi-dense point correspondences between two views. The extensions make the match propagation applicable for wide baseline matching, i.e., for cases where the camera pose can vary a lot between the views. Our first extension is to use a local affine model for the geometric transformation between the images. The estimate of the local transformation is obtained from affine covariant interest regions which are used as seed matches. The second extension is to use the second order intensity moments to adapt the current estimate of the local affine transformation during the propagation. This allows a single seed match to propagate into regions where the local transformation between the views differs from the initial one. The experiments with real data show that the proposed techniques improve both the quality and coverage of the quasi-dense disparity map. Juho Kannala, Sami S. Brandt |
CVPR | 1 |
| 2006 | Algorithms for Computing a Planar Homography from Conics in CorrespondenceabstractThis paper presents two new algorithms for computing a planar homography from conic correspondences. Firstly, we propose a linear algorithm for computing the homography when there are three or more conic correspondences. In this case, we get an overdetermined set of linear equations and the solution that minimizes the algebraic distance is obtained by the singular value decomposition. Secondly, we propose another algorithm for determining the homography from only two conic correspondences. Unlike the previous algorithms our approach uses only linear algebra and does not require solving high-degree polynomial equations. Hence, the proposed formulation leads to an algorithm that is efficient and easy to implement. In addition, our approach incorporates the computation of the two projective invariants for a pair of conics. These invariants provide a condition for the existence of a homography between the pairs of conics. We evaluate the characteristics and robustness of the proposed algorithms in experiments with synthetic and real data. 1 Juho Kannala, Mikko Salo, Janne Heikkilä |
BMVC | 1 |
| 2006 | A Generic Camera Model and Calibration Method for Conventional, Wide-Angle, and Fish-Eye LensesabstractFish-eye lenses are convenient in such applications where a very wide angle of view is needed, but their use for measurement purposes has been limited by the lack of an accurate, generic, and easy-to-use calibration procedure. We hence propose a generic camera model, which is suitable for fish-eye lens cameras as well as for conventional and wide-angle lens cameras, and a calibration method for estimating the parameters of the model. The achieved level of calibration accuracy is comparable to the previously reported state-of-the-art. Juho Kannala, Sami S. Brandt |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Affine registration with multi-scale autoconvolutionabstractIn this paper we propose a novel method for the recovery of affine transformation parameters between two images. Registration is achieved without separate feature extraction by directly utilizing the intensity distribution of the images. The method can also be used for matching point sets under affine transformations. Our approach is based on the same probabilistic interpretation of the image function as the recently introduced multi-scale autoconvolution (MSA) transform. Here we describe how the framework may be used in image registration and present two variants of the method for practical implementation. The proposed method is experimented with binary and grayscale images and compared with other non-feature-based registration methods. The experiments show that the new method can efficiently align images of isolated objects and is relatively robust. Juho Kannala, Esa Rahtu, Janne Heikkilä |
ICIP (3) | 1 |