EDBT 2026 Demo / reviewers in the wild / expert
Janne Heikkilä
dblp:25/4802
· DBLP profile ↗
122ranked-venue papers
13as first author
27since 2021 · last 2026
0000-0003-0073-0866ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 96 · 9 first-author · 18 since 2021Artificial intelligence and machine learning · 67 · 11 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Systems, architecture and hardware · 2Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Making Minimal Solvers Inverse-Free Using Null Space ComputationabstractIn this paper, we propose a novel resultant-based method for solving polynomial systems of equations that are commonly encountered in computer vision, particularly as minimal problems. Unlike traditional algorithms that rely on matrix inversion, the primary merits of the proposed method are the numerical stability of its formulation and the lack of need to compute the inverse of a matrix by leveraging null space computations. Additionally, its formulation paves the way for more computations to be performed in the offline stage by using the sparsity of coefficient matrices, thereby reducing the computational load in the online stage. This inverse-free formulation is especially suited for sparse systems and offers improved robustness in scenarios where matrix inversion is unstable or infeasible. Experimental results demonstrate better accuracy compared to state-of-the-art methods such as SRBM and GAPS across a variety of camera geometry problems. Hassan Bozorgmanesh, Janne Heikkilä |
3DV | 2 |
| 2026 | An End-to-End Depth-Based Pipeline for Selfie Image RectificationabstractPortraits or selfie images taken from a close distance typically suffer from perspective distortion. In this paper, we propose an end-to-end deep learning-based rectification pipeline to mitigate the effects of perspective distortion. We learn to predict the facial depth by training a deep CNN. The estimated depth is utilized to adjust the camera-to-subject distance by moving the camera farther, increasing the camera focal length, and reprojecting the 3D image features to the new perspective. The reprojected features are then fed to an inpainting module to fill in the missing pixels. We leverage a differentiable renderer to enable end-to-end training of our depth estimation and feature extraction nets to improve the rectified outputs. To boost the results of the inpainting module, we incorporate an auxiliary module to predict the horizontal movement of the camera which decreases the area that requires hallucination of challenging face parts such as ears. Unlike previous works, we process the full-frame input image at once without cropping the subject's face and processing it separately from the rest of the body, eliminating the need for complex post-processing steps to attach the face back to the subject's body. To train our network, we utilize the popular game engine Unreal Engine to generate a large synthetic face dataset containing various subjects, head poses, expressions, eyewear, clothes, and lighting. Quantitative and qualitative results show that our rectification pipeline outperforms previous methods, and produces comparable results with a time-consuming 3D GAN-based method while being more than 260 times faster. Ahmed Alhawwary, Janne Mustaniemi, Phong Nguyen 0001, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Prompt is All You Need: Prompting Foundation Models for Large-Scale Self-Supervised Semantic SegmentationabstractThis paper addresses the important and challenging task of large-scale unsupervised semantic segmentation (LUSS). We present the first attempt to unleash the power of foundation models (FMs) for the challenging, dense prediction task LUSS, and our main objective is to present simple, effective yet efficient solutions for LUSS, namely Prompting foundation models for LUSS (PLUSS). Firstly, we proposed a cascade framework PLUSS$_\alpha$α by effectively marrying CLIPS, Grounding DINO, and SAM in a zero-shot manner. This cascade architecture automatically generates semantic and spatial prompts for SAM, establishing a strong baseline that significantly outperforms previous state-of-the-art methods. Building upon this foundation, we propose PLUSS$_\beta$β, which addresses the critical bottleneck of prompt quality through two novel tuner modules: a semantic tuner that enhances fine-grained category discrimination via visual prompt tuning, and a box tuner that improves object localization through cross-modal feature fusion. Both tuners are optimized by capitalizing on the knowledge already present within the foundation models themselves, deriving self-supervised signals from internal model consistency. This approach requires no external supervision or updates to the foundation models' parameters. Extensive experiments on ImageNet-S benchmarks demonstrate that PLUSS$_\beta$β achieves remarkable performance improvements, surpassing the previous best method by 39.6%, 27.3%, and 22.6% in mIoU for 50, 300, and 919 categories respectively. Our approach exhibits robust category-shape representation across varying object sizes and dataset scales, while maintaining strong generalization capabilities for open-vocabulary tasks. The proposed framework provides a solid baseline for adapting foundation models to downstream vision tasks. Jiaojiao Su, Qiwu Luo, Shuzhou Sun, Yuenan Hou, Xinyu Zhang 0010, Janne Heikkilä, Chunhua Yang 0001, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | GS-Pose: Generalizable Segmentation-Based 6D Object Pose Estimation with 3D Gaussian SplattingabstractThis paper introduces GS-Pose, a unified framework for localizing and estimating the 6D pose of novel objects. GS-Pose begins with a set of posed RGB images of a previously unseen object and builds three distinct representations stored in a database. At inference, GS-Pose operates sequentially by locating the object in the input image, estimating its initial 6D pose using a retrieval approach, and refining the pose with a render-and-compare method. The key insight is the application of the appropriate object representation at each stage of the process. In particular, for the refinement step, we leverage 3D Gaussian splatting, a novel differentiable rendering technique that offers high rendering speed and relatively low optimization time. Off-the-shelf toolchains and commodity hard-ware, such as mobile phones, can be used to capture new objects to be added to the database. Extensive evaluations on the LINEMOD and OnePose-LowTexture datasets demonstrate excellent performance, establishing the new state-of-the-art. The source code is publicly available at https://github.com/dingdingcai/GSPose. Dingding Cai, Janne Heikkilä, Esa Rahtu |
3DV | 2 |
| 2025 | Synthesizing Images with Different Exposure Settings for Low-Light Image Enhancement
Ahmed Alhawwary, Janne Mustaniemi, Janne Heikkilä |
CAIP (2) | 3 |
| 2025 | FaDeN: Fast Depth-Supervised NeRFs with RGB-D Cameras
Janne Mustaniemi, Li Liu 0002, Janne Heikkilä |
CAIP (1) | 3 |
| 2025 | A Conic Transformation Approach for Solving the Perspective-Three-Point ProblemabstractWe propose a conic transformation method to solve the Perspective-Three-Point (P3P) problem. In contrast to the current state-of-the-art solvers, which formulate the P3P problem by intersecting two conics and constructing a de-generate conic to find the intersection, our approach builds upon a new formulation based on a transformation that maps the two conics to a new coordinate system, where one of the conics becomes a standard parabola in a canonical form. This enables expressing one variable in terms of the other variable, and as a consequence, substantially simpli-fies the problem of finding the conic intersection. Moreover, the polynomial coefficients are fast to compute, and we only need to determine the real-valued intersection points, which avoids the requirement of using computationally expensive complex arithmetic. While the current state-of-the-art methods reduce the conic intersection problem to solving a univariate cubic equation, our approach, despite resulting in a quartic equation, is still faster thanks to this new simplified formulation. Extensive evaluations demonstrate that our method achieves higher speed while maintaining robustness and stability comparable to state-of-the-art methods. Haidong Wu, Snehal Bhayani, Janne Heikkilä |
WACV | 3 |
| 2025 | A Causal Adjustment Module for Debiasing Scene Graph GenerationabstractWhile recent debiasing methods for Scene Graph Generation (SGG) have shown impressive performance, these efforts often attribute model bias solely to the long-tail distribution of relationships, overlooking the more profound causes stemming from skewed object and object pair distributions. In this paper, we employ causal inference techniques to model the causality among these observed skewed distributions. Our insight lies in the ability of causal inference to capture the unobservable causal effects between complex distributions, which is crucial for tracing the roots of model bias. Specifically, we introduce the Mediator-based Causal Chain Model (MCCM), which, in addition to modeling causality among objects, object pairs, and relationships, incorporates mediator variables, i.e., cooccurrence distribution, for complementing the causality. Following this, we propose the Causal Adjustment Module (CAModule) to estimate the modeled causal structure, using variables from MCCM as inputs to produce a set of adjustment factors aimed at correcting biased model predictions. Moreover, our method enables the composition of zero-shot relationships, thereby enhancing the model's ability to recognize such relationships. Experiments conducted across various SGG backbones and popular benchmarks demonstrate that CAModule achieves state-of-the-art mean recall rates, with significant improvements also observed on the challenging zero-shot recall rate metric. Li Liu 0002, Shuzhou Sun, Shuaifeng Zhi, Fan Shi 0003, Zhen Liu 0004, Janne Heikkilä, Yongxiang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph GenerationabstractExisting two-stage Scene Graph Generation (SGG) frameworks typically incorporate a detector to extract relationship features and a classifier to categorize these relationships; therefore, the training paradigm follows a causal chain structure, where the detector's inputs determine the classifier's inputs, which in turn influence the final predictions. However, such a causal chain structure can yield spurious correlations between the detector's inputs and the final predictions, i.e., the prediction of a certain relationship may be influenced by other relationships. This influence can induce at least two observable biases: tail relationships are predicted as head ones, and foreground relationships are predicted as background ones; notably, the latter bias is seldom discussed in the literature. To address this issue, we propose reconstructing the causal chain structure into a reverse causal structure, wherein the classifier's inputs are treated as the confounder, and both the detector's inputs and the final predictions are viewed as causal variables. Specifically, we term the reconstructed causal paradigm as the Reverse causal Framework for SGG (RcSGG). RcSGG initially employs the proposed Active Reverse Estimation (ARE) to intervene on the confounder to estimate the reverse causality, i.e., the causality from final predictions to the classifier's inputs. Then, the Maximum Information Sampling (MIS) is suggested to enhance the reverse causality estimation further by considering the relationship information. Theoretically, RcSGG can mitigate the spurious correlations inherent in the SGG framework, subsequently eliminating the induced biases. Comprehensive experiments on popular benchmarks and diverse SGG frameworks show the state-of-the-art mean recall rate. Shuzhou Sun, Li Liu 0002, Tianpeng Liu, Shuaifeng Zhi, Ming-Ming Cheng, Janne Heikkilä, Yongxiang Liu |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2024 | Polynomial Solvers for mmWave Radio BeamformingabstractMillimeter (mmWave) beamforming is an integral component of fifth-generation (5G) and beyond radio commu-nications. 5G beamforming involves the initial beam selection procedure using a codebook with multiple radio beam directions. Conventional codebook-based alignment schemes involve exhaustive sweeping over the predefined beam directions, the number of which increases significantly with large numbers of antennas resulting in undesirable latency and communications signal overhead. In this paper, we propose a novel algebraic-based codebook using Gröbner basis polynomial solvers to reduce the signal overhead during beam alignment. We also analyze the complexity-performance tradeoff between the proposed algebraic-based codebook and the exhaustive-based beam alignment across different monomial thresholds, multiple antenna configurations and radio contextual location information. Our results show that the proposed approach reduces the beam-search overhead at an average complexity reduction ratio of 73.95% with a performance tradeoff error of 32.25%. Praneeth Susarla, Snehal Bhayani, S. S. Krishna Chaitanya Bulusu, Miguel Bordallo López, Janne Heikkilä, Markku Juntti, Olli Silvén |
ICC | 5 |
| 2024 | Mitigating SAR Out-of-Distribution Overconfidence Based on Evidential UncertaintyabstractSynthetic aperture radar (SAR) automatic target recognition (ATR) is extensively applied in both military and civilian sectors. Nevertheless, test and training data distribution may differ in the open world. Therefore, SAR out-of-distribution (OOD) detection is important because it enhances the reliability and adaptability of SAR systems. However, most OOD detection models are based on maximum likelihood estimation (MLE) and overlook the impact of data uncertainty, leading to overconfidence output for both in-distribution (ID) and OOD data. To address this issue, we consider the effect of data uncertainty on prediction probabilities, treating these probabilities as random variables and modeling them using Dirichlet distribution. Building on this, we propose an evidential uncertainty aware mean squared error (UMSE) loss function to guide the model in learning highly distinguishable output between ID and OOD data. Furthermore, to comprehensively evaluate OOD detection performance, we have compiled and organized some publicly available data and constructed a new SAR OOD detection dataset named SAR-OOD. Experimental results on SAR-OOD demonstrate that the UMSE approach achieves state-of-the-art (SOTA) performance. The code and data are available at:https://github.com/Xiaoyan-Zhou/UMSE-SAR-OOD-Detection. Tao Tang 0006, Zhongzhen Sun, Gangyao Kuang, Janne Heikkilä, Li Liu 0002 |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Cascaded and Generalizable Neural Radiance Fields for Fast View SynthesisabstractWe present CG-NeRF, a cascade and generalizable neural radiance fields method for view synthesis. Recent generalizing view synthesis methods can render high-quality novel views using a set of nearby input views. However, the rendering speed is still slow due to the nature of uniformly-point sampling of neural radiance fields. Existing scene-specific methods can train and render novel views efficiently but can not generalize to unseen data. Our approach addresses the problems of fast and generalizing view synthesis by proposing two novel modules: a coarse radiance fields predictor and a convolutional-based neural renderer. This architecture infers consistent scene geometry based on the implicit neural fields and renders new views efficiently using a single GPU. We first train CG-NeRF on multiple 3D scenes of the DTU dataset, and the network can produce high-quality and accurate novel views on unseen real and synthetic data using only photometric losses. Moreover, our method can leverage a denser set of reference images of a single scene to produce accurate novel views without relying on additional explicit representations and still maintains the high-speed rendering of the pre-trained model. Experimental results show that CG-NeRF outperforms state-of-the-art generalizable neural rendering methods on various synthetic and real datasets. Phong Nguyen 0001, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Toward Verifiable and Reproducible Human Evaluation for Text-to-Image GenerationabstractHuman evaluation is critical for validating the performance of text-to-image generative models, as this highly cognitive process requires deep comprehension of text and images. However, our survey of 37 recent papers reveals that many works rely solely on automatic measures (e.g., FID) or perform poorly described human evaluations that are not reliable or repeatable. This paper proposes a standardized and well-defined human evaluation protocol to facilitate verifiable and reproducible human evaluation in future works. In our pilot data collection, we experimentally show that the current automatic measures are incompatible with human perception in evaluating the performance of the text-to-image generation results. Furthermore, we provide insights for designing human evaluation experiments reliably and conclusively. Finally, we make several resources publicly available to the community to facilitate easy and fast implementations. Mayu Otani, Riku Togashi, Yu Sawai, Ryosuke Ishigami, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Shin'ichi Satoh 0001 |
CVPR | 7 |
| 2023 | Evidential Uncertainty and Diversity Guided Active Learning for Scene Graph Generation
Shuzhou Sun, Shuaifeng Zhi, Janne Heikkilä, Li Liu 0002 |
ICLR | 3 |
| 2023 | Partially calibrated semi-generalized pose from hybrid point correspondencesabstractWe study the problem of estimating the semi-generalized pose of a partially calibrated camera, i.e., the pose of a perspective camera with unknown focal length w.r.t. a generalized camera, from a hybrid set of 2D-2D and 2D-3D point correspondences. We study all possible camera configurations within the generalized camera system. To derive practical solvers to previously unsolved challenging configurations, we test different parameterizations as well as different solving strategies based on state-of-the-art methods for generating efficient polynomial solvers. We evaluate the three most promising solvers, i.e., the H51f solver with five 2D-2D correspondences and one 2D-3D match viewed by the same camera inside the generalized camera, the H32f solver with three 2D-2D and two 2D-3D correspondences, and the H13f solver with one 2D-2D and three 2D-3D matches, on synthetic and real data. We show that in the presence of noise in the 3D points these solvers provide better estimates than the corresponding absolute pose solvers. Snehal Bhayani, Torsten Sattler, Viktor Larsson, Janne Heikkilä, Zuzana Kukelova |
WACV | 4 |
| 2023 | Unbiased Scene Graph Generation via Two-Stage Causal ModelingabstractDespite the impressive performance of recent unbiased Scene Graph Generation (SGG) methods, the current debiasing literature mainly focuses on the long-tailed distribution problem, whereas it overlooks another source of bias, i.e., semantic confusion, which makes the SGG model prone to yield false predictions for similar relationships. In this paper, we explore a debiasing procedure for the SGG task leveraging causal inference. Our central insight is that the Sparse Mechanism Shift (SMS) in causality allows independent intervention on multiple biases, thereby potentially preserving head category performance while pursuing the prediction of high-informative tail relationships. However, the noisy datasets lead to unobserved confounders for the SGG task, and thus the constructed causal models are always causal-insufficient to benefit from SMS. To remedy this, we propose Two-stage Causal Modeling (TsCM) for the SGG task, which takes the long-tailed distribution and semantic confusion as confounders to the Structural Causal Model (SCM) and then decouples the causal intervention into two stages. The first stage is causal representation learning, where we use a novel Population Loss (P-Loss) to intervene in the semantic confusion confounder. The second stage introduces the Adaptive Logit Adjustment (AL-Adjustment) to eliminate the long-tailed distribution confounder to complete causal calibration learning. These two stages are model agnostic and thus can be used in any SGG model that seeks unbiased predictions. Comprehensive experiments conducted on the popular SGG backbones and benchmarks show that our TsCM can achieve state-of-the-art performance in terms of mean recall rate. Furthermore, TsCM can maintain a higher recall rate than other debiasing methods, which indicates that our method can achieve a better tradeoff between head and tail relationships. Shuzhou Sun, Shuaifeng Zhi, Qing Liao 0001, Janne Heikkilä, Li Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | SC6D: Symmetry-agnostic and Correspondence-free 6D Object Pose EstimationabstractThis paper presents an efficient symmetry-agnostic and correspondence-free framework, referred to as SC6D, for 6D object pose estimation from a single monocular RGB image. SC6D requires neither the 3D CAD model of the object nor any prior knowledge of the symmetries. The pose estimation is decomposed into three sub-tasks: a) object 3D rotation representation learning and matching; b) estimation of the 2D location of the object center; and c) scale-invariant distance estimation (the translation along the z-axis) via classification. SC6D is evaluated on three benchmark datasets, T-LESS, YCB-V, and ITODD, and results in state-of-the-art performance on the T-LESS dataset. More-over, SC6D is computationally much more efficient than the previous state-of-the-art method SurfEmb. The implementation and pre-trained models are publicly available at https://github.com/dingdingcai/SC6D-pose. Dingding Cai, Janne Heikkilä, Esa Rahtu |
3DV | 2 |
| 2022 | PatchFlow: A Two-Stage Patch-Based Approach for Lightweight Optical Flow Estimation
Ahmed Alhawwary, Janne Mustaniemi, Janne Heikkilä |
ACCV (7) | 3 |
| 2022 | OVE6D: Object Viewpoint Encoding for Depth-based 6D Object Pose EstimationabstractThis paper proposes a universal framework, called OVE6D, for model-based 6D object pose estimation from a single depth image and a target object mask. Our model is trained using purely synthetic data rendered from ShapeNet, and, unlike most of the existing methods, it generalizes well on new real-world objects without any fine-tuning. We achieve this by decomposing the 6D pose into viewpoint, in-plane rotation around the camera optical axis and translation, and introducing novel lightweight modules for estimating each component in a cascaded manner. The resulting network contains less than 4M parameters while demon-strating excellent performance on the challenging T-LESS and Occluded LINEMOD datasets without any dataset-specific training. We show that OVE6D outperforms some contemporary deep learning-based pose estimation methods specifically trained for individual objects or datasets with real-world training data. The implementation is available at https://github.com/dingdingcai/OVE6D-pose. Dingding Cai, Janne Heikkilä, Esa Rahtu |
CVPR | 2 |
| 2022 | Optimal Correction Cost for Object Detection EvaluationabstractMean Average Precision (mAP) is the primary evaluation measure for object detection. Although object detection has a broad range of applications, mAP evaluates detectors in terms of the performance of ranked instance retrieval. Such the assumption for the evaluation task does not suit some downstream tasks. To alleviate the gap between downstream tasks and the evaluation scenario, we propose Optimal Correction Cost (OC-cost), which assesses detection accuracy at image level. OC-cost computes the cost of correcting detections to ground truths as a measure of accuracy. The cost is obtained by solving an optimal transportation problem between the detections and the ground truths. Unlike mAp, OC-cost is designed to penalize false positive and false negative detections properly, and every image in a dataset is treated equally. Our experimental result validates that OCscost has better agreement with human preference than a ranking-based measure, i.e., mAP for a single image. We also show that detectors' rankings by OC-cost are more consistent on different data splits than mAP. Our goal is not to replace mAP with OC-cost but provide an additional tool to evaluate detectors from another aspect. To help future researchers and developers choose a target measure, we provide a series of experiments to clarify how mAP and OC-cost differ. Mayu Otani, Riku Togashi, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Shin'ichi Satoh 0001 |
CVPR | 5 |
| 2022 | AxIoU: An Axiomatically Justified Measure for Video Moment RetrievalabstractEvaluation measures have a crucial impact on the direction of research. Therefore, it is of utmost importance to develop appropriate and reliable evaluation measures for new applications where conventional measures are not well suited. Video Moment Retrieval (VMR) is one such application, and the current practice is to use R@K,$\theta$for evaluating VMR systems. However, this measure has two disadvantages. First, it is rank-insensitive: It ignores the rank positions of successfully localised moments in the top-K ranked list by treating the list as a set. Second, it binarizes the Intersection over Union (IoU) of each retrieved video moment using the threshold$\theta$and thereby ignoring fine-grained localisation quality of ranked moments. We propose an alternative measure for evaluating VMR, called Average Max IoU (AxIoU), which is free from the above two problems. We show that AxIoU satisfies two important axioms for VMR evaluation, namely, Invariance against Redundant Moments and Monotonicity with respect to the Best Moment, and also that R@ K,$\theta$satisfies the first axiom only. We also empirically examine how Ax-IoU agrees with R@K,$\theta$, as well as its stability with respect to change in the test data and human-annotated temporal boundaries. Riku Togashi, Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Tetsuya Sakai |
CVPR | 5 |
| 2022 | Free-Viewpoint RGB-D Human Performance Capture and Rendering
Phong Nguyen 0001, Nikolaos Sarafianos, Christoph Lassner, Janne Heikkilä, Tony Tung |
ECCV (16) | 4 |
| 2022 | Lightweight Monocular Depth with a Novel Neural Architecture Search MethodabstractThis paper presents a novel neural architecture search method, called LiDNAS, for generating lightweight monocular depth estimation models. Unlike previous neural architecture search (NAS) approaches, where finding optimized networks is computationally demanding, the introduced novel Assisted Tabu Search leads to efficient architecture exploration. Moreover, we construct the search space on a pre-defined backbone network to balance layer diversity and search space size. The LiDNAS method outperforms the state-of-the-art NAS approach, proposed for disparity and depth estimation, in terms of search efficiency and output model performance. The LiDNAS optimized models achieve result superior to compact depth estimation state-of-the-art on NYU-Depth-v2, KITTI, and ScanNet, while being 7%-500% more compact in size, i.e the number of model parameters. Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
WACV | 5 |
| 2021 | Monocular Depth Estimation Primed by Salient Point Detection and Normalized Hessian LossabstractDeep neural networks have recently thrived on single image depth estimation. That being said, current developments on this topic highlight an apparent compromise between accuracy and network size. This work proposes an accurate and lightweight framework for monocular depth estimation based on a self-attention mechanism stemming from salient point detection. Specifically, we utilize a sparse set of keypoints to train a FuSaNet model that consists of two major components: Fusion-Net and Saliency-Net. In addition, we introduce a normalized Hessian loss term invariant to scaling and shear along the depth direction, which is shown to substantially improve the accuracy. The proposed method achieves state-of-the-art results on NYU-Depth-v2 and KITTI while using 3.1-38.4 times smaller model in terms of the number of parameters than baseline approaches. Experiments on the SUN-RGBD further demonstrate the generalizability of the proposed method. Lam Huynh, Matteo Pedone, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
3DV | 6 |
| 2021 | RGBD-Net: Predicting Color and Depth Images for Novel Views SynthesisabstractWe propose a new cascaded architecture for novel view synthesis, called RGBD-Net, which consists of two core components: a hierarchical depth regression network and a depth-aware generator network. The former one predicts depth maps of the target views by using adaptive depth scaling, while the latter one leverages the predicted depths and renders spatially and temporally consistent target images. In the experimental evaluation on standard datasets, RGBD-Net not only outperforms the state-of-the-art by a clear margin, but it also generalizes well to new scenes without per-scene optimization. Moreover, we show that RGBD-Net can be optionally trained without depth supervision while still retaining high-quality rendering. Thanks to the depth regression network, RGBD-Net can be also used for creating dense 3D point clouds that are more accurate than those produced by some state-of-the-art multi-view stereo methods. Phong Nguyen 0001, Animesh Karnewar, Lam Huynh, Esa Rahtu, Jiri Matas, Janne Heikkilä |
3DV | 6 |
| 2021 | Calibrated and Partially Calibrated Semi-Generalized HomographiesabstractIn this paper, we propose the first minimal solutions for estimating the semi-generalized homography given a perspective and a generalized camera. The proposed solvers use five 2D-2D image point correspondences induced by a scene plane. One group of solvers assumes the perspective camera to be fully calibrated, while the other estimates the unknown focal length together with the absolute pose parameters. This setup is particularly important in structure-from-motion and visual localization pipelines, where a new camera is localized in each step with respect to a set of known cameras and 2D-3D correspondences might not be available. Thanks to a clever parametrization and the elimination ideal method, our solvers only need to solve a univariate polynomial of degree five or three, respectively a system of polynomial equations in two variables. All proposed solvers are stable and efficient as demonstrated by a number of synthetic and real-world experiments. Snehal Bhayani, Torsten Sattler, Daniel Barath, Patrik Beliansky, Janne Heikkilä, Zuzana Kukelova |
ICCV | 5 |
| 2021 | Boosting Monocular Depth Estimation with Lightweight 3D Point FusionabstractIn this paper, we propose enhancing monocular depth estimation by adding 3D points as depth guidance. Unlike existing depth completion methods, our approach performs well on extremely sparse and unevenly distributed point clouds, which makes it agnostic to the source of the 3D points. We achieve this by introducing a novel multi-scale 3D point fusion network that is both lightweight and efficient. We demonstrate its versatility on two different depth estimation problems where the 3D points have been acquired with conventional structure-from-motion and Li-DAR. In both cases, our network performs on par with state-of-the-art depth completion methods and achieves significantly higher accuracy when only a small number of points is used while being more compact in terms of the number of parameters. We show that our method outperforms some contemporary deep learning based multi-view stereo and structure-from-motion methods both in accuracy and in compactness. Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
ICCV | 5 |
| 2020 | Sequential View Synthesis with Transformer
Phong Nguyen 0001, Lam Huynh, Esa Rahtu, Janne Heikkilä |
ACCV (4) | 4 |
| 2020 | LSD_2 - Joint Denoising and Deblurring of Short and Long Exposure Images with CNNs
Janne Mustaniemi, Juho Kannala, Jiri Matas, Simo Särkkä, Janne Heikkilä |
BMVC | 5 |
| 2020 | Uncovering Hidden Challenges in Query-Based Video Moment Retrieval
Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä |
BMVC | 4 |
| 2020 | A Sparse Resultant Based Method for Efficient Minimal SolversabstractMany computer vision applications require robust and efficient estimation of camera geometry. The robust estimation is usually based on solving camera geometry problems from a minimal number of input data measurements, i.e. solving minimal problems in a RANSAC framework. Minimal problems often result in complex systems of polynomial equations. Many state-of-the-art efficient polynomial solvers to these problems are based on Gröbner basis and the action-matrix method that has been automatized and highly optimized in recent years. In this paper we study an alternative algebraic method for solving systems of polynomial equations, i.e., the sparse resultant-based method and propose a novel approach to convert the resultant constraint to an eigenvalue problem. This technique can significantly improve the efficiency and stability of existing resultant-based solvers. We applied our new resultant-based method to a large variety of computer vision problems and show that for most of the considered problems, the new method leads to solvers that are the same size as the the best available Gröbner basis solvers and of similar accuracy. For some problems the new sparse-resultant based method leads to even smaller and more stable solvers than the state-of-the-art Gröbner basis solvers. Our new method can be fully automatized and incorporated into existing tools for automatic generation of efficient polynomial solvers and as such it represents a competitive alternative to popular Gröbner basis methods for minimal problems in computer vision. Snehal Bhayani, Zuzana Kukelova, Janne Heikkilä |
CVPR | 3 |
| 2020 | Guiding Monocular Depth Estimation Using Depth-Attention Volume
Lam Huynh, Phong Nguyen 0001, Jiri Matas, Esa Rahtu, Janne Heikkilä |
ECCV (26) | 5 |
| 2020 | Computing stable resultant-based minimal solvers by hiding a variableabstractMany computer vision applications require robust and efficient estimation of camera geometry. The robust estimation is usually based on solving camera geometry problems from a minimal number of input data measurements, i.e., solving minimal problems, in a RANSAC-style framework. Minimal problems often result in complex systems of polynomial equations. The existing state-of-the-art methods for solving such systems are either based on Gröbner bases and the action matrix method, which have been extensively studied and optimized in the recent years or recently proposed approach based on a resultant computation using an extra variable. In this paper, we study an interesting alternative resultant-based method for solving sparse systems of polynomial equations by hiding one variable. This approach results in a larger eigenvalue problem than the action matrix and extra variable resultant-based methods; however, it does not need to compute an inverse or elimination of large matrices that may be numerically unstable. The proposed approach includes several improvements to the standard sparse resultant algorithms, which significantly improves the efficiency and stability of the hidden variable resultant-based solvers as we demonstrate on several interesting computer vision problems. We show that for the studied problems, our sparse resultant based approach leads to more stable solvers than the state-of-the-art Gröbner basis as well as existing resultant-based solvers, especially in close to critical configurations. Our new method can be fully automated and incorporated into existing tools for the automatic generation of efficient minimal solvers. Snehal Bhayani, Zuzana Kukelova, Janne Heikkilä |
ICPR | 3 |
| 2020 | Learning non-rigid surface reconstruction from spatia-temporal image patchesabstractWe present a method to reconstruct a dense spatiotemporal depth map of a non-rigidly deformable object directly from a video sequence. The estimation of depth is performed locally on spatio-temporal patches of the video, and then the full depth video of the entire shape is recovered by combining them together. Since the geometric complexity of a local spatiotemporal patch of a deforming non-rigid object is often simple enough to be faithfully represented with a parametric model, we artificially generate a database of small deforming rectangular meshes rendered with different material properties and light conditions, along with their corresponding depth videos, and use such data to train a convolutional neural network. We tested our method on both synthetic and Kinect data and experimentally observed that the reconstruction error is significantly lower than the one obtained using conventional non-rigid structure from motion approaches. Matteo Pedone, Abdelrahman Mostafa, Janne Heikkilä |
ICPR | 3 |
| 2019 | Rethinking the Evaluation of Video SummariesabstractVideo summarization is a technique to create a short skim of the original video while preserving the main stories/content. There exists a substantial interest in automatizing this process due to the rapid growth of the available material. The recent progress has been facilitated by public benchmark datasets, which enable easy and fair comparison of methods. Currently the established evaluation protocol is to compare the generated summary with respect to a set of reference summaries provided by the dataset. In this paper, we will provide in-depth assessment of this pipeline using two popular benchmark datasets. Surprisingly, we observe that randomly generated summaries achieve comparable or better performance to the state-of-the-art. In some cases, the random summaries outperform even the human generated summaries in leave-one-out experiments. Moreover, it turns out that the video segmentation, which is often considered as a fixed pre-processing method, has the most significant impact on the performance measure. Based on our observations, we propose alternative approaches for assessing the importance scores as well as an intuitive visualization of correlation between the estimated scoring and human annotations. Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä |
CVPR | 4 |
| 2019 | Gyroscope-Aided Motion Deblurring with Deep NetworksabstractWe propose a deblurring method that incorporates gyroscope measurements into a convolutional neural network (CNN). With the help of such measurements, it can handle extremely strong and spatially-variant motion blur. At the same time, the image data is used to overcome the limitations of gyro-based blur estimation. To train our network, we also introduce a novel way of generating realistic training data using the gyroscope. The evaluation shows a clear improvement in visual quality over the state-of-the-art while achieving real-time performance. Furthermore, the method is shown to improve the performance of existing feature detectors and descriptors against the motion blur. Janne Mustaniemi, Juho Kannala, Simo Särkkä, Jiri Matas, Janne Heikkilä |
WACV | 5 |
| 2019 | 3D Multi-Resolution Optical Flow Analysis of Cardiovascular Pulse Propagation in Human BrainabstractThe brain is cleaned from waste by glymphatic clearance serving a similar purpose as the lymphatic system in the rest of the body. Impairment of the glymphatic brain clearance precedes protein accumulation and reduced cognitive function in Alzheimer's disease (AD). Cardiovascular pulsations are a primary driving force of the glymphatic brain clearance. We developed a method to quantify cardiovascular pulse propagation in the human brain with magnetic resonance encephalography (MREG). We extended a standard optical flow estimation method to three spatial dimensions, with a multi-resolution processing scheme. We added application-specific criteria for discarding inaccurate results. With the proposed method, it is now possible to estimate the propagation of cardiovascular pulse wavefronts from the whole brain MREG data sampled at 10 Hz. The results show that on average the cardiovascular pulse propagates from major arteries via cerebral spinal fluid spaces into all tissue compartments in the brain. We present an example, that cardiovascular pulsations are significantly altered in AD: coefficient of variation and sample entropy of the pulse propagation speed in the lateral ventricles change in AD. These changes are in line with the theory of glymphatic clearance impairment in AD. The proposed non-invasive method can assess a performance indicator related to the glymphatic clearance in the human brain. Zalan Rajna, Lauri Raitamaa, Timo Tuovinen, Janne Heikkilä, Vesa Kiviniemi, Tapio Seppänen |
IEEE Trans. Medical Imaging | 4 |
| 2019 | Dynamic Texture Classification Using Unsupervised 3D Filter Learning and Local Binary EncodingabstractLocal binary descriptors, such as local binary pattern (LBP) and its various variants, have been studied extensively in texture and dynamic texture analysis due to their outstanding characteristics, such as grayscale invariance, low computational complexity and good discriminability. Most existing local binary feature extraction methods extract spatio-temporal features from three orthogonal planes of a spatio-temporal volume by viewing a dynamic texture in 3D space. For a given pixel in a video, only a proportion of its surrounding pixels is incorporated in the local binary feature extraction process. We argue that the ignored pixels contain discriminative information that should be explored. To fully utilize the information conveyed by all the pixels in a local neighborhood, we propose extracting local binary features from the spatio-temporal domain with 3D filters that are learned in an unsupervised manner so that the discriminative features along both the spatial and temporal dimensions are captured simultaneously. The proposed approach consists of three components: 1) 3D filtering; 2) binary hashing; and 3) joint histogramming. Densely sampled 3D blocks of a dynamic texture are first normalized to have zero mean and are then filtered by 3D filters that are learned in advance. To preserve more of the structure information, the filter response vectors are decomposed into two complementary components, namely, the signs and the magnitudes, which are further encoded separately into binary codes. The local mean pixels of the 3D blocks are also converted into binary codes. Finally, three types of binary codes are combined via joint or hybrid histograms for the final feature representation. Extensive experiments are conducted on three commonly used dynamic texture databases: 1) UCLA; 2) DynTex; and 3) YUVL. The proposed method provides comparable results to, and even outperforms, many state-of-the-art methods. Xiaochao Zhao, Yaping Lin, Li Liu 0002, Janne Heikkilä, Wenming Zheng |
IEEE Trans. Multim. | 4 |
| 2018 | Fast Motion Deblurring for Feature Detection and Matching Using Inertial MeasurementsabstractMany computer vision and image processing applications rely on local features. It is well-known that motion blur decreases the performance of traditional feature detectors and descriptors. We propose an inertial-based deblurring method for improving the robustness of existing feature detectors and descriptors against the motion blur. Unlike most deblurring algorithms, the method can handle spatially-variant blur and rolling shutter distortion. Furthermore, it is capable of running in real-time contrary to state-of-the-art algorithms. The limitations of inertial-based blur estimation are taken into account by validating the blur estimates using image data. The evaluation shows that when the method is used with traditional feature detector and descriptor, it increases the number of detected keypoints, provides higher repeatability and improves the localization accuracy. We also demonstrate that such features will lead to more accurate and complete reconstructions when used in the application of 3D visual reconstruction. Janne Mustaniemi, Juho Kannala, Simo Särkkä, Jiri Matas, Janne Heikkilä |
ICPR | 5 |
| 2018 | Accurate 3-D Reconstruction with RGB-D Cameras using Depth Map Fusion and Pose RefinementabstractDepth map fusion is an essential part in both stereo and RGB-D based 3- D reconstruction pipelines. Whether produced with a passive stereo reconstruction or using an active depth sensor, such as Microsoft Kinect, the depth maps have noise and may have poor initial registration. In this paper, we introduce a method which is capable of handling outliers, and especially, even significant registration errors. The proposed method first fuses a sequence of depth maps into a single non-redundant point cloud so that the redundant points are merged together by giving more weight to more certain measurements. Then, the original depth maps are re-registered to the fused point cloud to refine the original camera extrinsic parameters. The fusion is then performed again with the refined extrinsic parameters. This procedure is repeated until the result is satisfying or no significant changes happen between iterations. The method is robust to outliers and erroneous depth measurements as well as even significant depth map registration errors due to inaccurate initial camera poses. Markus Ylimäki, Janne Heikkilä, Juho Kannala |
ICPR | 2 |
| 2018 | Dynamic Texture Recognition Using Volume Local Binary Count Patterns With an Application to 2D Face Spoofing DetectionabstractIn this paper, a local spatiotemporal descriptor, namely, the volume local binary count (VLBC), is proposed for the representation and recognition of dynamic texture. This descriptor, which is similar in spirit to the volume local binary pattern (VLBP), extracts histograms of thresholded local spatiotemporal volumes using both appearance and motion features to describe dynamic texture. Unlike VLBP using binary encoding, VLBC does not exploit the local structure information and only counts the number of 1s in the thresholded codes. Thus, VLBC can include more neighboring pixels without exponentially increasing the feature dimension as VLBP does. Furthermore, a completed version of VLBC (CVLBC) is also proposed to enhance the performance of dynamic texture recognition with additional information about local contrast and central pixel intensities. The proposed method is not only efficient to compute but also effective for dynamic texture representation. In experiments with three dynamic texture databases, namely, UCLA, DynTex, and DynTex++, the proposed method produces classification rates that are comparable to those produced by the state-of-the-art approaches. In addition to dynamic texture recognition, we propose utilizing CVLBC for 2-D face spoofing detection. As an effective spatiotemporal descriptor, CVLBC can well describe the differences between facial videos of valid users and impostors, thus achieving good performance for face spoofing detection. For comparison with other methods, the proposed method is evaluated on three face antispoofing databases: Print-Attack, Replay-Attack, and CAS Face Antispoofing. The experimental results demonstrate the effectiveness of CVLBC for 2-D face spoofing detection. Xiaochao Zhao, Yaping Lin, Janne Heikkilä |
IEEE Trans. Multim. | 3 |
| 2017 | Using Sparse Elimination for Solving Minimal Problems in Computer VisionabstractFinding a closed form solution to a system of polynomial equations is a common problem in computer vision as well as in many other areas of engineering and science. Gröbner basis techniques are often employed to provide the solution, but implementing an efficient Gröbner basis solver to a given problem requires strong expertise in algebraic geometry. One can also convert the equations to a polynomial eigenvalue problem (PEP) and solve it using linear algebra, which is a more accessible approach for those who are not so familiar with algebraic geometry. In previous works PEP has been successfully applied for solving some relative pose problems in computer vision, but its wider exploitation is limited by the problem of finding a compact monomial basis. In this paper, we propose a new algorithm for selecting the basis that is in general more compact than the basis obtained with a state-of-the-art algorithm making PEP a more viable option for solving polynomial equations. Another contribution is that we present two minimal problems for camera self-calibration based on homography, and demonstrate experimentally using synthetic and real data that our algorithm can provide a numerically stable solution to the camera focal length from two homographies of unknown planar scene. Janne Heikkilä |
ICCV | 1 |
| 2017 | Dynamic texture recognition using multiscale PCA-learned filtersabstractIn this paper, we propose a novel method for dynamic texture recognition using multiscale PCA-learned filters. PCA is utilized to learn multiscale filters from image sequences on three orthogonal planes (XY, XT and YT). Filter responses that contain both spatial and temporal information at multiple scales are then encoded into a descriptor named MPCAF-TOP. The proposed method is simple to derive and implement, and also very effective for dynamic texture recognition. The proposed method is evaluated on two benchmark databases, namely UCLA and DynTex++. Experimental results show that the proposed approach is comparable to state-of-the-art methods. Xiaochao Zhao, Yaping Lin, Janne Heikkilä |
ICIP | 3 |
| 2017 | Inertial-based scale estimation for structure from motion on mobile devicesabstractStructure from motion algorithms have an inherent limitation that the reconstruction can only be determined up to the unknown scale factor. Modern mobile devices are equipped with an inertial measurement unit (IMU), which can be used for estimating the scale of the reconstruction. We propose a method that recovers the metric scale given inertial measurements and camera poses. In the process, we also perform a temporal and spatial alignment of the camera and the IMU. Therefore, our solution can be easily combined with any existing visual reconstruction software. The method can cope with noisy camera pose estimates, typically caused by motion blur or rolling shutter artifacts, via utilizing a Rauch-Tung-Striebel (RTS) smoother. Furthermore, the scale estimation is performed in the frequency domain, which provides more robustness to inaccurate sensor time stamps and noisy IMU samples than the previously used time domain representation. In contrast to previous methods, our approach has no parameters that need to be tuned for achieving a good performance. In the experiments, we show that the algorithm outperforms the state-of-the-art in both accuracy and convergence speed of the scale estimate. The accuracy of the scale is around 1% from the ground truth depending on the recording. We also demonstrate that our method can improve the scale accuracy of the Project Tango's build-in motion tracking. Janne Mustaniemi, Juho Kannala, Simo Särkkä, Jiri Matas, Janne Heikkilä |
IROS | 5 |
| 2016 | Video Summarization Using Deep Semantic Features
Mayu Otani, Yuta Nakashima, Esa Rahtu, Janne Heikkilä, Naokazu Yokoya |
ACCV (5) | 4 |
| 2016 | Cell proposal network for microscopy image analysisabstractRobust cell detection plays a key role in the development of reliable methods for automated analysis of microscopy images. It is a challenging problem due to low contrast, variable fluorescence, weak boundaries, conjoined and overlapping cells, causing most cell detection methods to fail in difficult situations. One approach for overcoming these challenges is to use cell proposals, which enable the use of more advanced features from ambiguous regions and/or information from adjacent frames to make better decisions. However, most current methods rely on simple proposal generation and scoring methods, which limits the performance they can reach. In this paper, we propose a convolutional neural network based method which generates cell proposals to facilitate cell detection, segmentation and tracking. We compare our method against commonly used proposal generation and scoring methods and show that our method generates significantly better proposals, and achieves higher final recall and average precision. Saad Ullah Akram, Juho Kannala, Lauri Eklund, Janne Heikkilä |
ICIP | 4 |
| 2016 | Deep learning for magnification independent breast cancer histopathology image classificationabstractMicroscopic analysis of breast tissues is necessary for a definitive diagnosis of breast cancer which is the most common cancer among women. Pathology examination requires time consuming scanning through tissue images under different magnification levels to find clinical assessment clues to produce correct diagnoses. Advances in digital imaging techniques offers assessment of pathology images using computer vision and machine learning methods which could automate some of the tasks in the diagnostic pathology workflow. Such automation could be beneficial to obtain fast and precise quantification, reduce observer variability, and increase objectivity. In this work, we propose to classify breast cancer histopathology images independent of their magnifications using convolutional neural networks (CNNs). We propose two different architectures; single task CNN is used to predict malignancy and multi-task CNN is used to predict both malignancy and image magnification level simultaneously. Evaluations and comparisons with previous results are carried out on BreaKHis dataset. Experimental results show that our magnification independent CNN approach improved the performance of magnification specific model. Our results in this limited set of training data are comparable with previous state-of-the-art results obtained by hand-crafted features. However, unlike previous methods, our approach has potential to directly benefit from additional training data, and such additional data could be captured with same or different magnification levels than previous data. Neslihan Bayramoglu, Juho Kannala, Janne Heikkilä |
ICPR | 3 |
| 2016 | Geometry based exhaustive line correspondence determinationabstractIn this paper we propose a purely geometric approach to establish correspondence between 3D line segments in a given model and 2D line segments detected in an image. Contrary to the existing methods which use strong assumptions on camera pose, we perform exhaustive search in order to compute maximum number of geometrically permitted correspondences between a 3D model and 2D lines. We present a novel theoretical framework in which we sample the space of camera axis direction (which is bounded and hence can be densely sampled unlike the unbounded space of camera position) and show that geometric constraints arising from it reduce rest of the computation to simple operations of finding camera position as the intersection of 3 planes. These geometric constraints can be represented using indexed arrays which accelerate it further. The algorithm returns all sets of correspondences and associated camera poses having high geometric consensus. The obtained experimental results show that our method has better asymptotic behavior than conventional approach. We also show that with the inclusion of additional sensor information our method can be used to initialize pose in just few seconds in many practical situations. K. K. Srikrishna Bhat, Utpala Musti, Janne Heikkilä |
ICRA | 3 |
| 2016 | Forget the checkerboard: Practical self-calibration using a planar sceneabstractWe introduce a camera self-calibration method using a planar scene of unknown texture. Planar surfaces are everywhere but checkerboards are not, thus the method can be more easily applied outside the lab. We demonstrate that the accuracy is equivalent to a checkerboard-based calibration, so there is no need for printing checkerboards any more. Moreover, the use of a planar scene provides improved robustness and stronger constraints than a self-calibration with an arbitrary scene. We utilize a closed-form initialization of the focal length with minimal and practical assumptions. The method recovers the intrinsic and extrinsic parameters of the camera and the metric structure of the planar scene. The method is implemented in a real-time application for non-expert users that provides an easy and practical process to obtain high accuracy calibrations. Daniel Herrera C., Juho Kannala, Janne Heikkilä |
WACV | 3 |
| 2016 | MORE - a multimodal observation and analysis system for social interaction research
Anja Keskinarkaus, Sami Huttunen, Antti Siipo, Jukka Holappa, Magda Laszlo, Ilkka Juuso, Eero Väyrynen, Janne Heikkilä, Matti Lehtihalmes, Tapio Seppänen, Seppo J. Laukka |
Multim. Tools Appl. | 8 |
| 2016 | Guest Editorial: Immersive Audio/Visual Systems
Lei Xie 0001, Longbiao Wang, Janne Heikkilä, Peng Zhang 0005 |
Multim. Tools Appl. | 3 |
| 2016 | Parallax correction via disparity estimation in a multi-aperture camera
Janne Mustaniemi, Juho Kannala, Janne Heikkilä |
Mach. Vis. Appl. | 3 |
| 2015 | Human Epithelial Type 2 cell classification with convolutional neural networksabstractAutomated cell classification in Indirect Immunofluorescence (IIF) images has potential to be an important tool in clinical practice and research. This paper presents a framework for classification of Human Epithelial Type 2 cell IIF images using convolutional neural networks (CNNs). Previuos state-of-the-art methods show classification accuracy of 75.6% on a benchmark dataset. We conduct an exploration of different strategies for enhancing, augmenting and processing training data in a CNN framework for image classification. Our proposed strategy for training data and pre-training and fine-tuning the CNN network led to a significant increase in the performance over other approaches that have been used until now. Specifically, our method achieves a 80.25% classification accuracy. Source code and models to reproduce the experiments in the paper is made publicly available. Neslihan Bayramoglu, Juho Kannala, Janne Heikkilä |
BIBE | 3 |
| 2015 | Disparity Estimation for Image Fusion in a Multi-aperture Camera
Janne Mustaniemi, Juho Kannala, Janne Heikkilä |
CAIP (2) | 3 |
| 2015 | Optimizing the Accuracy and Compactness of Multi-view Reconstructions
Markus Ylimäki, Juho Kannala, Janne Heikkilä |
CAIP (2) | 3 |
| 2015 | A novel feature descriptor based on microscopy image statisticsabstractIn this paper, we propose a novel feature description algorithm based on image statistics. The pipeline first performs independent component analysis on training image patches to obtain basis vectors (filters) for a lower dimensional representation. Then for a given image, a set of filter responses at each pixel is computed. Finally, a histogram representation, which considers the signs and magnitudes of the responses as well as the number of filters, is applied on local image patches. We propose to apply this idea to a microscopy image pixel identification system based on a learning framework. Experimental results show that the proposed algorithm performs better than the state-of-the-art descriptors in biomedical images of different microscopy modalities. Neslihan Bayramoglu, Juho Kannala, Malin Akerfelt, Mika Kaakinen, Lauri Eklund, Matthias Nees, Janne Heikkilä |
ICIP | 7 |
| 2015 | Focal length change compensation for monocular slamabstractIn this paper, we propose a method for handling focal length changes in the SLAM algorithm. Our method is designed as a pre-processing step to first estimate the change of the camera focal length, and then compensate for the zooming effects before running the actual SLAM algorithm. By using our method, camera zooming can be used in the existing SLAM algorithms with minor modifications. In the experiments, the effectiveness of the proposed method was quantitatively evaluated. The results indicate that the method can successfully deal with abrupt changes of the camera focal length. Takafumi Taketomi, Janne Heikkilä |
ICIP | 2 |
| 2015 | Zoom factor compensation for monocular SLAMabstractSLAM algorithms are widely used in augmented reality applications for registering virtual objects. Most SLAM algorithms estimate camera poses and 3D positions of feature points using known intrinsic camera parameters that are calibrated and fixed in advance. This assumption means that the algorithm does not allow changing the intrinsic camera parameters during runtime. We propose a method for handling focal length changes in the SLAM algorithm. Our method is designed as a pre-processing step for the SLAM algorithm input. In our method, the change of the focal length is estimated before the tracking process of the SLAM algorithm. Camera zooming effects in the input camera images are compensated for by using the estimated focal length change. By using our method, camera zooming can be used in the existing SLAM algorithms such as PTAM [4] with minor modifications. In the experiment, the effectiveness of the proposed method was quantitatively evaluated. The results indicate that the method can successfully deal with abrupt changes of the camera focal length. Takafumi Taketomi, Janne Heikkilä |
VR | 2 |
| 2015 | Fast and accurate multi-view reconstruction by multi-stage prioritised matchingabstractIn this study, the authors propose a multi‐view stereo reconstruction method which creates a three‐dimensional point cloud of a scene from multiple calibrated images captured from different viewpoints. The method is based on a prioritised match expansion technique, which starts from a sparse set of seed points, and iteratively expands them into neighbouring areas by using multiple expansion stages. Each seed point represents a surface patch and has a position and a surface normal vector. The location and surface normal of the seeds are optimised using a homography‐based local image alignment. The propagation of seeds is performed in a prioritised order in which the most promising seeds are expanded first and removed from the list of seeds. The first expansion stage proceeds until the list of seeds is empty. In the following expansion stages, the current reconstruction may be further expanded by finding new seeds near the boundaries of the current reconstruction. The prioritised expansion strategy allows efficient generation of accurate point clouds and their experiments show its benefits compared with non‐prioritised expansion. In addition, a comparison to the widely used patch‐based multi‐view stereo software shows that their method is significantly faster and produces more accurate and complete reconstructions. Markus Ylimäki, Juho Kannala, Jukka Holappa, Sami S. Brandt, Janne Heikkilä |
IET Comput. Vis. | 5 |
| 2015 | Quaternion Wiener Deconvolution for Noise Robust Color Image RegistrationabstractIn this letter, we propose a global method for registering color images with respect to translation. Our approach is based on the idea of representing translations as convolutions with unknown shifted delta functions, and performing Wiener deconvolution in order to recover the shift between two images. We then derive a quaternionic version of the Wiener deconvolution filter in order to register color images. The use of Wiener filter also allows us to explicitly take into account the effect of noise. We prove that the well-known algorithm of phase correlation is a special case of our method, and we experimentally demonstrate the advantages of our approach by comparing it to other known generalizations of the phase correlation algorithm. Matteo Pedone, Eduardo Bayro-Corrochano, Jan Flusser, Janne Heikkilä |
IEEE Signal Process. Lett. | 4 |
| 2015 | Registration of Images With N-Fold Dihedral BlurabstractIn this paper, we extend our recent registration method designed specifically for registering blurred images. The original method works for unknown blurs, assuming the blurring point-spread function (PSF) exhibits an N -fold rotational symmetry. Here, we also generalize the theory to the case of dihedrally symmetric blurs, which are produced by the PSFs having both rotational and axial symmetries. Such kind of blurs are often found in unfocused images acquired by digital cameras, as in out-of-focus shots the PSF typically mimics the shape of the shutter aperture. This makes our registration algorithm particularly well-suited in applications where blurred image registration must be used as a preprocess step of an image fusion algorithm, and where common registration methods fail, due to the amount of blur. We demonstrate that the proposed method leads to an improvement of the registration performance, and we show its applicability to real images by providing successful examples of blurred image registration followed by depth-of-field extension and multichannel blind deconvolution. Matteo Pedone, Jan Flusser, Janne Heikkilä |
IEEE Trans. Image Process. | 3 |
| 2014 | Line Matching and Pose Estimation for Unconstrained Model-to-Image AlignmentabstractThis paper has two contributions in the context of line based camera pose estimation, 1) We propose a purely geometric approach to establish correspondence between 3D line segments in a given model and 2D line segments detected in an image, 2) We eliminate a degenerate case due to the type of rotation representation in arguably the best line based pose estimation method currently available. For establishing line correspondences we perform exhaustive search on the space of camera pose values till we obtain a pose (position and rotation) which is geometrically consistent with the given set of 2D, 3D lines. For this highly complex search we design a strategy which performs precomputations on the 3D model using separate set of constraints on position and rotation values. During runtime, the set of different rotation values are ranked independently and combined with each position values in the order of their ranking. Then successive geometric constraints which are much simpler when compared to computing reprojection error are used to eliminate incorrect pose values. We show that the ranking of rotation values reduces the number of trials needed by a huge factor and the simple geometric constraints avoid the need for computing the reprojection error in most cases. Though the execution time for the current MATLAB implementation is far from real time requirement, our method can be accelerated significantly by exploiting simplicity and parallelizability of the operations we employ. For eliminating the degenerate case in the state of art pose estimation method, we reformulate the rotation representation. We use unit quaternions instead of CGR parameters used by the method. K. K. Srikrishna Bhat, Janne Heikkilä |
3DV | 2 |
| 2014 | DT-SLAM: Deferred Triangulation for Robust SLAMabstractObtaining a good baseline between different video frames is one of the key elements in vision-based monocular SLAM systems. However, if the video frames contain only a few 2D feature correspondences with a good baseline, or the camera only rotates without sufficient translation in the beginning, tracking and mapping becomes unstable. We introduce a real-time visual SLAM system that incrementally tracks individual 2D features, and estimates camera pose by using matched 2D features, regardless of the length of the baseline. Triangulating 2D features into 3D points is deferred until key frames with sufficient baseline for the features are available. Our method can also deal with pure rotational motions, and fuse the two types of measurements in a bundle adjustment step. Adaptive criteria for key frame selection are also introduced for efficient optimization and dealing with multiple maps. We demonstrate that our SLAM system improves camera pose estimates and robustness, even with purely rotational motions. Daniel Herrera C., Juho Kannala, Kari Pulli, Janne Heikkilä |
3DV | 5 |
| 2014 | Segmentation of Cells from Spinning Disk Confocal Images Using a Multi-stage Approach
Saad Ullah Akram, Juho Kannala, Mika Kaakinen, Lauri Eklund, Janne Heikkilä |
ACCV (3) | 5 |
| 2014 | Detection of Tumor Cell Spheroids from Co-cultures Using Phase Contrast Images and Machine Learning ApproachabstractAutomated image analysis is demanded in cell biology and drug development research. The type of microscopy is one of the considerations in the trade-offs between experimental setup, image acquisition speed, molecular labelling, resolution and quality of images. In many cases, phase contrast imaging gets higher weights in this optimization. And it comes at the price of reduced image quality in imaging 3D cell cultures. For such data, the existing state-of-the-art computer vision methods perform poorly in segmenting specific cell type. Low SNR, clutter and occlusions are basic challenges for blind segmentation approaches. In this study we propose an automated method, based on a learning framework, for detecting particular cell type in cluttered 2D phase contrast images of 3D cell cultures that overcomes those challenges. It depends on local features defined over super pixels. The method learns appearance based features, statistical features, textural features and their combinations. Also, the importance of each feature is measured by employing Random Forest classifier. Experiments show that our approach does not depend on training data and the parameters. Neslihan Bayramoglu, Mika Kaakinen, Lauri Eklund, Malin Akerfelt, Matthias Nees, Juho Kannala, Janne Heikkilä |
ICPR | 7 |
| 2014 | Emotional Valence Recognition, Analysis of Salience and Eye MovementsabstractThis paper studies the performance of recorded eye movements and computational visual attention models (i.e. saliency models) in the recognition of emotional valence of an image. In the first part of this study, it employs eye movement data (fixation & saccade) to build image content descriptors and use them with support vector machines to classify the emotional valence. In the second part, it examines if the human saliency map can be substituted with the state-of-the-art computational visual attention models in the task of valence recognition. The results indicate that the eye movement based descriptors provide significantly better performance compared to the baselines, which apply low-level visual cues (e.g. color, texture and shape). Furthermore, it will be shown that the current computational models for visual attention are not able to capture the emotional information in similar extent as the real eye movements. Hamed Rezazadegan Tavakoli, Victoria Yanulevskaya, Esa Rahtu, Janne Heikkilä, Nicu Sebe |
ICPR | 4 |
| 2013 | A Learned Joint Depth and Intensity Prior Using Markov Random FieldsabstractWe present a joint prior that takes intensity and depth information into account. The prior is defined using a flexible Field-of-Experts model and is learned from a dataBase of natural images. It is a generative model and has an efficient method for sampling. We use sampling from the model to perform in painting and up sampling of depth maps when intensity information is available. We show that including the intensity information in the prior improves the results obtained from the model. We also compare to another two-channel in painting approach and show superior results. Daniel Herrera C., Juho Kannala, Peter F. Sturm, Janne Heikkilä |
3DV | 4 |
| 2013 | Spherical Center-Surround for Video Saliency Detection Using Sparse Sampling
Hamed Rezazadegan Tavakoli, Esa Rahtu, Janne Heikkilä |
ACIVS | 3 |
| 2013 | Stochastic bottom-up fixation prediction and saccade generation
Hamed Rezazadegan Tavakoli, Esa Rahtu, Janne Heikkilä |
Image Vis. Comput. | 3 |
| 2013 | Blur Invariant Translational Image Registration for N -fold Symmetric BlursabstractIn this paper, we propose a new registration method designed particularly for registering differently blurred images. Such a task cannot be successfully resolved by traditional approaches. Our method is inspired by traditional phase correlation, which is now applied to certain blur-invariant descriptors instead of the original images. This method works for unknown blurs assuming the blurring PSF exhibits an N-fold rotational symmetry. It does not require any landmarks. We have experimentally proven its good performance, which is not dependent on the amount of blur. In this paper, we explicitly address only registration with respect to translation, but the method can be readily generalized to rotation and scaling. Matteo Pedone, Jan Flusser, Janne Heikkilä |
IEEE Trans. Image Process. | 3 |
| 2012 | Local phase quantization descriptors for blur robust and illumination invariant recognition of color textures
Matteo Pedone, Janne Heikkilä |
ICPR | 2 |
| 2012 | Robust and accurate multi-view reconstruction by prioritized matching
Markus Ylimäki, Juho Kannala, Jukka Holappa, Janne Heikkilä, Sami S. Brandt |
ICPR | 4 |
| 2012 | Local phase quantization for blur-insensitive image analysis
Esa Rahtu, Janne Heikkilä, Ville Ojansivu, Timo Ahonen |
Image Vis. Comput. | 2 |
| 2012 | Joint Depth and Color Camera Calibration with Distortion CorrectionabstractWe present an algorithm that simultaneously calibrates two color cameras, a depth camera, and the relative pose between them. The method is designed to have three key features: accurate, practical, and applicable to a wide range of sensors. The method requires only a planar surface to be imaged from various poses. The calibration does not use depth discontinuities in the depth image, which makes it flexible and robust to noise. We apply this calibration to a Kinect device and present a new depth distortion model for the depth sensor. We perform experiments that show an improved accuracy with respect to the manufacturer's calibration. Daniel Herrera C., Juho Kannala, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Accurate and Practical Calibration of a Depth and Color Camera Pair
Daniel Herrera C., Juho Kannala, Janne Heikkilä |
CAIP (2) | 3 |
| 2010 | Isotropic Granularity-tunable gradients partition (IGGP) descriptors for human detectionabstractThis paper presents a new descriptor for human detection in still images. It is referred to as isotropic granularity-tunable gradients partition (IGGP), which is extended from granularity-tunable gradients partition (GGP) descriptors. The isotropic representation is achieved by aligning the features with different orientation channels according to their principal angles. The benefits of this extension are two folds: firstly, since the partitions’ sizes of all the orientation channels are equal, the noise introduce by the small partitions in the original GGP descriptors is eliminated and the performance can be essentially improved; secondly, the integral image based fast computation is applied and more than 20 times speedup has been achieved. In addition, we introduce a new human dataset HIMA. Unlike the previous available human datasets which are mainly captured on the street views for automobile safety or robotics, HIMA dataset is captured on the outdoor work fields for industry safety. The major challenges include: extreme light conditions, occlusion and strong noise. We benchmark several promising detection systems, providing an overview of state-of-the-art performance on the HIMA set. Experimental results show that the proposed method can yield very competitive results in both the detection speed and accuracy. Yazhou Liu, Janne Heikkilä |
BMVC | 2 |
| 2010 | Spatial-Temporal Granularity-Tunable Gradients Partition (STGGP) Descriptors for Human Detection
Yazhou Liu, Shiguang Shan, Xilin Chen 0001, Janne Heikkilä, Wen Gao 0001, Matti Pietikäinen |
ECCV (1) | 4 |
| 2010 | Segmenting Salient Objects from Images and Videos
Esa Rahtu, Juho Kannala, Mikko Salo, Janne Heikkilä |
ECCV (5) | 4 |
| 2010 | Improved Blur Insensitivity for Decorrelated Local Phase QuantizationabstractThis paper presents a novel blur tolerant decor relation scheme for local phase quantization (LPQ) texture descriptor. As opposed to previous methods, the introduced model can be applied with virtually any kind of blur regardless of the point spread function. The new technique takes also into account the changes in the image characteristics originating from the blur itself. The implementation does not suffer from multiple solutions like the decor relation in original LPQ, but still retains the same run-time computational complexity. The texture classification experiments illustrate considerable improvements in the performance of LPQ descriptors in the case of blurred images and show only negligible loss of accuracy with sharp images. Janne Heikkilä, Ville Ojansivu, Esa Rahtu |
ICPR | 1 |
| 2010 | A Human Detection Framework for Heavy MachineryabstractA stereo camera based human detection framework for heavy machinery is proposed. The framework allows easy integration of different human detection and image segmentation methods. This integration is essential for diverge and challenging work machine environments, in which traditional, one detector based human detection approaches has been found to be insufficient. The framework is based on the idea of pixel-wise human probabilities, which are obtained by several separate detection trials following binomial distribution. The framework has been evaluated with extensive image sequences of authentic work machine environments, and it has proven to be feasible. Promising detection performance was achieved by utilizing publically available human detectors. Teuvo Antero Heimonen, Janne Heikkilä |
ICPR | 2 |
| 2010 | Compressing Sparse Feature Vectors Using Random Ortho-ProjectionsabstractIn this paper we investigate the usage of random ortho-projections in the compression of sparse feature vectors. The study is carried out by evaluating the compressed features in classification tasks instead of concentrating on reconstruction accuracy. In the random ortho-projection method, the mapping for the compression can be obtained without any further knowledge of the original features. This makes the approach favorable if training data is costly or impossible to obtain. The independence from the data also enables one to embed the compression scheme directly into the computation of the original features. Our study is inspired by the results in compressive sensing, which state that up to a certain compression ratio and with high probability, such projections result in no loss of information. In comparison to learning based compression, namely principal component analysis (PCA), the random projections resulted in comparable performance already at high compression ratios depending on the sparsity of the original features. Esa Rahtu, Mikko Salo, Janne Heikkilä |
ICPR | 3 |
| 2008 | On Bin Configuration of Shape Context Descriptors in Human Silhouette Classification
Mark Barnard, Janne Heikkilä |
ACIVS | 2 |
| 2008 | Blur and Contrast Invariant Fast Stereo Matching
Matteo Pedone, Janne Heikkilä |
ACIVS | 2 |
| 2008 | Object recognition and segmentation by non-rigid quasi-dense matchingabstractIn this paper, we present a non-rigid quasi-dense matching method and its application to object recognition and segmentation. The matching method is based on the match propagation algorithm which is here extended by using local image gradients for adapting the propagation to smooth non-rigid deformations of the imaged surfaces. The adaptation is based entirely on the local properties of the images and the method can be hence used in non-rigid image registration where global geometric constraints are not available. Our approach for object recognition and segmentation is directly built on the quasi-dense matching. The quasi-dense pixel matches between the model and test images are grouped into geometrically consistent groups using a method which utilizes the local affine transformation estimates obtained during the propagation. The number and quality of geometrically consistent matches is used as a recognition criterion and the location of the matching pixels directly provides the segmentation. The experiments demonstrate that our approach is able to deal with extensive background clutter, partial occlusion, large scale and viewpoint changes, and notable geometric deformations. Juho Kannala, Esa Rahtu, Sami S. Brandt, Janne Heikkilä |
CVPR | 4 |
| 2008 | Multi-object tracking using binary masksabstractIn this paper, we introduce a new method for tracking multiple objects. The method combines Kalman filtering and the Expectation Maximization (EM) algorithm in a novel way to deal with observations that obey a Gaussian mixture model instead of a unimodal distribution that is assumed by the ordinary Kalman filter. It also involves a new approach to measuring the object locations using a series of morphological operations with binary masks. The benefit of this approach is that soft assignment of the measurements to corresponding objects can be performed automatically using their a posteriori probabilities. This is a general approach for multi-object tracking, and there are basically various ways to segment the objects, but in this paper we use simple color features simply to demonstrate the feasibility of the concept. Sami Huttunen, Janne Heikkilä |
ICIP | 2 |
| 2008 | Face Tracking for Spatially Aware Mobile User Interfaces
Jari Hannuksela, Pekka Sangi, Markus Turtinen, Janne Heikkilä |
ICISP | 4 |
| 2008 | Blur Insensitive Texture Classification Using Local Phase Quantization
Ville Ojansivu, Janne Heikkilä |
ICISP | 2 |
| 2008 | Body part segmentation of noisy human silhouette imagesabstractIn this paper we propose a solution to the problem of body part segmentation in noisy silhouette images. In developing this solution we revisit the issue of insufficient labeled training data, by investigating how synthetically generated data can be used to train general statistical models for shape classification. In our proposed solution we produce sequences of synthetically generated images, using three dimensional rendering and motion capture information. Each image in these sequences is labeled automatically as it is generated and this labeling is based on the hand labeling of a single initial image.We use shape context features and Hidden Markov Models trained based on this labeled synthetic data. This model is then used to segment silhouettes into four body parts; arms, legs, body and head. Importantly, in all the experiments we conducted the same model is employed with no modification of any parameters after initial training. Mark Barnard, Matti Matilainen, Janne Heikkilä |
ICME | 3 |
| 2008 | Recognition of blurred faces using Local Phase QuantizationabstractIn this paper, recognition of blurred faces using the recently introduced Local Phase Quantization (LPQ) operator is proposed. LPQ is based on quantizing the Fourier transform phase in local neighborhoods. The phase can be shown to be a blur invariant property under certain commonly fulfilled conditions. In face image analysis, histograms of LPQ labels computed within local regions are used as a face descriptor similarly to the widely used Local Binary Pattern (LBP) methodology for face image description. The experimental results on CMU PIE and FRGC 1.0.4 datasets show that the LPQ descriptor is highly tolerant to blur but still very descriptive outperforming LBP both with blurred and sharp images. Timo Ahonen, Esa Rahtu, Ville Ojansivu, Janne Heikkilä |
ICPR | 4 |
| 2008 | Rotation invariant local phase quantization for blur insensitive texture analysisabstractThis paper introduces a rotation invariant extension to the blur insensitive local phase quantization texture descriptor. The new method consists of two stages, the first of which estimates the local characteristic orientation, and the second one extracts a binary descriptor vector. Both steps of the algorithm apply the phase of the locally computed Fourier transform coefficients, which can be shown to be insensitive to centrally symmetric image blurring. The new descriptors are assessed in comparison with the well known texture descriptors, local binary patterns (LBP) and Gabor filtering. The results illustrate that the proposed method has superior performance in those cases where the image contains blur and is slightly better even with sharp images. Ville Ojansivu, Esa Rahtu, Janne Heikkilä |
ICPR | 3 |
| 2008 | Adaptive Motion-Based Gesture Recognition Interface for Mobile Phones
Jari Hannuksela, Mark Barnard, Pekka Sangi, Janne Heikkilä |
ICVS | 4 |
| 2008 | Measuring and modelling sewer pipes from video
Juho Kannala, Sami S. Brandt, Janne Heikkilä |
Mach. Vis. Appl. | 3 |
| 2007 | A New Rotation Search for Dependent Rate-Distortion Optimization in Video CodingabstractWe present a novel search algorithm which is suitable for optimizing functions with a high-dimensional discrete-valued parameter vector. The algorithm is designed to find a function local optimum with the minimal number of evaluated points without requiring function derivatives. The algorithm is applied to frame-level rate-distortion (R-D) optimization using Lagrangian relaxation to the rate constraints and to block motion estimation in H.264-based video coding. The R-D optimization is further accelerated by finding a good starting point by the golden section search. The results show excellent near-optimal R-D performance while computation is reduced by 99% compared to the quadratic coordinate-wise steepest descent algorithm. In motion estimation, the new algorithm requires 7-13% less checking points than the small diamond search algorithm with only a small penalty in prediction quality. Tuukka Toivonen, Loren Merritt, Ville Ojansivu, Janne Heikkilä |
ICASSP (1) | 4 |
| 2007 | Vision-based motion estimation for interaction with mobile devices
Jari Hannuksela, Pekka Sangi, Janne Heikkilä |
Comput. Vis. Image Underst. | 3 |
| 2007 | Image Registration Using Blur-Invariant Phase CorrelationabstractIn this paper, we propose an image registration method, which is invariant to centrally symmetric blur. The method utilizes the phase of the images and has its roots on phase correlation (PC) registration. We show how the even powers of the normalized Fourier transform of an image are invariant to centrally symmetric blur, such as motion or out-of-focus blur. We then use these results to propose blur-invariant phase correlation. The method has been compared to PC registration with excellent results. With a subpixel extension of PC registration, the method achieves subpixel accuracy for even heavily blurred images. Ville Ojansivu, Janne Heikkilä |
IEEE Signal Process. Lett. | 2 |
| 2006 | Motion Blur Concealment of Digital Video Using Invariant Features
Ville Ojansivu, Janne Heikkilä |
ACIVS | 2 |
| 2006 | An Active Head Tracking System for Distance Education and Videoconferencing ApplicationsabstractWe present a system for automatic head tracking with a single pan-tilt-zoom (PTZ) camera. In distance education the PTZ tracking system developed can be used to follow a teacher actively when s/he moves in the classroom. In other videoconferencing applications the system can be utilized to provide a close-up view of the person all the time. Since the color features used in tracking are selected and updated online, the system can adapt to changes rapidly. The information received from the tracking module is used to actively control the PTZ camera in order to keep the person in the camera view. In addition, the system implemented is able to recover from erroneous situations. Preliminary experiments indicate that the PTZ system can perform well under different lighting conditions and large scale changes. Sami Huttunen, Janne Heikkilä |
AVSS | 2 |
| 2006 | Algorithms for Computing a Planar Homography from Conics in CorrespondenceabstractThis paper presents two new algorithms for computing a planar homography from conic correspondences. Firstly, we propose a linear algorithm for computing the homography when there are three or more conic correspondences. In this case, we get an overdetermined set of linear equations and the solution that minimizes the algebraic distance is obtained by the singular value decomposition. Secondly, we propose another algorithm for determining the homography from only two conic correspondences. Unlike the previous algorithms our approach uses only linear algebra and does not require solving high-degree polynomial equations. Hence, the proposed formulation leads to an algorithm that is efficient and easy to implement. In addition, our approach incorporates the computation of the two projective invariants for a pair of conics. These invariants provide a condition for the existence of a homography between the pairs of conics. We evaluate the characteristics and robustness of the proposed algorithms in experiments with synthetic and real data. 1 Juho Kannala, Mikko Salo, Janne Heikkilä |
BMVC | 3 |
| 2006 | A New Affine Invariant Image Transform Based on RidgeletsabstractIn this paper we present a new affine invariant image transform, based on ridgelets. The proposed transform is directly applicable to segmented image patches. The new method has some similarities with the previously proposed Multiscale Autoconvolution, but it will offer a more general framework and possibilities for variations. The obtained transform coefficients can be used in affine invariant pattern classification, and as shown in the experiments, already a small subset of them is enough for reliable recognition of complex patterns. The new method is assessed in several experiments and it is observed to perform well under many nonaffine distortions. 1 Esa Rahtu, Janne Heikkilä, Mikko Salo |
BMVC | 2 |
| 2006 | Multiscale Autoconvolution Histograms for Affine Invariant Pattern RecognitionabstractIn this paper we present a new way of producing affine invariant histograms from images. The approach is based on a probabilistic interpretation of the image function as in the multiscale autoconvolution (MSA) transform, but the histograms extract much more information of the image than traditional MSA. The new histograms can be considered as generalizations of the image gray scale histogram, encoding also the spatial information. It turns out that the proposed method can be efficiently computed using the Fast Fourier Transform, and it will be shown to have essentially the same computational load as MSA. The experiments performed indicate that the new invariants are capable of reliable classification of complex patterns, outperforming MSA and many other methods. 1 Esa Rahtu, Mikko Salo, Janne Heikkilä |
BMVC | 3 |
| 2006 | Improved Unsymmetric-Cross Multi-Hexagon-Grid Search Algorithm for Fast Block Motion EstimationabstractWe develop a set of new motion estimation (ME) algorithms based mainly on unsymmetric-cross multi-hexagon-grid search (UMH). The original algorithms are improved by applying the successive elimination algorithm (SEA) and subsampling the image blocks while computing the matching criterion. We also improve SEA by adding a small constant to the lower bound before trying to eliminate the current checking point. The motion compensated results stay in most cases similar to the original UMH algorithm, while computation is decreased by up to 95%. The new algorithms outperform in both image quality and computational efficiency other well-known fast ME algorithms such as three step search (TSS), diamond search (DS), and hexagon-based search (HEXBS). Tuukka Toivonen, Janne Heikkilä |
ICIP | 2 |
| 2006 | A New Convexity Measure Based on a Probabilistic Interpretation of ImagesabstractIn this paper, we present a novel convexity measure for object shape analysis. The proposed method is based on the idea of generating pairs of points from a set and measuring the probability that a point dividing the corresponding line segments belongs to the same set. The measure is directly applicable to image functions representing shapes and also to gray-scale images which approximate image binarizations. The approach introduced gives rise to a variety of convexity measures which make it possible to obtain more information about the object shape. The proposed measure turns out to be easy to implement using the Fast Fourier Transform and we will consider this in detail. Finally, we illustrate the behavior of our measure in different situations and compare it to other similar ones. Esa Rahtu, Mikko Salo, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2006 | Video filtering with Fermat number theoretic transforms using residue number systemabstractWe investigate image and video convolutions based on Fermat number transform (FNT) modulo q=2/sup M/+1 where M is an integer power of two. These transforms are found to be ideal for image convolutions, except that the choices for the word length, restricted by the transform modulus, are rather limited. We discuss two methods to overcome this limitation. First, we allow M to be an arbitrary integer. This gives much wider variety in possible moduli, at the cost of decreased transform length of 16 or 32 points for M<32. Nevertheless, the transform length appears still to be useful especially with block-based image and video filtering applications. We call these transforms the generalized FNT (GFNT). The second solution is to use a residue number system (RNS) to enlarge the effective modulus, while performing actual number theoretic transforms with smaller moduli. This approach appears to be particularly useful with moduli q/sub 1/=2/sup 16/+1 and q/sub 2/=2/sup 8/+1, which allow transforms up to 256 points with a dynamic range of about 24 bits. We design an efficient reconstruction circuit based on mixed radix conversion for converting the result from diminished-1 RNS into normal binary code. The circuit is implemented in VHDL and found to be very small in area. We also discuss the necessary steps in performing convolutions with the GFNT and evaluate the integrated circuit implementation cost for various elementary operations. Tuukka Toivonen, Janne Heikkilä |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | Affine registration with multi-scale autoconvolutionabstractIn this paper we propose a novel method for the recovery of affine transformation parameters between two images. Registration is achieved without separate feature extraction by directly utilizing the intensity distribution of the images. The method can also be used for matching point sets under affine transformations. Our approach is based on the same probabilistic interpretation of the image function as the recently introduced multi-scale autoconvolution (MSA) transform. Here we describe how the framework may be used in image registration and present two variants of the method for practical implementation. The proposed method is experimented with binary and grayscale images and compared with other non-feature-based registration methods. The experiments show that the new method can efficiently align images of isolated objects and is relatively robust. Juho Kannala, Esa Rahtu, Janne Heikkilä |
ICIP (3) | 3 |
| 2005 | A likelihood function for block-based motion analysisabstractIn this paper, the computation of likelihood of block motion candidates is considered. The method is based on the evaluation of the sum of squared differences (SSD) measure for local displacements and probabilistic interpretation of these values using local gradient information. Simulated motion data is used to estimate parameters of conditional SSD distributions. The application of our novel likelihood function is demonstrated in a task of dominant motion estimation, where particle filtering is used to maintain a set of global motion hypotheses. In this task, the block motion likelihood function is used as a basis for hypothesis testing, which provides a means for evaluating global motion hypotheses. Pekka Sangi, Janne Heikkilä, Olli Silvén |
ICIP (1) | 2 |
| 2005 | Affine Invariant Pattern Recognition Using Multiscale AutoconvolutionabstractThis paper presents a new affine invariant image transform called Multiscale Autoconvolution (MSA). The proposed transform is based on a probabilistic interpretation of the image function. The method is directly applicable to isolated objects and does not require extraction of boundaries or interest points, and the computational load is significantly reduced using the Fast Fourier Transform. The transform values can be used as descriptors for affine invariant pattern classification and, in this article, we illustrate their performance in various object classification tasks. As shown by a comparison with other affine invariant techniques, the new method appears to be suitable for problems where image distortions can be approximated with affine transformations. Esa Rahtu, Mikko Salo, Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | A Texture-based Method for Detecting Moving ObjectsabstractThe detection of moving objects from video frames plays an important and often very critical role in different kinds of machine vision applications including human detection and tracking, traffic monitoring, humanmachine interfaces and military applications, since it usually is one of the first phases in a system architecture. A common way to detect moving objects is background subtraction. In background subtraction, moving objects are detected by comparing each video frame against an existing model of the scene background. In this paper, we propose a novel block-based algorithm for background subtraction. The algorithm is based on the Local Binary Pattern (LBP) texture measure. Each image block is modelled as a group of weighted adaptive LBP histograms. The algorithm operates in real-time under the assumption of a stationary camera with fixed focal length. It can adapt to inherent changes in scene background and can also handle multimodal backgrounds. Marko Heikkilä, Matti Pietikäinen, Janne Heikkilä |
BMVC | 3 |
| 2004 | Selection of the Lagrange multiplier for block-based motion estimation criteriaabstractIn hybrid video coding, motion vectors used for motion compensation constitute an important set of decisions. Cost functions for block motion estimation that take the smoothness of the resulting motion vector field into account, in addition to the motion compensated prediction error, have been proposed. Computationally simple derivatives of sum of absolute differences and sum of squared differences-based criteria are studied in this paper. Cost functions are based on Lagrangian rate-distortion formulation, and the basic question is how the Lagrangian multiplier involved should be selected. Assumptions behind these cost functions are discussed, and a new method is derived for determining the multiplier. Comparisons with other strategies are made with experiments. The results show that the selection of the multiplier is not critical. Pekka Sangi, Janne Heikkilä, Olli Silvén |
ICASSP (3) | 2 |
| 2004 | Image scale and rotation from the phase-only bispectrum
Janne Heikkilä |
ICIP | 1 |
| 2004 | Fast full search block motion estimation for H.264/AVC with multilevel successive elimination algorithm
Tuukka Toivonen, Janne Heikkilä |
ICIP | 2 |
| 2004 | A real-time system for monitoring of cyclists and pedestrians
Janne Heikkilä, Olli Silvén |
Image Vis. Comput. | 1 |
| 2004 | Pattern matching with affine moment descriptors
Janne Heikkilä |
Pattern Recognit. | 1 |
| 2004 | A new class of shift-invariant operatorsabstractThis letter proposes a class of operators with a shift invariance property. These operators are derived from two-dimensional (2-D) complex moment invariants based on the observation that there is a duality between rotation invariance and shift invariance. A general form of the shift invariants belonging to this class is presented, which shows that polyspectral invariants such as the power spectrum and the bispectrum are members of the class. Methods for computing shift invariants for one-dimensional (1-D) and 2-D signals are also presented. The examples given in the paper suggest that the higher order operators can preserve the original signal waveform better than autocorrelation. Janne Heikkilä |
IEEE Signal Process. Lett. | 1 |
| 2003 | A new rate-minimizing matching criterion and a fast algorithm for block motion estimationabstractA new block matching criterion for motion estimation in video coding that will give better encoded video quality than the commonly used sum of absolute differences (SAD) or even sum of squared differences (SSD) criteria is presented. The new criterion tends to concentrate the discrete cosine transformed block energy into DC frequency which may allow coding the AC coefficients with less bits. The criterion gives best results on sequences which have varying lighting conditions. Furthermore, we modify the successive elimination algorithm (SEA) and multilevel successive elimination algorithm (MSEA) to be usable with the new criterion by deriving a new tighter lower bound for the SSD criterion. The new bound can be used either directly with the SSD criterion or with the new bit-rate minimizing criterion. Tuukka Toivonen, Janne Heikkilä |
ICIP (2) | 2 |
| 2000 | Camera Motion Estimation from Non-Stationary Scenes Using EM-Based Motion SegmentationabstractAn algorithm for recovering 3-D camera motion from sequences of images is proposed. The algorithm has four stages. In the first stage, the motion vector field is segmented using an EM-based method. The resulting segments are compared and the coherent regions are merged in the second stage. The candidates for the background regions are determined and finally used for 3-D motion estimation in the last two stages. Unlike most of the other methods, this approach tolerates also non-rigid motion in the scene. The experiments performed show that in some cases more information or reasoning is needed for selecting plausible motion parameters from several hypotheses. Janne Heikkilä, Pekka Sangi, Olli Silvén |
ICPR | 1 |
| 2000 | Intensity Independent Color Models and Visual TrackingabstractSome intensity independent color models are studied experimentally in the scope of visual tracking to introduce robustness to illumination changes. Also, a plain color background model to allow modest camera motion is presented. The background is represented as a Gaussian mixture in color space. The EM algorithm is applied to find the decomposition and minimum description length principle is proposed to determine the number of mixture components. A simple tracking algorithm is outlined and the achieved results are introduced. Mika Korhonen, Janne Heikkilä, Olli Silvén |
ICPR | 2 |
| 2000 | Geometric Camera Calibration Using Circular Control PointsabstractModern CCD cameras are usually capable of a spatial accuracy greater than 1/50 of the pixel size. However, such accuracy is not easily attained due to various error sources that can affect the image formation process. Current calibration methods typically assume that the observations are unbiased, the only error is the zero-mean independent and identically distributed random noise in the observed image coordinates, and the camera model completely explains the mapping between the 3D coordinates and the image coordinates. In general, these conditions are not met, causing the calibration results to be less accurate than expected. In the paper, a calibration procedure for precise 3D computer vision applications is described. It introduces bias correction for circular control points and a nonrecursive method for reversing the distortion model. The accuracy analysis is presented and the error sources that can reduce the theoretical accuracy are discussed. The tests with synthetic images indicate improvements in the calibration results in limited error conditions. In real images, the suppression of external error sources becomes a prerequisite for successful calibration. Janne Heikkilä |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1998 | Moment and curvature preserving technique for accurate ellipse boundary detectionabstractCircles and their elliptic projections are very commonly used image features in computer vision applications. Thus, it is very important to be able to determine their location in an accurate manner. A technique for determining an ellipse boundary with subpixel precision is proposed. The technique, called the moment and curvature preserving detection (MCP), utilizes the first three intensity moments and the intensity gradient of the image. The idea of using moments for subpixel edge detection is not new, but in the case of ellipses the moments do not provide sufficient information for precise detection. However, if the local curvature is augmented to the observations, the ellipse boundary can be determined reliably. Janne Heikkilä |
ICPR | 1 |
| 1998 | Linear motion estimation for image sequence based accurate 3-D measurementsabstractWe present a method for making accurate 3-D measurements from monocular image sequences. The process of determining camera motion is completely separated from 3-D structure estimation. The algorithm has two steps: elimination of rotations and estimation of the camera translation. Elimination of rotations is based on pre-calibration, and estimation of the camera translation is based on locating the focus of expansion from image disparities. The method proposed utilizes the total least squares estimation technique. By using the motion data, the 3-D coordinates of the measurement points can be solved linearly up to a scale factor. Due to the nonrecursive nature of the method, it provides a fast approach for processing long image sequences in an accurate manner. Janne Heikkilä, Olli Silvén |
ICPR | 1 |
| 1997 | A Four-step Camera Calibration Procedure with Implicit Image CorrectionabstractIn geometrical camera calibration the objective is to determine a set of camera parameters that describe the mapping between 3-D reference coordinates and 2-D image coordinates. Various methods for camera calibration can be found from the literature. However surprisingly little attention has been paid to the whole calibration procedure, i.e., control point extraction from images, model fitting, image correction, and errors originating in these stages. The main interest has been in model fitting, although the other stages are also important. In this paper we present a four-step calibration procedure that is an extension to the two-step method. There is an additional step to compensate for distortion caused by circular features, and a step for correcting the distorted image coordinates. The image correction is performed with an empirical inverse model that accurately compensates for radial and tangential distortions. Finally, a linear method for solving the parameters of the inverse model is presented. Janne Heikkilä, Olli Silvén |
CVPR | 1 |
| 1996 | Calibration procedure for short focal length off-the-shelf CCD camerasabstractA camera calibration procedure intended for a 3D measurement application is presented, paying attention to the various error sources. The error may be measurement noise that is random by nature, but it may also be systematic originating from the calibration target used, geometrical distortions and illumination. In order to obtain good calibration results, the systematic error sources should be eliminated or their effects compensated for. Then, the camera parameters can be determined by fitting the corrected measurements to the camera model which in our case is a combination of a pinhole camera and lens distortion models. We also notice that a more complete camera model is needed to explain all the error components. Janne Heikkilä, Olli Silvén |
ICPR | 1 |
| 1996 | accurate 3-D Measurement Using a Single Video CameraabstractWe present a straightforward technique for determining the 3-D locations of feature points using sequences of monocular image frames captured by a moving camera. The motion of the camera is estimated simultaneously. In practice, only the camera needs careful calibration. Based on experiments, the repeatability is currently about 1/3500 and accuracy 1/2500. This approach has potential for high speed, as hundreds of points may be measured from the same image sequence. Janne Heikkilä, Olli Silvén |
Int. J. Pattern Recognit. Artif. Intell. | 1 |