Joni-Kristian Kämäräinen

dblp:k/JoniKristianKamarainen · also Joni-Kristian Kamarainen · DBLP profile ↗
← Back
78ranked-venue papers
5as first author
15since 2021 · last 2024
0000-0002-5801-4371ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 64 · 4 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 42 · 3 first-author · 8 since 2021Systems, architecture and hardware · 11 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5Human-computer interaction and ubiquitous computing · 3
YearPublicationVenuePosition
2024 Probabilistic Subgoal Representations for Hierarchical Reinforcement Learning
abstract
In goal-conditioned hierarchical reinforcement learning (HRL), a high-level policy specifies a subgoal for the low-level policy to reach. Effective HRL hinges on a suitable subgoal representation function, abstracting state space into latent subgoal space and inducing varied low-level behaviors. Existing methods adopt a subgoal representation that provides a deterministic mapping from state space to latent subgoal space. Instead, this paper utilizes Gaussian Processes (GPs) for the first probabilistic subgoal representation. Our method employs a GP prior on the latent subgoal space to learn a posterior distribution over the subgoal representation functions while exploiting the long-range correlation in the state space through learnable kernels. This enables an adaptive memory that integrates long-range subgoal information from prior planning steps allowing to cope with stochastic uncertainties. Furthermore, we propose a novel learning objective to facilitate the simultaneous learning of probabilistic subgoal representations and policies within a unified framework. In experiments, our approach outperforms state-of-the-art baselines in standard benchmarks but also in environments with stochastic elements and under diverse reward conditions. Additionally, our model shows promising capabilities in transferring low-level policies across different tasks.
Vivienne Huiling Wang, Tinghuai Wang, Wenyan Yang, Joni-Kristian Kämäräinen, Joni Pajarinen
ICML4
2024 PlaceNav: Topological Navigation through Place Recognition
abstract
Recent results suggest that splitting topological navigation into robot-independent and robot-specific components improves navigation performance by enabling the robot-independent part to be trained with data collected by robots of different types. However, the navigation methods’ performance is still limited by the scarcity of suitable training data and they suffer from poor computational scaling. In this work, we present PlaceNav, subdividing the robot-independent part into navigation-specific and generic computer vision components. We utilize visual place recognition for the subgoal selection of the topological navigation pipeline. This makes subgoal selection more efficient and enables leveraging large-scale datasets from non-robotics sources, increasing training data availability. Bayesian filtering, enabled by place recognition, further improves navigation performance by increasing the temporal consistency of subgoals. Our experimental results verify the design and the new method obtains a 76 % higher success rate in indoor and 23 % higher in outdoor navigation tasks with higher computational efficiency.
Lauri Suomela, Jussi Kalliola, Harry Edelman, Joni-Kristian Kämäräinen
ICRA4
2024 Single Pixel Spectral Color Constancy
abstract
Abstract Color constancy is still one of the biggest challenges in camera color processing. Convolutional neural networks have been able to improve the situation but there are still problems in many conditions, especially in scenes where a single color is dominating. In this work, we approach the problem from a slightly different setting. What if we could have some other information than the raw RGB image data. What kind of information would help to bring significant improvements while still be feasible in a mobile device. These questions sparked an idea for a novel approach for computational color constancy. Instead of raw RGB images used by the existing algorithms to estimate the scene white points, our approach is based on the scene’s average color spectra-single pixel spectral measurement. We show that as few as 10–14 spectral channels are sufficient. Notably, the sensor output has five orders of magnitude less data than in raw RGB images of a 10MPix camera. The spectral sensor captures the “spectral fingerprints” of different light sources and the illuminant white point can be accurately estimated by a standard regressor. The regressor can be trained with generated measurements using the existing RGB color constancy datasets. For this purpose, we propose a spectral data generation pipeline that can be used if the dataset camera model is known and thus its spectral characterization can be obtained. To verify the results with real data, we collected a real spectral dataset with a commercial spectrometer. On all datasets the proposed Single Pixel Spectral Color Constancy obtains the highest accuracy in the both single and cross-dataset experiments. The method is particularly effective for the difficult scenes for which the average improvements are 40–70% compared to state-of-the-arts. The approach can be extended to multi-illuminant case for which the experimental results also provide promising results.
Samu Koskinen, Erman Acar, Joni-Kristian Kämäräinen
Int. J. Comput. Vis.3
2023 State-Conditioned Adversarial Subgoal Generation
abstract
Hierarchical reinforcement learning (HRL) proposes to solve difficult tasks by performing decision-making and control at successively higher levels of temporal abstraction. However, off-policy HRL often suffers from the problem of a non-stationary high-level policy since the low-level policy is constantly changing. In this paper, we propose a novel HRL approach for mitigating the non-stationarity by adversarially enforcing the high-level policy to generate subgoals compatible with the current instantiation of the low-level policy. In practice, the adversarial learning is implemented by training a simple state conditioned discriminator network concurrently with the high-level policy which determines the compatibility level of subgoals. Comparison to state-of-the-art algorithms shows that our approach improves both learning efficiency and performance in challenging continuous control tasks.
Vivienne Huiling Wang, Joni Pajarinen, Tinghuai Wang, Joni-Kristian Kämäräinen
AAAI4
2023 Seq2Seq Imitation Learning for Tactile Feedback-based Manipulation
abstract
Robot control for tactile feedback based manip-ulation can be difficult due to modeling of physical contacts, partial observability of the environment, and noise in perception and control. This work focuses on solving partial observability of contact-rich manipulation tasks as a Sequence-to-Sequence (Seq2Seq) Imitation Learning (IL) problem. The proposed Seq2Seq model first produces a robot-environment interaction sequence to estimate the partially observable environment state variables, and then, the observed interaction sequence is transformed to a control sequence for the task itself. The proposed Seq2Seq IL for tactile feedback based manipulation is experimentally validated on a door-open task in a simulated environment and a snap-on insertion task with a real robot. The model is able to learn both tasks from only 50 expert demonstrations while state-of-the-art reinforcement learning and imitation learning methods fail.
Wenyan Yang, Alexandre Angleraud, Roel Pieters, Joni Pajarinen, Joni-Kristian Kämäräinen
ICRA5
2023 Benchmarking Visual Localization for Autonomous Navigation
abstract
This work introduces a simulator-based benchmark for visual localization in the autonomous navigation context. The dynamic benchmark enables investigation of how variables such as the time of day, weather, and camera perspective affect the navigation performance of autonomous agents that utilize visual localization for closed-loop control. The experimental part of the paper studies the effects of four such variables by evaluating state-of-the-art visual localization methods as part of the motion planning module of an autonomous navigation stack. The results show major variation in the suitability of the different methods for vision-based navigation. To the authors' best knowledge, the proposed benchmark is the first to study modern visual localization methods as part of a complete navigation stack. We make the benchmark available at https://github.com/lasuomela/carla_vloc_benchmark.
Lauri Suomela, Jussi Kalliola, Atakan Dag, Harry Edelman, Joni-Kristian Kämäräinen
WACV5
2022 Long-term Visual Place Recognition
abstract
In this work, we study the long-term performance of visual place recognition in urban outdoor environment. A long-term benchmark is constructed from the Oxford RobotCar dataset. It contains sequences of the same route traversed over a period of approx. 500 days. We carefully selected three gallery sequences, one training sequence and 15 query sequences that cover different seasons, times of day and weather. The RobotCar sequences from the first half year have several problems, for example, only partial routes and inaccurate location data. We circumvent these problems by reversing the time. In the benchmark dataset the gallery and training images are the latest and the query sequences go gradually back in time. Our experiments provide the following findings. 1) the selected gallery sequence has strong impact on performance, and 2) additional training sequences help to mitigate differences between the gallery sequences. In addition, results indicate that 3) there is a long-term trend of performance degradation over time. The degradation can be quantified as about 6 percentage points per 100 days and, therefore, the initial performance of 40% eventually drops below 20% at the end.
Farid Alijani, Jukka Peltomäki, Jussi Puura, Heikki Huttunen, Joni-Kristian Kämäräinen, Esa Rahtu
ICPR5
2022 Active Short-Long Exposure Deblurring
abstract
Mobile phones can capture image bursts to produce high quality still photographs. The simplest form of a burst is two frame short-long (S-L) exposure. S-L exposure is particularly suitable in low light conditions where short exposure frames are sharp but noisy and dark, and long exposure frames are affected by motion blur but have better scene chromaticity and luminance. In this work, we take a step further and define active short-long exposure deblurring where the viewfinder frames before the burst are used to optimize the S-L exposure parameters. We introduce deep architectures and data generation for active S-L exposure deblurring. The approach is experimentally validated with realistic data and it shows clear improvements. For the most difficult scenes (worst 5%) the PSNR is improved by +1.39dB.
Dan Yang 0012, Samu Koskinen, Joni-Kristian Kämäräinen
ICPR3
2021 Single Pixel Spectral Color Constancy
Samu Koskinen, Erman Acar, Joni-Kristian Kämäräinen
BMVC3
2021 Depth-only Object Tracking
Ales Leonardis, Joni-Kristian Kämäräinen
BMVC4
2021 DepthTrack: Unveiling the Power of RGBD Tracking
abstract
RGBD (RGB plus depth) object tracking is gaining momentum as RGBD sensors have become popular in many application fields such as robotics. However, the best RGBD trackers are extensions of the state-of-the-art deep RGB trackers. They are trained with RGB data and the depth channel is used as a sidekick for subtleties such as occlusion detection. This can be explained by the fact that there are no sufficiently large RGBD datasets to 1) train "deep depth trackers" and to 2) challenge RGB trackers with sequences for which the depth cue is essential. This work introduces a new RGBD tracking dataset - Depth-Track - that has twice as many sequences (200) and scene types (40) than in the largest existing dataset, and three times more objects (90). In addition, the average length of the sequences (1473), the number of deformable objects (16) and the number of annotated tracking attributes (15) have been increased. Furthermore, by running the SotA RGB and RGBD trackers on DepthTrack, we propose a new RGBD tracking baseline, namely DeT, which reveals that deep RGBD tracking indeed benefits from genuine training data. The code and dataset is available at https://github.com/xiaozai/DeT.
Jani Käpylä, Feng Zheng 0001, Ales Leonardis, Joni-Kristian Kämäräinen
ICCV6
2021 Neural Network Controller for Autonomous Pile Loading Revised
abstract
We have recently proposed two pile loading controllers that learn from human demonstrations: a neural network (NNet) [1] and a random forest (RF) controller [2]. In the field experiments the RF controller obtained clearly better success rates. In this work, the previous findings are drastically revised by experimenting summer time trained controllers in winter conditions. The winter experiments revealed a need for additional sensors, more training data, and a controller that can take advantage of these. Therefore, we propose a revised neural controller (NNetV2) which has a more expressive structure and uses a neural attention mechanism to focus on important parts of the sensor and control signals. Using the same data and sensors to train and test the three controllers, NNetV2 achieves better robustness against drastically changing conditions and superior success rate. To the best of our knowledge, this is the first work testing a learning-based controller for a heavy-duty machine in drastically varying outdoor conditions and delivering high success rate in winter, being trained in summer.
Wenyan Yang, Nataliya Strokina, Nikolay Serbenyuk, Joni Pajarinen, Reza Ghabcheloo, Juho Vihonen, Mohammad M. Aref, Joni-Kristian Kämäräinen
ICRA8
2021 Monolithic vs. hybrid controller for multi-objective Sim-to-Real learning
abstract
Simulation to real (Sim-to-Real) is an attractive approach to construct controllers for robotic tasks that are easier to simulate than to analytically solve. Working Sim-to-Real solutions have been demonstrated for tasks with a clear single objective such as "reach the target". Real world applications, however, often consist of multiple simultaneous objectives such as "reach the target" but "avoid obstacles". A straightforward solution in the context of reinforcement learning (RL) is to combine multiple objectives into a multi-term reward function and train a single monolithic controller. Recently, a hybrid solution based on pre-trained single objective controllers and a switching rule between them was proposed. In this work, we compare these two approaches in the multi-objective setting of a robot manipulator to reach a target while avoiding an obstacle. Our findings show that the training of a hybrid controller is easier and obtains a better success-failure trade-off than a monolithic controller. The controllers trained in simulator were verified by a real set-up.
Atakan Dag, Alexandre Angleraud, Wenyan Yang, Nataliya Strokina, Roel Pieters, Minna Lanz, Joni-Kristian Kämäräinen
IROS7
2021 Evaluation of Long-term LiDAR Place Recognition
abstract
We compare a state-of-the-art deep image retrieval and a deep place recognition method for place recognition using LiDAR data. Place recognition aims to detect previously visited locations and thus provides an important tool for navigation, mapping, and localisation. Experimental comparisons are conducted using challenging outdoor and indoor datasets, Oxford Radar RobotCar and COLD, in the "long-term" setting where the test conditions differ substantially from the training and gallery data. Based on our results the image retrieval methods using LiDAR depth images can achieve accurate localization (the single best match recall 80%) within 5.00 m in urban outdoors. In office indoors the comparable accuracy is 50 cm but is more sensitive to changes in the environment.
Jukka Peltomäki, Farid Alijani, Jussi Puura, Heikki Huttunen, Esa Rahtu, Joni-Kristian Kämäräinen
IROS6
2021 Fast Fourier Intrinsic Network
abstract
We address the problem of decomposing an image into albedo and shading. We propose the Fast Fourier Intrinsic Network, FFI-Net in short, that operates in the spectral domain, splitting the input into several spectral bands. Weights in FFI-Net are optimized in the spectral domain, allowing faster convergence to a lower error. FFI-Net is lightweight and does not need auxiliary networks for training. The network is trained end-to-end with a novel spectral loss which measures the global distance between the network prediction and corresponding ground truth. FFI-Net achieves state-of-the-art performance on MPI-Sintel, MIT Intrinsic, and IIW datasets.
Yanlin Qian, Miaojing Shi, Joni-Kristian Kämäräinen, Jiri Matas
WACV3
2020 Cross-dataset Color Constancy Revisited Using Sensor-to-Sensor Transfer
Samu Koskinen, Dan Yang 0012, Joni-Kristian Kämäräinen
BMVC3
2020 Loop-closure detection by LiDAR scan re-identification
abstract
In this work, loop-closure detection from LiDAR scans is defined as an image re-identification problem. Reidentification is performed by computing Euclidean distances of a query scan to a gallery set of previous scans. The distances are computed in a feature embedding space where the scans are mapped by a convolutional neural network (CNN). The network is trained using the triplet loss training strategy. In our experiments we compare different backbone networks, variants of the triplet loss and generic and LiDAR specific data augmentation techniques. With a realistic indoor dataset the best architecture obtains the mean average precision (mAP) above 0.94.
Jukka Peltomäki, Xingyang Ni, Jussi Puura, Joni-Kristian Kämäräinen, Heikki Huttunen
ICPR4
2020 DAL: A Deep Depth-Aware Long-term Tracker
abstract
The best RGBD trackers provide high accuracy but are slow to run. On the other hand, the best RGB trackers are fast but clearly inferior on the RGBD datasets. In this work, we propose a deep depth-aware long-term tracker that achieves state-of-the-art RGBD tracking performance and is fast to run. We reformulate deep discriminative correlation filter (DCF) to embed the depth information into deep features. Moreover, the same depth-aware correlation filter is used for target redetection. Comprehensive evaluations show that the proposed tracker achieves state-of-the-art performance on the Princeton RGBD, STC, and the newly-released CDTB benchmarks and runs 20 fps.
Yanlin Qian, Alan Lukezic, Matej Kristan, Joni-Kristian Kämäräinen, Jiri Matas
ICPR5
2020 Silhouette Body Measurement Benchmarks
abstract
Anthropometric body measurements are important for industrial design, garment fitting, medical diagnosis and ergonomics. A number of methods have been proposed to estimate the body measurements from images, but progress has been slow due to the lack of realistic and publicly available datasets. The existing works train and test on silhouettes of 3D body meshes obtained by fitting a human body model to the commercial CAESAR scans. In this work, we introduce the BODY-fit dataset that contains fitted meshes of 2,675 female and 1,474 male 3D body scans. We unify evaluation on the CAESAR-fit and BODY-fit datasets by computing body measurements from geodesic surface paths as the ground truth and by generating two-view silhouette images. We also introduce BODY-rgb - a realistic dataset of 86 male and 108 female subjects captured with an RGB camera and manually tape measured ground truth. We propose a simple yet effective deep CNN architecture as a baseline method which obtains competitive accuracy on the three datasets.
Johan Wirta, Joni-Kristian Kämäräinen
ICPR3
2020 Learning a Pile Loading Controller from Demonstrations
abstract
This work introduces a learning-based pile loading controller for autonomous robotic wheel loaders. Controller parameters are learnt from a small number of demonstrations for which low level sensor (boom angle, bucket angle and hydrostatic driving pressure), egocentric video frames and control signals are recorded. Application specific deep visual features are learnt from demonstrations using a Siamese network architecture and a combination of cross-entropy and contrastive loss. The controller is based on a Random Forest (RF) regressor that provides robustness against changes in field conditions (loading distance, soil type, weather and illumination). The controller is deployed to a real autonomous robotic wheel loader and it outperforms prior art with a clear margin.
Wenyan Yang, Nataliya Strokina, Nikolay Serbenyuk, Reza Ghabcheloo, Joni-Kristian Kämäräinen
ICRA5
2020 Anthropometric clothing measurements from 3D body scans
abstract
Abstract We propose a full processing pipeline to acquire anthropometric measurements from 3D measurements. The first stage of our pipeline is a commercial point cloud scanner. In the second stage, a pre-defined body model is fitted to the captured point cloud. We have generated one male and one female model from the SMPL library. The fitting process is based on non-rigid iterative closest point algorithm that minimizes overall energy of point distance and local stiffness energy terms. In the third stage, we measure multiple circumference paths on the fitted model surface and use a nonlinear regressor to provide the final estimates of anthropometric measurements. We scanned 194 male and 181 female subjects, and the proposed pipeline provides mean absolute errors from 2.5 to 16.0 mm depending on the anthropometric measurement.
Johan Wirta, Joni-Kristian Kämäräinen
Mach. Vis. Appl.3
2019 Object Tracking by Reconstruction With View-Specific Discriminative Correlation Filters
abstract
Standard RGB-D trackers treat the target as a 2D structure, which makes modelling appearance changes related even to out-of-plane rotation challenging. This limitation is addressed by the proposed long-term RGB-D tracker called OTR - Object Tracking by Reconstruction. OTR performs online 3D target reconstruction to facilitate robust learning of a set of view-specific discriminative correlation filters (DCFs). The 3D reconstruction supports two performance- enhancing features: (i) generation of an accurate spatial support for constrained DCF learning from its 2D projection and (ii) point-cloud based estimation of 3D pose change for selection and storage of view-specific DCFs which robustly localize the target after out-of-view rotation or heavy occlusion. Extensive evaluation on the Princeton RGB-D tracking and STC Benchmarks shows OTR outperforms the state-of-the-art by a large margin.
Ugur Kart, Alan Lukezic, Matej Kristan, Joni-Kristian Kämäräinen, Jiri Matas
CVPR4
2019 On Finding Gray Pixels
abstract
We propose a novel grayness index for finding gray pixels and demonstrate its effectiveness and efficiency in illumination estimation. The grayness index, GI in short, is derived using the Dichromatic Reflection Model and is learning-free. GI allows to estimate one or multiple illumination sources in color-biased images. On standard single-illumination and multiple-illumination estimation benchmarks, GI outperforms state-of-the-art statistical methods and many recent deep methods. GI is simple and fast, written in a few dozen lines of code, processing a 1080p image in ~0.4 seconds with a non-optimized Matlab code.
Yanlin Qian, Joni-Kristian Kämäräinen, Jarno Nikkanen, Jiri Matas
CVPR2
2019 CDTB: A Color and Depth Visual Object Tracking Dataset and Benchmark
abstract
We propose a new color-and-depth general visual object tracking benchmark (CDTB). CDTB is recorded by several passive and active RGB-D setups and contains indoor as well as outdoor sequences acquired in direct sunlight. The CDTB dataset is the largest and most diverse dataset in RGB-D tracking, with an order of magnitude larger number of frames than related datasets. The sequences have been carefully recorded to contain significant object pose change, clutter, occlusion, and periods of long-term target absence to enable tracker evaluation under realistic conditions. Sequences are per-frame annotated with 13 visual attributes for detailed analysis. Experiments with RGB and RGB-D trackers show that CDTB is more challenging than previous datasets. State-of-the-art RGB trackers outperform the recent RGB-D trackers, indicating a large gap between the two fields, which has not been previously detected by the prior benchmarks. Based on the results of the analysis we point out opportunities for future research in RGB-D tracker design.
Alan Lukezic, Ugur Kart, Jani Käpylä, Ahmed Durmush, Joni-Kristian Kämäräinen, Jiri Matas, Matej Kristan
ICCV5
2019 Reverse Imaging Pipeline for Raw RGB Image Augmentation
abstract
We propose a reverse camera imaging pipeline to convert arbitrary images to raw RGB responses of a specific camera. The pipeline requires only that the camera's RGB responses are characterized. The reversed pipeline helps camera developers to generate camera specific raw images and use them to train learning-based imaging pipeline algorithms. In our experiments, three recent deep color constancy architectures achieve superior results in the cross-dataset setting using generated images.
Samu Koskinen, Dan Yang 0012, Joni-Kristian Kämäräinen
ICIP3
2019 Flash Lightens Gray Pixe
abstract
In the real world, a scene is usually cast by multiple illuminants and herein we address the problem of spatial illumination estimation. Our solution is based on detecting gray pixels with the help of flash photography. We show that flash photography significantly improves the performance of gray pixel detection without illuminant prior, training data or calibration of the flash. We also introduce a novel flash photography dataset generated from the MIT intrinsic dataset.
Yanlin Qian, Joni-Kristian Kämäräinen, Jiri Matas
ICIP3
2019 Portrait Instance Segmentation for Mobile Devices
abstract
Accurate and efficient portrait instance segmentation has become a crucial enabler for many multimedia applications on mobile devices. We present a novel convolutional neural network (CNN) architecture to explicitly address the long standing problems in portrait segmentation, i.e., semantic coherence and boundary localization. Specifically, we propose a cross-granularity categorical attention mechanism leveraging the deep supervisions to close the semantic gap of CNN feature hierarchy by imposing consistent category-oriented information across layers. Furthermore, a cross-granularity boundary enhancement module is proposed to boost the boundary awareness of deep layers by integrating the shape context cues from shallow layers of the network. We further propose a novel and efficient non-parametric affinity model to achieve efficient instance segmentation on mobile devices. We present a portrait image dataset with instance level annotations dedicated to evaluating portrait instance segmentation algorithms. We evaluate our approach on challenging datasets which obtains state-of-the-art results.
Lingyu Zhu 0001, Tinghuai Wang, Emre Aksu, Joni-Kristian Kämäräinen
ICME4
2019 Neural Network Pile Loading Controller Trained by Demonstration
abstract
This paper presents the development and testing of end-to-end Neural Network (NN) controllers for automated pile loading with a robotic wheel loader. NNs were trained using the Learning from Demonstration approach, i.e. by first recording sensor and control signals during manually-driven pile loading actions. Training made use of three input signals: boom angle, bucket angle and hydrostatic driving pressure; and three output signals: boom control, bucket control and the gas command. Most testing was conducted using NNs with 5 neurons in a single hidden layer, which were able to fill the bucket reasonably well. Qualitative comparisons were made to ascertain how the amount of training data and number of hidden neurons affects bucket filling performance, for NNs trained using both the Levenberg-Marquardt and Bayesian Regularization backpropagation algorithms. Different NNs trained with the same data were also compared. An additional pile transfer experiment compared the performance of an NN controller with a heuristic automated controller and manual human control. By estimating the total volume of material transferred using 3D laser scans, human control was found to have the highest performance, though the NN outperformed the heuristic controller. This indicated that end-to-end NN control trained by demonstration could offer improvement over current heuristic methods for automated pile loading.
Eric Halbach, Joni-Kristian Kämäräinen, Reza Ghabcheloo
ICRA2
2019 Proof of concept of a projection-based safety system for human-robot collaborative engine assembly
abstract
In the past years human-robot collaboration has gained interest among industry and production environments. While there is interest towards the topic, there is a lack of industrially relevant cases utilizing novel methods and technologies. The feasibility of the implementation, worker safety and production efficiency are the key questions in the field. The aim of the proposed work is to provide a conceptual safety system for context-dependent, multi-modal communication in human-robot collaborative assembly, which will contribute to safety and efficiency of the collaboration. The approach we propose offers an addition to traditional interfaces like push buttons installed at fixed locations. We demonstrate an approach and corresponding technical implementation of the system with projected safety zones based on the dynamically updated depth map and a graphical user interface (GUI). The proposed interaction is a simplified two-way communication between human and the robot to allow both parties to notify each other, and for the human to coordinate the operations.
Antti Hietanen, Alireza Changizi, Minna Lanz, Joni-Kristian Kämäräinen, Pallab Ganguly, Roel Pieters, Jyrki Latokartano
RO-MAN4
2019 Performance analysis of single-query 6-DoF camera pose estimation in self-driving setups
Junsheng Fu, Said Pertuz, Jiri Matas, Joni-Kristian Kämäräinen
Comput. Vis. Image Underst.4
2019 Cumulative attribute space regression for head pose estimation and color constancy
abstract
Two-stage Cumulative Attribute (CA) regression has been found effective in regression problems of computer vision such as facial age and crowd density estimation. The first stage regression maps input features to cumulative attributes that encode correlations between target values. The previous works have dealt with single output regression. In this work, we propose cumulative attribute spaces for 2- and 3-output (multivariate) regression. We show how the original CA space can be generalized to multiple output by the Cartesian product (CartCA). However, for target spaces with more than two outputs the CartCA becomes computationally infeasible and therefore we propose an approximate solution - multi-view CA (MvCA) - where CartCA is applied to output pairs. We experimentally verify improved performance of the CartCA and MvCA spaces in 2D and 3D face pose estimation and three-output (RGB) illuminant estimation for color constancy.
Ke Chen 0004, Kui Jia, Heikki Huttunen, Jiri Matas, Joni-Kristian Kämäräinen
Pattern Recognit.5
2019 Convolutional low-resolution fine-grained classification
Dingding Cai, Ke Chen 0004, Yanlin Qian, Joni-Kristian Kämäräinen
Pattern Recognit. Lett.4
2018 Depth Masked Discriminative Correlation Filter
abstract
Depth information provides a strong cue for occlusion detection and handling, but has been largely omitted in generic object tracking until recently due to lack of suitable benchmark datasets and applications. In this work, we propose a Depth Masked Discriminative Correlation Filter (DM-DCF) which adopts novel depth segmentation based occlusion detection that stops correlation filter updating and depth masking which adaptively adjusts the spatial support for correlation filter. In Princeton RGBD Tracking Benchmark, our DM-DCF is among the state-of-the-art in overall ranking and the winner on multiple categories. Moreover, since it is based on DCF, “DM-DCF” runs an order of magnitude faster than its competitors making it suitable for time constrained applications.
Ugur Kart, Joni-Kristian Kämäräinen, Jiri Matas, Lixin Fan, Francesco Cricri
ICPR2
2018 Object Detection in Equirectangular Panorama
abstract
We introduce a high-resolution equirectangular panorama (aka 360-degree, virtual reality, VR) dataset for object detection and propose a multi-projection variant of the YOLO detector. The main challenges with equirectangular panorama images are i) the lack of annotated training data, ii) high-resolution imagery and iii) severe geometric distortions of objects near the panorama projection poles. In this work, we solve the challenges by I) using training examples available in the “conventional datasets” (ImageNet and COCO), II) employing only low resolution images that require only moderate GPU computing power and memory, and III) our multi-projection YOLO handles projection distortions by making multiple stereographic sub-projections. In our experiments, YOLO outperforms the other state-of-the-art detector, Faster R-CNN, and our multi-projection YOLO achieves the best accuracy with low-resolution input.
Wenyan Yang, Yanlin Qian, Joni-Kristian Kämäräinen, Francesco Cricri, Lixin Fan
ICPR3
2018 Hierarchical Sliding Slice Regression for Vehicle Viewing Angle Estimation
abstract
We propose a novel hierarchical sliding slice regression which in a coarse-to-fine manner represents global circular target space with a number of ordinally localized and overlapping subspaces. Our method is particularly suitable for visual regression problems where the regression target is circular (e.g., car viewing angle) and visual similarity inconsistent over the target space (e.g., repetitive appearance). A good application example is the camera-based car viewing angle estimation problem, where visual similarity of different views is highly inconsistent-front and back views and left and right side views are pair-wise similar, but appear at the far ends of the circular view angle space. In practice, the problem is even more complicated due to large visual variation of objects (e.g., different car models). We perform extensive experiments on the Lausanne Federal of Institute of Technology Multi-view Car and KITTI Data Sets as well as the Technische Universitat Darmstadt Multi-view Pedestrians Data Set and achieve superior performance as compared to the state-of-the-art algorithms.
Dan Yang 0012, Yanlin Qian, Ke Chen 0004, Eleni Berki, Joni-Kristian Kämäräinen
IEEE Trans. Intell. Transp. Syst.5
2017 Recurrent Color Constancy
abstract
We introduce a novel formulation of temporal color constancy which considers multiple frames preceding the frame for which illumination is estimated. We propose an end-to-end trainable recurrent color constancy network – the RCC-Net – which exploits convolutional LSTMs and a simulated sequence to learn compositional representations in space and time. We use a standard single frame color constancy benchmark, the SFU Gray Ball Dataset, which can be adapted to a temporal setting. Extensive experiments show that the proposed method consistently outperforms single-frame state-of-the-art methods and their temporal variants.
Yanlin Qian, Ke Chen 0004, Jarno Nikkanen, Joni-Kristian Kämäräinen, Jiri Matas
ICCV4
2017 Region-based depth recovery for highly sparse depth maps
abstract
The accurate recovery of missing values in depth maps is an important problem in computer vision and image processing. In depth maps with large, irregular missing regions (i.e., sparse depth maps) inaccuracies arise when depth values of known pixels are used to recover depth near object edges and depth discontinuities (leakage). In order to overcome this problem, we propose an iterative region-based depth recovery method. In the proposed approach, the depth recovery problem is solved iteratively for each region of the segmented image in order to reduce the effect of leakage. Quantitative and qualitative experiments conducted on real data sets show promising results when comparing the proposed approach with state-of-the-art methods.
Said Pertuz, Joni-Kristian Kämäräinen
ICIP2
2017 Robustifying correspondence based 6D object pose estimation
abstract
We propose two methods to robustify point correspondence based 6D object pose estimation. The first method, curvature filtering, is based on the assumption that low curvature regions provide false matches, and removing points in these regions improves robustness. The second method, region pruning, is more general by making no assumptions about local surface properties. Our region pruning segments a model point cloud into cluster regions and searches good region combinations using a validation set. The robustifying methods are general and can be used with any correspondence based method. For the experiments, we evaluated three correspondence selection methods, Geometric Consistency (GC) [1], Hough Grouping (HG) [2] and Search of Inliers (SI) [3] and report systematic improvements for their robustified versions with two distinct datasets.
Antti Hietanen, Jussi Halme, Anders Glent Buch, Jyrki Latokartano, Joni-Kristian Kämäräinen
ICRA5
2017 Cross-Granularity Graph Inference for Semantic Video Object Segmentation
abstract
We address semantic video object segmentation via a novel cross-granularity hierarchical graphical model to integrate tracklet and object proposal reasoning with superpixel labeling. Tracklet characterizes varying spatial-temporal relations of video object which, however, quite often suffers from sporadic local outliers. In order to acquire high-quality tracklets, we propose a transductive inference model which is capable of calibrating short-range noisy object tracklets with respect to long-range dependencies and high-level context cues. In the center of this work lies a new paradigm of semantic video object segmentation beyond modeling appearance and motion of objects locally, where the semantic label is inferred by jointly exploiting multi-scale contextual information and spatial-temporal relations of video object. We evaluate our method on two popular semantic video object segmentation benchmarks and demonstrate that it advances the state-of-the-art by achieving superior accuracy performance than other leading methods.
Tinghuai Wang, Ke Chen 0004, Joni-Kristian Kämäräinen
IJCAI4
2017 Urban 3D segmentation and modelling from street view images and LiDAR point clouds
abstract
3D urban maps with semantic labels and metric information are not only essential for the next generation robots such autonomous vehicles and city drones, but also help to visualize and augment local environment in mobile user applications. The machine vision challenge is to generate accurate urban maps from existing data with minimal manual annotation. In this work, we propose a novel methodology that takes GPS registered LiDAR (Light Detection And Ranging) point clouds and street view images as inputs and creates semantic labels for the 3D points clouds using a hybrid of rule-based parsing and learning-based labelling that combine point cloud and photometric features. The rule-based parsing boosts segmentation of simple and large structures such as street surfaces and building facades that span almost 75% of the point cloud data. For more complex structures, such as cars, trees and pedestrians, we adopt boosted decision trees that exploit both structure (LiDAR) and photometric (street view) features. We provide qualitative examples of our methodology in 3D visualization where we construct parametric graphical models from labelled data and in 2D image segmentation where 3D labels are back projected to the street view images. In quantitative evaluation we report classification accuracy and computing times and compare results to competing methods with three popular databases: NAVTEQ True, Paris-Rue-Madame and TLS (terrestrial laser scanned) Velodyne.
Pouria Babahajiani, Lixin Fan, Joni-Kristian Kämäräinen, Moncef Gabbouj
Mach. Vis. Appl.3
2017 Spectral attribute learning for visual regression
Ke Chen 0004, Kui Jia, Zhaoxiang Zhang 0001, Joni-Kristian Kämäräinen
Pattern Recognit.4
2016 Learning with Ambiguous Label Distribution for Apparent Age Estimation
Ke Chen 0004, Joni-Kristian Kämäräinen
ACCV (3)2
2016 Deep structured-output regression learning for computational color constancy
abstract
The color constancy problem is addressed by structured-output regression on the values of the fully-connected layers of a convolutional neural network. The AlexNet and the VGG are considered and VGG slightly outperformed AlexNet. Best results were obtained with the first fully-connected “fc6” layer and with multi-output support vector regression. Experiments on the SFU Color Checker and Indoor Dataset benchmarks demonstrate that our method achieves competitive performance, outperforming the state of the art on the SFU indoor benchmark.
Yanlin Qian, Ke Chen 0004, Joni-Kristian Kämäräinen, Jarno Nikkanen, Jiri Matas
ICPR3
2016 Facial Age Estimation Using Robust Label Distribution
abstract
Facial age estimation, to predict the persons' exact ages given facial images, usually encounters the data sparsity problem due to the difficulties in data annotation. To mitigate the suffering from sparse data, a recent label distribution learning (LDL) algorithm attempts to embed label correlation into a classification based framework. However, the conventional label distribution learning framework only considers correlations across the neighbouring variables (ages), which omits the intrinsic complexity of age classes during different ageing periods (age groups). In the light of this, we introduce a novel concept of robust label distribution for scalar-valued labels, which is designed to encode the age scalars into label distribution matrices, i.e. two-dimensional Gaussian distributions along age classes and age groups respectively. Overcoming the limitations of conventional hard group boundaries in age grouping and capturing intrinsic inter-group dependency, our framework achieves robust and competitive performance over the conventional algorithms on two popular benchmarks for human age estimation.
Ke Chen 0004, Joni-Kristian Kämäräinen, Zhaoxiang Zhang 0001
ACM Multimedia2
2016 A comparison of feature detectors and descriptors for object class matching
Antti Hietanen, Jukka Lankinen, Joni-Kristian Kämäräinen, Anders Glent Buch, Norbert Krüger
Neurocomputing3
2016 Pedestrian Density Analysis in Public Scenes With Spatiotemporal Tensor Features
abstract
Pedestrian density estimation is one of the key problems in intelligent transportation systems and has been widely applied to a number of applications in other fields of engineering. Counting-by-regression methods are more favorable for coping with such a problem owing to their robustness against interperson occlusion and relaxing the impractical requirement of a high video frame rate, compared to counting-by-detection and counting-by-clustering methods. However, imagery features in the existing counting-by-regression approaches are extracted from the whole region or spatially localized cells/pixels of each single video frame, which omits the unique motion patterns of the same pedestrians across the neighboring frames. In the light of this, this paper exploits a novel tensor-formed spatiotemporal feature representation and applies it in a multilinear regression learning framework, which can capture spatially distributed dynamic crowd patterns by discovering the latent multidimensional structural correlations of tensor features along both spatial (i.e., horizontal and vertical) and temporal dimensions. Extensive evaluation with the public UCSD and Shopping Mall benchmarks demonstrate superior performance of our approach to the state-of-the-art counting methods even when the surveillance data has a low frame rate.
Ke Chen 0004, Joni-Kristian Kämäräinen
IEEE Trans. Intell. Transp. Syst.2
2015 Unsupervised visual alignment with similarity graphs
abstract
Alignment of semantically meaningful visual patterns, such as object classes, is an important pre-processing step for a number of applications such as object detection and image categorization. Considering the expensive manpower spent on the annotation for supervised alignment methods, unsupervised alignment techniques are more favorable especially for large-scale problems. Fine adjustment can be effectively and efficiently achieved with image congealing methods, but they require moderately good initialization which is largely invalid in practice. Alignment of visual class examples with large view point changes remains as an open problem. Feature-based methods can solve the problem to some degree, but require manual selection of a good seed image and omit the fact that examples of a semantic class can be visually very different (e.g., Harley-Davidsons and Scooters in “motorbikes”). In this work, we overcome the aforementioned drawbacks by defining visual similarity under the generalized assignment problem which is solved by fast approximation and non-linear optimization. From pair-wise image similarities we construct an image graph which is used to step-wise align, “morph”, an image to another by graph traveling. We automatically find a suitable seed by novel centrality measure which identifies “similarity hubs” in the graph. The proposed approach in the unsupervised manner outperforms the state-of-the-art methods with classes from the popular benchmark datasets.
Fatemeh Shokrollahi Yancheshmeh, Ke Chen 0004, Joni-Kristian Kämäräinen
CVPR3
2015 Flow feature extraction for underwater robot localization: Preliminary results
abstract
Underwater robots conventionally use vision and sonar sensors for perception purposes, but recently bio-inspired sensors that can sense flow have been developed. In literature, flow sensing has been shown to provide useful information about an underwater object and its surroundings. In the light of this, we develop an underwater landmark recognition technique which is based on the extraction and comparison of compact flow features. The proposed features are based on frequency spectrum of a pressure signal acquired by a piezo-resistive sensor. We report experiments in semi-natural (human-made flume with obstacles) and natural (river) underwater conditions where the proposed technique successfully recognizes previously visited locations.
Muhammad Naveed 0003, Nataliya Strokina, Gert Toming, Jeffrey A. Tuhtan, Joni-Kristian Kämäräinen, Maarja Kruusmaa
ICRA5
2015 Generative part-based Gabor object detector
Ekaterina Riabchenko, Joni-Kristian Kämäräinen
Pattern Recognit. Lett.2
2014 Learning to Count with Back-propagated Information
abstract
Error back-propagation is one of the principled learning strategies widely used in pattern recognition and machine learning, e.g. neural networks. The existing frameworks employed back-propagated error as a performance criteria (or termed, object function) aiming for supervising model-learning. Inspired by the recent success achieved by learning with the privileged information (LPI), we propose a novel regression-based framework by extending the concept of back-propagation in supervised learning methods to high-level guiding the model learning, so the proposed model is able to mine the importance of samples contributed to the fitting performance, which is missed in the existing regression techniques. To verify the effectiveness of the proposed learning paradigm, both low-level imagery features and intermediary semantic attributes are adopted in this paper. Extensive evaluations on pedestrian counting with public UCSD and Mall benchmarks demonstrate that the effectiveness of the proposed framework.
Ke Chen 0004, Joni-Kristian Kämäräinen
ICPR2
2014 Learning Generative Models of Object Parts from a Few Positive Examples
abstract
A number of computer vision problems such as object detection, pose estimation, and face recognition utilise local parts to represent objects, which include the distinguished information of objects. In this work, we introduce a novel probabilistic framework which automatically learns class-specific object parts (landmarks) in generative-learning manner. Encouraged by the success in learning and detecting facial landmarks, we employ bio-inspired multi-resolution Gabor features in the proposed framework. Specifically, complex-valued Gabor filter responses are first transformed to landmark specific likelihoods using Gaussian Mixture Models (GMM), and then efficient response matrix shift operations provide detection over orientations and scales. We avoid the undesirable characteristic of generative learning, a large number of training instances, with the novel concept of randomised Gaussian mixture model. Extensive experiments with public benchmarking Caltech-101 and BioID datasets demonstrate the effectiveness of our proposed method for localising object landmarks.
Ekaterina Riabchenko, Joni-Kristian Kämäräinen, Ke Chen 0004
ICPR2
2014 Density-Aware Part-Based Object Detection with Positive Examples
abstract
Part-based models have become the mainstream approach for visual object classification and detection. The key tools adopted by the most methods are interest point detectors and descriptors, shared codes for object parts (visual codebook) and discriminative learning using positive and negative class examples. Distinction of our method from the existing part-based methods for object detection is the use of sparse class-specific landmarks with semantic meaning. The landmarks are the additional distinguished information of object location in the proposed framework. Additionally, localising semantic and discriminative landmarks (object parts) is significant in other related applications of computer vision, such as facial expression recognition and pose/orientation estimation of objects. Therefore, we propose a model which deviates from the mainstream by the fact that the object parts' appearance and spatial variation, constellation, are explicitly modelled in a generative probabilistic manner. With using only positive examples our method can achieve object detection accuracy comparable to state-of-the-art discriminative method.
Ekaterina Riabchenko, Joni-Kristian Kämäräinen, Ke Chen 0004
ICPR2
2014 Live RGB-D camera tracking for television production studios
Tommi Tykkala, Andrew I. Comport, Joni-Kristian Kämäräinen, Hannu Hartikainen
J. Vis. Commun. Image Represent.3
2014 Best of 18th Scandinavian Conference On Image Analysis 2013
Joni-Kristian Kämäräinen, Markus Koskela
Pattern Recognit. Lett.1
2013 Pose estimation using local structure-specific shape and appearance context
abstract
We address the problem of estimating the alignment pose between two models using structure-specific local descriptors. Our descriptors are generated using a combination of 2D image data and 3D contextual shape data, resulting in a set of semi-local descriptors containing rich appearance and shape information for both edge and texture structures. This is achieved by defining feature space relations which describe the neighborhood of a descriptor. By quantitative evaluations, we show that our descriptors provide high discriminative power compared to state of the art approaches. In addition, we show how to utilize this for the estimation of the alignment pose between two point sets. We present experiments both in controlled and real-life scenarios to validate our approach.
Anders Glent Buch, Dirk Kraft, Joni-Kristian Kämäräinen, Henrik Gordon Petersen, Norbert Krüger
ICRA3
2013 Photorealistic 3D mapping of indoors by RGB-D scanning process
abstract
In this work, a RGB-D input stream is utilized for GPU-boosted 3D reconstruction of textured indoor environments. The goal is to develop a process which produces standard 3D models from indoors to explore them virtually. Camera motion is tracked in 3D space by registering the current view with a reference view. Depending on the trajectory shape, the reference is either fetched from a concurrently built keyframe model or from a previous RGB-D measurement. Realtime tracking (30Hz) is executed on a low-end GPU, which is possible because structural data is not fused concurrently. After camera poses have been estimated, both trajectory and structure are refined in post-processing. The global point cloud is compressed into a watertight polygon mesh by using Poisson reconstruction method. The Poisson method is well-suited, because it compressed the raw data without introducing multiple geometries and also fills holes efficiently. Holes are typically introduced at occluded regions. Texturing is generated by backprojecting the nearest RGB image onto the mesh. The final model is stored in a standard 3D model format to allow easy user exploration and navigation in virtual 3D environment.
Tommi Tykkala, Andrew I. Comport, Joni-Kristian Kämäräinen
IROS3
2012 Visual saliency and categorisation of abstract images
Mari Laine-Hernandez, Teemu Kinnunen, Joni-Kristian Kämäräinen, Lasse Lensu, Heikki Kälviäinen, Pirkko Oittinen
ICPR3
2012 A comparison of local feature detectors and descriptors for visual object categorization by intra-class repeatability and matching
Jukka Lankinen, Ville Kangas, Joni-Kristian Kämäräinen
ICPR3
2012 Unsupervised object discovery via self-organisation
Teemu Kinnunen, Joni-Kristian Kämäräinen, Lasse Lensu, Heikki Kälviäinen
Pattern Recognit. Lett.2
2011 Local Feature Based Unsupervised Alignment of Object Class Images
abstract
Alignment of objects is a predominant problem in visual object categorisation (VOC). State-of-the-art part-based VOC methods try to automatically learn object parts and their spatial variation, which is difficult for objects in arbitrary poses. A straightforward solution is to annotate images with a set of “object landmarks”, but due to laborious work required, less supervised methods are preferred. Effective semi-supervised VOC methods have been introduced, but none of them explicitly define an alignment procedure or study its effect to overall VOC performance. Unsupervised alignment has been recognised as its own problem referred to as “spatial image congealing” and a number of congealing methods have been proposed. These methods are mainly seminal work to Learned-Miller [3, 4] extending and improving the original algorithm. The main drawback of the congealing methods is that they are iterative optimisation methods operating on pixel-level and thus require at least moderate initial alignment to converge. Our approach [2] deviates from the congealing works by the fact that we utilise local features instead of pixel level processing, i.e. featurebased congealing. Our solution is more similar to those used in the partbased VOC methods, but we explicitly define the alignment algorithm and measure its performance.
Jukka Lankinen, Joni-Kristian Kämäräinen
BMVC2
2011 Bayesian network model of overall print quality: Construction and structural optimisation
Tuomas Eerola, Lasse Lensu, Joni-Kristian Kämäräinen, Tuomas Leisti, Risto Ritala, Göte Nyman, Heikki Kälviäinen
Pattern Recognit. Lett.3
2010 Learning and Detection of Object Landmarks in Canonical Object Space
abstract
This work contributes to part-based object detection and recognition by introducing an enhanced method for local part detection. The method is based on complex-valued multiresolution Gabor features and their ranking using multiple hypothesis testing. In the present work, our main contribution is the introduction of a canonical object space, where objects are represented in their ``expected pose and visual appearance''. The canonical space circumvents the problem of geometric image normalisation prior to feature extraction. In addition, we define a compact set of Gabor filter parameters, from where the optimal values can be easily devised. These enhancements make our method an attractive landmark detector for part-based object detection and recognition methods.
Joni-Kristian Kämäräinen, Jarmo Ilonen
ICPR1
2010 Unsupervised Visual Object Categorisation via Self-organisation
abstract
Visual object categorisation (VOC) has become one of the most actively investigated topic in computer vision. In the mainstream studies, the topic is considered as a supervised problem, but recently, the ultimate challenge has been posed: Unsupervised visual object categorisation. Hitherto only a few methods have been published, all of them being computationally demanding successors of their supervised counterparts. In this study, we address this problem with a simple and effective method: competitive learning leading to self organisation (self-categorisation). The unsupervised competitive learning approach is implemented using the Kohonen self-organising map algorithm (SOM). The SOM is used to perform the both unsupervised codebook generation and object categorisation. We present our method in detail and compare results to the supervised approach.
Teemu Kinnunen, Joni-Kristian Kämäräinen, Lasse Lensu, Heikki Kälviäinen
ICPR2
2010 Making Visual Object Categorization More Challenging: Randomized Caltech-101 Data Set
abstract
Visual object categorization is one of the most active research topics in computer vision, and Caltech-101 data set is one of the standard benchmarks for evaluating the method performance. Despite of its wide use, the data set has certain weaknesses: (i) the objects are practically in a standard pose and scale in the middle of the images and (ii) background varies too little in certain categories making it more discriminative than the foreground objects. In this work, we demonstrate how these weaknesses bias the evaluation results in an undesired manner. In addition, we reduce the bias effect by replacing the backgrounds with random landscape images from Google and by applying random Euclidean transformations to the foreground objects. We demonstrate how the proposed randomization process makes visual object categorization more challenging improving the relative results of methods which categorize objects by their visual appearance and are invariant to pose changes. The new data set is made publicly available for other researchers.
Teemu Kinnunen, Joni-Kristian Kämäräinen, Lasse Lensu, Jukka Lankinen, Heikki Kälviäinen
ICPR2
2008 Is there hope for predicting human visual quality experience?
abstract
One of the most important research goals in media science is a computational model for the human perception of visual quality, that is, how to predict the subjective visual quality experience. This research area has converged to developing new and investigating existing lower-level measurable quantities, physical, visual or computational, which could explain the high level experience. A principal research question, whether the prediction of the visual quality experience based on any lower-level objective measurements is possible at all, has received much less attention. This question is investigated in this study. First, we describe a large psychological experiment where true factors of the human quality experience are pair-wise resolved for dedicatedly selected samples. Second, we describe a ranking measure which reveals the relationship between selected measurable quantities and the human evaluation. Finally, the presented ranking method is used to provide quantitative evidence that visual quality experience can be predicted using lower-level measurable quantities. This result is novel and by simultaneously revealing the underlying lower-level factors it should re-direct the future research towards the true model.
Tuomas Eerola, Joni-Kristian Kämäräinen, Tuomas Leisti, Raisa Halonen, Lasse Lensu, Heikki Kälviäinen, Göte Nyman, Pirkko Oittinen
SMC2
2008 Finding best measurable quantities for predicting human visual quality experience
abstract
The literature of visual quality is mainly concentrated on devising new physical, visual, or computational quality features which could indirectly reflect ldquotrue visual qualityrdquo. The problem is that the true visual quality is always a subjective and context sensitive judgement of a single individual or a group of individuals. Therefore, the developed methods are only loosely connected to this ultimate objective, and the existing de facto and official standards have been designed by forming a consensus among experts of a specific field (e.g., in the printing industry). In this study, we describe a large psychological experiment where true factors of the human quality experience are pair-wise resolved for dedicatedly selected samples. Then we describe a ranking measure which reveals the relationship between selected measurable quantities and the human evaluation trial. Finally by using the above framework, we devise the best combinations from a set of well-known measurable quantities. The devised combinations can be considered as optimal when agreement with the human visual quality experience is desired, and therefore, they also reveal completely novel information about measuring visual quality.
Tuomas Eerola, Joni-Kristian Kämäräinen, Tuomas Leisti, Raisa Halonen, Lasse Lensu, Heikki Kälviäinen, Pirkko Oittinen, Göte Nyman
SMC2
2008 Detection of irregularities in regular patterns
Jarkko Vartiainen, Albert Sadovnikov, Joni-Kristian Kämäräinen, Lasse Lensu, Heikki Kälviäinen
Mach. Vis. Appl.3
2008 Image Feature Localization by Multiple Hypothesis Testing of Gabor Features
abstract
Several novel and particularly successful object and object category detection and recognition methods based on image features, local descriptions of object appearance, have recently been proposed. The methods are based on a localization of image features and a spatial constellation search over the localized features. The accuracy and reliability of the methods depend on the success of both tasks: image feature localization and spatial constellation model search. In this paper, we present an improved algorithm for image feature localization. The method is based on complex-valued multi resolution Gabor features and their ranking using multiple hypothesis testing. The algorithm provides very accurate local image features over arbitrary scale and rotation. We discuss in detail issues such as selection of filter parameters, confidence measure, and the magnitude versus complex representation, and show on a large test sample how these influence the performance. The versatility and accuracy of the method is demonstrated on two profoundly different challenging problems (faces and license plates).
Jarmo Ilonen, Joni-Kristian Kämäräinen, Pekka Paalanen, Miroslav Hamouz, Josef Kittler, Heikki Kälviäinen
IEEE Trans. Image Process.2
2007 The DIARETDB1 Diabetic Retinopathy Database and Evaluation Protocol
abstract
Automatic diagnosis of diabetic retinopathy from digital fundus images has been an active research topic in the medical image processing community. The research interest is justified by the excellent potential for new products in the medical industry and significant reductions in health care costs. However, the maturity of proposed algorithms cannot be judged due to the lack of commonly accepted and representative image database with a verified ground truth and strict evaluation protocol. In this study, an evaluation methodology is proposed and an image database with ground truth is described. The database is publicly available for benchmarking diagnosis algorithms. With the proposed database and protocol, it is possible to compare different algorithms, and correspondingly, analyse their maturity for technology transfer from the research laboratories to the medical practice.
Tomi Kauppi, Valentina Kalesnykiene, Joni-Kristian Kämäräinen, Lasse Lensu, Iiris Sorri, A. Raninen, R. Voutilainen, Hannu Uusitalo, Heikki Kälviäinen, Juhani Pietilä
BMVC3
2007 Object Localisation Using Generative Probability Model for Spatial Constellation and Local Image Features
abstract
In this paper we apply state-of-the-art approach to object detection and localisation by incorporating local descriptors and their spatial configuration into a generative probability model. In contrast to the recent semi- supervised methods we do not utilise interest point detectors, but apply a supervised approach where local image features (landmarks) are annotated in a training set and therefore their appearance and spatial variation can be learnt. Our method enables working in purely probabilistic search spaces providing a MAP estimate of object location, and in contrast to the recent methods, no background class needs to be formed. Using the training set we can estimate pdfs for both spatial constellation and local feature appearance. By applying an inference bias that the largest pdf mode has probability one, we are able to combine prior information (spatial configuration of the features) and observations (image feature appearance) into posterior distribution which can be generatively sampled, e.g. using MCMC techniques. The MCMC methods are sensitive to initialisation, but as a solution, we also propose a very efficient and accurate RANSAC-based method for finding good initial hypotheses of object poses. The complete method can robustly and accurately detect and localise objects under any homography.
Joni-Kristian Kämäräinen, Miroslav Hamouz, Josef Kittler, Pekka Paalanen, Jarmo Ilonen, Alexander Drobchenko
ICCV1
2006 Object categorization using self-organization over visual appearance
abstract
We propose an object categorization method which utilizes a feature structure, capturing object visual appearance, and the self-organizing map (SOM). The feature structure combines a set of spatially distant local receptive field responses with a constellation model which represents spatial relationships between the responses. The receptive field responses capture local appearance information and the spatial model generates a complete description of an object. The combination allows accurate representation of objects and their deformations. By the self-organization procedure unsupervised categorization over visual appearance of objects can be constructed. In addition, the proposed feature structure provides a reconstruction property, and thus, categorization can be used to visualize modalities of visual appearance. Categorization of real objects is demonstrated with human face images.
Jarmo Ilonen, Joni-Kristian Kämäräinen
IJCNN2
2006 Feature representation and discrimination based on Gaussian mixture model probability densities - Practices and algorithms
Pekka Paalanen, Joni-Kristian Kämäräinen, Jarmo Ilonen, Heikki Kälviäinen
Pattern Recognit.2
2006 Invariance properties of Gabor filter-based features-overview and applications
abstract
For almost three decades the use of features based on Gabor filters has been promoted for their useful properties in image processing. The most important properties are related to invariance to illumination, rotation, scale, and translation. These properties are based on the fact that they are all parameters of Gabor filters themselves. This is especially useful in feature extraction, where Gabor filters have succeeded in many applications, from texture analysis to iris and face recognition. This study provides an overview of Gabor filters in image processing, a short literature survey of the most significant results, and establishes invariance properties and restrictions to the use of Gabor filters in feature extraction. Results are demonstrated by application examples.
Joni-Kristian Kämäräinen, Ville Kyrki, Heikki Kälviäinen
IEEE Trans. Image Process.1
2005 Quantified and Perceived Unevenness of Solid Printed Areas
Albert Sadovnikov, Lasse Lensu, Joni-Kristian Kämäräinen, Heikki Kälviäinen
CIARP3
2005 Feature-Based Affine-Invariant Localization of Faces
abstract
We present a novel method for localizing faces in person identification scenarios. Such scenarios involve high resolution images of frontal faces. The proposed algorithm does not require color, copes well in cluttered backgrounds, and accurately localizes faces including eye centers. An extensive analysis and a performance evaluation on the XM2VTS database and on the realistic BioID and BANCA face databases is presented. We show that the algorithm has precision superior to reference methods.
Miroslav Hamouz, Josef Kittler, Joni-Kristian Kämäräinen, Pekka Paalanen, Heikki Kälviäinen, Jiri Matas
IEEE Trans. Pattern Anal. Mach. Intell.3
2004 Simple Gabor feature space for invariant object recognition
Ville Kyrki, Joni-Kristian Kämäräinen, Heikki Kälviäinen
Pattern Recognit. Lett.2
2003 Differential Evolution Training Algorithm for Feed Forward Neural Networks
Jarmo Ilonen, Joni-Kristian Kämäräinen, Jouni Lampinen
Neural Process. Lett.2
2003 Improving similarity measures of histograms using smoothing projections
Joni-Kristian Kämäräinen, Ville Kyrki, Jarmo Ilonen, Heikki Kälviäinen
Pattern Recognit. Lett.1