VLDB 2026 Research / reviewers in the wild / expert
Heikki Kälviäinen
dblp:03/6103
· DBLP profile ↗
71ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-0790-6847ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 4 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5Human-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Cross-Modal Learning for Plankton Recognition
Joona Kareinen, Veikka Immonen, Tuomas Eerola, Lumi Haraguchi, Lasse Lensu, Kaisa Kraft, Sanna Suikkanen, Heikki Kälviäinen |
ICPR (5) | 8 |
| 2026 | On Combining Animal Re-Identification Models to Address Small DatasetsabstractAbstract Recent advancements in the automatic re-identification of animal individuals from images have opened up new possibilities for studying wildlife through camera traps and citizen science projects. Existing methods leverage distinct and permanent visual body markings, such as fur patterns or scars, and typically employ one of two approaches: local features or end-to-end learning. The end-to-end learning-based methods outperform local feature-based methods given a sufficient amount of good-quality training data, but the challenge of gathering such datasets for wildlife animals means that local feature-based methods remain a more practical approach for many species. In this study, we aim to achieve two goals: (1) to obtain a better understanding of the impact of training-set size on animal re-identification, and (2) to explore ways to combine various methods to leverage the advantages of their approaches for re-identification. In the work, we conduct comprehensive experiments across six different methods and six animal species with various training set sizes. Furthermore, we propose a simple yet effective combination strategy and show that a properly selected method combinations outperform the individual methods with both small and large training sets up to 30%. Additionally, the proposed combination strategy offers a generalizable framework to improve accuracy across species and address the challenges posed by small datasets, which are common in ecological research. This work lays the foundation for more robust and accessible tools to support wildlife conservation, population monitoring, and behavioral studies. Aleksandr Algasov, Ekaterina A. Nepovinnykh, Fedor Zolotarev, Tuomas Eerola, Heikki Kälviäinen, Charles V. Stewart, Lasha Otarashvili, Jason Holmberg |
Int. J. Comput. Vis. | 5 |
| 2026 | EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction TuningabstractAbstract Facial expression recognition (FER) has emerged as an important research topic in recent years. However, current FER paradigms face challenges in generalization, lack semantic information aligned with natural language, and struggle to process both images and videos within a unified framework. Multimodal Large Language Models (MLLMs) have recently achieved success, offering advantages in addressing these issues and potentially overcoming the limitations of current FER paradigms. Nonetheless, directly applying pre-trained MLLMs to FER remains challenging due to insufficient instruction datasets and the inability of vision encoders to extract fine-grained facial information. Our zero-shot evaluations of existing open-source MLLMs on FER reveal a significant performance gap compared to GPT-4V and state-of-the-art supervised methods. In this paper, we aim to enhance MLLMs’ capabilities in understanding facial expressions. We first introduce a facial expression recognition instruction dataset ( FERID ), which has 376k category instructions and 339k conversational instructions. We then propose a novel MLLM, named EMO-LLaMA , which incorporates facial priors from a pretrained facial analysis network to enhance its understanding of human facial information. Specifically, we design a Face Info Mining module to extract both global and local facial information. Furthermore, we utilize a handcrafted prompt to introduce age-gender-race attributes, considering the emotional differences across diverse human groups. Extensive experiments show that EMO-LLaMA achieves results comparable to or competitive with SOTA on both static and dynamic FER datasets. The instruction dataset and code will be available at https://github.com/xxtars/EMO-LLaMA . Bohao Xing, Zitong Yu, Xin Liu 0012, Kaishen Yuan, Qilang Ye, Weicheng Xie 0001, Huanjing Yue, Heikki Kälviäinen |
Int. J. Comput. Vis. | 9 |
| 2026 | Context-aware railway video segmentation with unsupervised evaluationabstractAbstract In this work, we address the problem of segmenting railway tracks from video recordings captured from a train engine, focusing on the rails the train travels on. Our approach is motivated by the physical properties of railway rails, which typically exhibit a distinctive reflective appearance and remain visually and geometrically consistent across consecutive video frames. Moreover, the rails do not undergo significant shape changes or spatial displacement between frames. These properties enable the transfer of contextual information between frames and allow the use of composite models with a reduced number of trainable parameters. We introduce a composite model that incorporates contextual information from previous frames, reflecting the limited temporal variability of rail geometry and appearance, thereby reducing overfitting when training on small annotated datasets. We evaluate three neural network-based segmentation approaches: a convolutional neural network (CNN), a CNN combined with post-processing using morphological transformations and connected component labelling, and the proposed context-aware composite model. Furthermore, we propose an unsupervised evaluation methodology that assesses segmentation quality based on the distinguishability of segment colours from the background, geometric complexity, and colour variance within segments. All models are trained exclusively on synthetic data, with evaluation performed on real-world video sequences. Experimental results demonstrate that incorporating physically interpretable temporal context improves segmentation consistency compared to frame-wise methods. The synthetic dataset created for this work is publicly released. Vojtech Drahý, Radek Marik, Heikki Kälviäinen |
Pattern Anal. Appl. | 3 |
| 2026 | Identity-Free Artificial Emotional Intelligence via Micro-Gesture UnderstandingabstractIn this work, we focus on a special group of human body language — themicro-gesture (MG), which differs from the range of ordinary illustrative gestures in that they are not intentional behaviors performed to convey information to others, but rather unintentional behaviors driven by inner feelings. This characteristic introduces two novel challenges regarding micro-gestures that are worth rethinking. The first is whether strategies designed for other action recognition are entirely applicable to micro-gestures. The second is whether micro-gestures, as supplementary data, can provide additional insights for emotional understanding. In recognizing micro-gestures, we explore various augmentation strategies that take into account the subtle spatial and brief temporal characteristics of micro-gestures, often accompanied by repetitiveness, to determine more suitable augmentation methods. Considering the significance of temporal domain information for micro-gestures, we introduce a simple and efficient spatiotemporal balancing fusion method. We not only study our method on the considered micro-gesture dataset but also conduct experiments on mainstream gesture/action datasets. The results show that our approach performs well in micro-gesture recognition and on other datasets, achieving state-of-the-art performance compared to previous micro-gesture recognition methods. For emotional understanding based on micro-gestures, we construct complex emotional reasoning scenarios. Our evaluation, conducted with large language models, shows that micro-gestures play a significant and positive role in enhancing comprehensive emotional understanding. We confirm that our new insights contribute to advancing research in micro-gesture and emotional artificial intelligence. Rong Gao 0005, Xin Liu 0012, Bohao Xing, Zitong Yu, Björn W. Schuller, Heikki Kälviäinen |
IEEE Trans. Affect. Comput. | 6 |
| 2026 | MSF-Mamba: Motion-Aware State Fusion Mamba for Efficient Micro-Gesture RecognitionabstractMicro-gesture recognition (MGR) targets the identification of subtle and fine-grained human motions and requires accurate modeling of both long-range and local spatiotemporal dependencies. While convolutional neural networks (CNNs) are effective at capturing local patterns, they struggle with long-range dependencies due to their limited receptive fields. Transformer-based models address this limitation through self-attention mechanisms but suffer from high computational costs. Recently, Mamba has shown promise as an efficient model, leveraging state space models (SSMs) to enable linear-time processing. However, directly applying the vanilla Mamba to MGR may not be optimal. This is because Mamba processes inputs as 1D sequences, with state updates relying solely on the previous state, and thus lacks the ability to model local spatiotemporal dependencies. In addition, previous methods lack a design of motion-awareness, which is crucial in MGR. To overcome these limitations, we propose motion-aware state fusion mamba (MSF-Mamba), which enhances Mamba with local spatiotemporal modeling by fusing local contextual neighboring states. Our design introduces a motion-aware state fusion module based on central frame difference (CFD). Furthermore, a multiscale version named MSF-Mamba$^{+}$has been proposed. Specifically, MSF-Mamba$^{+}$supports multiscale motion-aware state fusion, as well as an adaptive scale weighting module that dynamically weighs the fused states across different scales. These enhancements explicitly address the limitations of vanilla Mamba by enabling motion-aware local spatiotemporal modeling, allowing MSF-Mamba and MSF-Mamba$^{+}$to effectively capture subtle motion cues for MGR. Experiments on two public MGR datasets (i.e., SMG and iMiGUE) demonstrate that even the lightweight version, namely, MSF-Mamba, achieves state-of-the-art performance, outperforming existing CNN-, Transformer-, and SSM-based models while maintaining high efficiency. For example, MSF-Mamba improves Top-1 accuracy by +2.2% and +1.5% over VideoMamba on SMG and iMiGUE, respectively. MSF-Mamba$^{+}$outperforms VideoMamba on SMG and iMiGUE, achieving Top-1 accuracy improvements of 2.9% and 3.0%, respectively. The code will be accessible onhttps://github.com/Leedeng/MSF-Mamba Deng Li 0002, Bohao Xing, Rong Gao 0005, Bihan Wen, Heikki Kälviäinen, Xin Liu 0012 |
IEEE Trans. Multim. | 6 |
| 2025 | FSBench: A Figure Skating Benchmark for Advancing Artistic Sports UnderstandingabstractFigure skating, known as the "Art on Ice," is among the most artistic sports, challenging to understand due to its blend of technical elements (like jumps and spins) and overall artistic expression. Existing figure skating datasets mainly focus on single tasks, such as action recognition or scoring, lacking comprehensive annotations for both technical and artistic evaluation. Current sports research is largely centered on ball games, with limited relevance to artistic sports like figure skating. To address this, we introduce FSAnno, a large-scale dataset advancing artistic sports understanding through figure skating. FSAnno includes an open-access training and test dataset, alongside a benchmark dataset, FSBench, for fair model evaluation. FSBench consists of FSBench-Text, with multiple-choice questions and explanations, and FSBench-Motion, containing multimodal data and Question and Answer (QA) pairs, supporting tasks from technical analysis to performance commentary. Initial tests on FSBench reveal significant limitations in existing models’ understanding of artistic sports. We hope FSBench will become a key tool for evaluating and enhancing model comprehension of figure skating. All data, models, and more details are available at: https://github.com/Moomin-Fin/Ano. Rong Gao 0005, Xin Liu 0012, Zhuozhao Hu, Bohao Xing, Baiqiang Xia, Zitong Yu, Heikki Kälviäinen |
CVPR | 7 |
| 2025 | AU-TTT: Vision Test-Time Training model for Facial Action Unit DetectionabstractFacial Action Units (AUs) detection is a cornerstone of objective facial expression analysis and a critical focus in affective computing. Despite its importance, AU detection faces significant challenges, such as the high cost of AU annotation and the limited availability of datasets. These constraints often lead to overfitting in existing methods, resulting in substantial performance degradation when applied across diverse datasets. Addressing these issues is essential for improving the reliability and generalizability of AU detection methods. Moreover, many current approaches leverage Transformers for their effectiveness in long-context modeling, but they are hindered by the quadratic complexity of self-attention. Recently, Test-Time Training (TTT) layers have emerged as a promising solution for long-sequence modeling. Additionally, TTT applies self-supervised learning for iterative updates during both training and inference, offering a potential pathway to mitigate the generalization challenges inherent in AU detection tasks. In this paper, we propose a novel vision backbone tailored for AU detection, incorporating bidirectional TTT blocks, named AU-TTT. Our approach introduces TTT Linear to the AU detection task and optimizes image scanning mechanisms for enhanced performance. Additionally, we design an AU-specific Region of Interest (RoI) scanning mechanism to capture fine-grained facial features critical for AU detection. Experimental results demonstrate that our method achieves competitive performance in both within-domain and cross-domain scenarios. Bohao Xing, Kaishen Yuan, Zitong Yu, Xin Liu 0012, Heikki Kälviäinen |
ICME | 5 |
| 2025 | DEEMO: De-identity Multimodal Emotion Recognition and ReasoningabstractEmotion understanding is a critical yet challenging task. Most existing approaches rely heavily on identity-sensitive information, such as facial expressions and speech, which raises concerns about personal privacy. To address this, we introduce the De-identity Multimodal Emotion Recognition and Reasoning ( DEEMO ), a novel task designed to enable emotion understanding using de-identified video and audio inputs. The DEEMO dataset consists of two subsets: DEEMO-NFBL , which includes rich annotations of Non-Facial Body Language (NFBL), and DEEMO-MER , an instruction dataset for Multimodal Emotion Recognition and Reasoning using identity-free cues. This design supports emotion understanding without compromising identity privacy. In addition, we propose DEEMO-LLaMA, a Multimodal Large Language Model (MLLM) that integrates de-identified audio, video, and textual information to enhance both emotion recognition and reasoning. Extensive experiments show that DEEMO-LLaMA achieves state-of-the-art performance on both tasks, outperforming existing MLLMs by a significant margin, achieving 74.49% accuracy and 74.45% F1-score in de-identity emotion recognition, and 6.20 clue overlap and 7.66 label overlap in de-identity emotion reasoning. Our work contributes to ethical AI by advancing privacy-preserving emotion understanding and promoting responsible affective computing. The dataset and codes will be available at https://github.com/Leedeng/DEEMO. Deng Li 0002, Bohao Xing, Xin Liu 0012, Baiqiang Xia, Bihan Wen, Heikki Kälviäinen |
ACM Multimedia | 6 |
| 2025 | Re-identification of patterned animals by multi-image feature aggregation and geometric similarityabstractAbstract Image‐based re‐identification of animal individuals allows gathering of information such as population size and migration patterns of the animals over time. This, together with large image volumes collected using camera traps and crowdsourcing, opens novel possibilities to study animal populations. For many species, the re‐identification can be done by analysing the permanent fur, feather, or skin patterns that are unique to each individual. In this paper, the authors study pattern feature aggregation based re‐identification and consider two ways of improving accuracy: (1) aggregating pattern image features over multiple images and (2) combining the pattern appearance similarity obtained by feature aggregation and geometric pattern similarity. Aggregation over multiple database images of the same individual allows to obtain more comprehensive and robust descriptors while reducing the computation time. On the other hand, combining the two similarity measures allows to efficiently utilise both the local and global pattern features, providing a general re‐identification approach that can be applied to a wide variety of different pattern types. In the experimental part of the work, the authors demonstrate that the proposed method achieves promising re‐identification accuracies for Saimaa ringed seals and whale sharks without species‐specific training or fine‐tuning. Ekaterina A. Nepovinnykh, Veikka Immonen, Tuomas Eerola, Charles V. Stewart, Heikki Kälviäinen |
IET Comput. Vis. | 5 |
| 2024 | DiffFAS: Face Anti-spoofing via Generative Diffusion Models
Xinxu Ge, Xin Liu 0012, Zitong Yu, Jingang Shi, Chun Qi, Heikki Kälviäinen |
ECCV (54) | 7 |
| 2024 | DAPlankton: Benchmark Dataset For Multi-Instrument Plankton Recognition Via Fine-Grained Domain AdaptationabstractPlankton recognition provides novel possibilities to study various environmental aspects and an interesting real-world context to develop domain adaptation (DA) methods. Different imaging instruments cause domain shift between datasets hampering the development of general plankton recognition methods. A promising remedy for this is DA allowing to adapt a model trained on one instrument to other instruments. In this paper, we present a new DA dataset called DAPlankton which consists of phytoplankton images obtained with different instruments. Phytoplankton provides a challenging DA problem due to the fine-grained nature of the task and high class imbalance in real-world datasets. DAPlankton consists of two subsets. $\mathrm{DAPlankton}_{\text {LAB }}$ contains images of cultured phytoplankton providing a balanced dataset with minimal label uncertainty. $\mathrm{DAPlankton}_{\text {SEA }}$ consists of images collected from the Baltic Sea providing challenging real-world data with large intra-class variance and class imbalance. We further present a benchmark comparison of three widely used DA methods. Daniel Batrakhanov, Tuomas Eerola, Kaisa Kraft, Lumi Haraguchi, Lasse Lensu, Sanna Suikkanen, María Teresa Camarena-Gómez, Jukka Seppälä, Heikki Kälviäinen |
ICIP | 9 |
| 2024 | Species-Agnostic Patterned Animal Re-identification by Aggregating Deep Local FeaturesabstractAbstract Access to large image volumes through camera traps and crowdsourcing provides novel possibilities for animal monitoring and conservation. It calls for automatic methods for analysis, in particular, when re-identifying individual animals from the images. Most existing re-identification methods rely on either hand-crafted local features or end-to-end learning of fur pattern similarity. The former does not need labeled training data, while the latter, although very data-hungry typically outperforms the former when enough training data is available. We propose a novel re-identification pipeline that combines the strengths of both approaches by utilizing modern learnable local features and feature aggregation. This creates representative pattern feature embeddings that provide high re-identification accuracy while allowing us to apply the method to small datasets by using pre-trained feature descriptors. We report a comprehensive comparison of different modern local features and demonstrate the advantages of the proposed pipeline on two very different species. Ekaterina A. Nepovinnykh, Ilja Chelak, Tuomas Eerola, Veikka Immonen, Heikki Kälviäinen, Maksim Kholiavchenko, Charles V. Stewart |
Int. J. Comput. Vis. | 5 |
| 2023 | Toward phytoplankton parasite detection using autoencodersabstractAbstract Phytoplankton parasites are largely understudied microbial components with a potentially significant ecological influence on phytoplankton bloom dynamics. To better understand the impact of phytoplankton parasites, improved detection methods are needed to integrate phytoplankton parasite interactions into monitoring of aquatic ecosystems. Automated imaging devices commonly produce vast amounts of phytoplankton image data, but the occurrence of anomalous phytoplankton data in such datasets is rare. Thus, we propose an unsupervised anomaly detection system based on the similarity between the original and autoencoder-reconstructed samples. With this approach, we were able to reach an overall F1 score of 0.75 in nine phytoplankton species, which could be further improved by species-specific fine-tuning. The proposed unsupervised approach was further compared with the supervised Faster R-CNN-based object detector. Using this supervised approach and the model trained on plankton species and anomalies, we were able to reach a highest F1 score of 0.86. However, the unsupervised approach is expected to be more universal as it can also detect unknown anomalies and it does not require any annotated anomalous data that may not always be available in sufficient quantities. Although other studies have dealt with plankton anomaly detection in terms of non-plankton particles or air bubble detection, our paper is, according to our best knowledge, the first that focuses on automated anomaly detection considering putative phytoplankton parasites or infections. Simon Bilik, Daniel Batrakhanov, Tuomas Eerola, Lumi Haraguchi, Kaisa Kraft, Silke Van den Wyngaert, Jonna Kangas, Conny Sjöqvist, Karin Madsen, Lasse Lensu, Heikki Kälviäinen, Karel Horák 0001 |
Mach. Vis. Appl. | 11 |
| 2020 | Resolving overlapping convex objects in silhouette images by concavity analysis and Gaussian processabstractThis paper introduces a novel method for segmentation of clustered partially overlapping convex objects in silhouette images. The proposed method involves three main steps: pre-processing, contour evidence extraction, and contour estimation. Contour evidence extraction starts by recovering contour segments from a binarized image by detecting concave points. After this the contour segments which belong to the same objects are grouped. The grouping is formulated as a combinatorial optimization problem and solved using the branch and bound algorithm. Finally, the full contours of the objects are estimated by a Gaussian process regression method. The experiments on a challenging dataset consisting of nanoparticles demonstrate that the proposed method outperforms three current state-of-art approaches in overlapping convex objects segmentation. The method relies only on edge information and can be applied to any segmentation problems where the objects are partially overlapping and have a convex shape. Sahar F. Zafar, Mariia Murashkina, Tuomas Eerola, Jouni Sampo, Heikki Kälviäinen, Heikki Haario |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Automated Segmentation of Nanoparticles in BF TEM Images by U-Net Binarization and Branch and Bound
Sahar F. Zafar, Tuomas Eerola, Heikki Kälviäinen, Alan C. Bovik |
CAIP (1) | 4 |
| 2019 | Timber Tracing with Multimodal Encoder-Decoder Networks
Fedor Zolotarev, Tuomas Eerola, Lasse Lensu, Heikki Kälviäinen, Heikki Haario, Jere Heikkinen, Tomi Kauppi |
CAIP (2) | 4 |
| 2018 | Two-Camera Synchronization and Trajectory Reconstruction for a Touch Screen Usability Experiment
Toni Kuronen, Tuomas Eerola, Lasse Lensu, Heikki Kälviäinen |
ACIVS | 4 |
| 2018 | Identification of Saimaa Ringed Seal Individuals Using Transfer Learning
Ekaterina A. Nepovinnykh, Tuomas Eerola, Heikki Kälviäinen, Gleb I. Radchenko |
ACIVS | 3 |
| 2018 | Comparison of Co-segmentation Methods for Wildlife Photo-identification
Anastasia Popova, Tuomas Eerola, Heikki Kälviäinen |
ACIVS | 3 |
| 2018 | Automatic individual identification of Saimaa ringed sealsabstractIn order to monitor an animal population and to track individual animals in a non‐invasive way, identification of individual animals based on certain distinctive characteristics is necessary. In this study, automatic image‐based individual identification of the endangered Saimaa ringed seal ( Phoca hispida saimensis ) is considered. Ringed seals have a distinctive permanent pelage pattern that is unique to each individual. This can be used as a basis for the identification process. The authors propose a framework that starts with segmentation of the seal from the background and proceeds to various post‐processing steps to make the pelage pattern more visible and the identification easier. Finally, two existing species independent individual identification methods are compared with a challenging data set of Saimaa ringed seal images. The results show that the segmentation and proposed post‐processing steps increase the identification performance. Tina Chehrsimin, Tuomas Eerola, Meeri Koivuniemi, Miina Auttila, Riikka Levänen, Marja Niemi, Mervi Kunnasranta, Heikki Kälviäinen |
IET Comput. Vis. | 8 |
| 2018 | Comparison of bubble detectors and size distribution estimatorsabstractDetection, counting and characterization of bubbles, that is, transparent objects in a liquid, is important in many industrial applications. These applications include monitoring of pulp delignification and multiphase dispersion processes common in the chemical, pharmaceutical, and food industries. Typically the aim is to measure the bubble size distribution. In this paper, we present a comprehensive comparison of bubble detection methods for challenging industrial image data. Moreover, we compare the detection-based methods to a direct bubble size distribution estimation method that does not require the detection of individual bubbles. The experiments showed that the approach based on a convolutional neural network (CNN) outperforms the other methods in detection accuracy. However, the boosting-based approaches were remarkably faster to compute. The power spectrum approach for direct bubble size distribution estimation produced accurate distributions and it is fast to compute, but it does not provide the spatial locations of the bubbles. Selecting the most suitable method depends on the specific application. Jarmo Ilonen, Roman Juránek, Tuomas Eerola, Lasse Lensu, Markéta Dubská, Pavel Zemcík, Heikki Kälviäinen |
Pattern Recognit. Lett. | 7 |
| 2017 | Towards Condition Analysis for Machine Vision Based Traffic Sign Inventory
Petri Hienonen, Lasse Lensu, Markus Melander, Heikki Kälviäinen |
ACIVS | 4 |
| 2017 | Multi-camera Finger Tracking and 3D Trajectory Reconstruction for HCI Studies
Vadim Lyubanenko, Toni Kuronen, Tuomas Eerola, Lasse Lensu, Heikki Kälviäinen, Jukka Häkkinen |
ACIVS | 5 |
| 2017 | Joint facial expression recognition and intensity estimation based on weighted votes of image sequences
Siti Khairuni Amalina Kamarol, Mohamed Hisham Jaward, Heikki Kälviäinen, Jussi Parkkinen, Rajendran Parthiban |
Pattern Recognit. Lett. | 3 |
| 2016 | Detection of bubbles as concentric circular arrangementsabstractThe paper proposes a method for the detection of bubble-like transparent objects in a liquid. The detection problem is non-trivial since bubble appearance varies considerably due to different lighting conditions causing contrast reversal and multiple interreflections. We formulate the problem as the detection of concentric circular arrangements (CCA). The CCAs are recovered in a hypothesize-optimize-verify framework. The hypothesis generation is based on sampling from the partially linked components of the non-maximum suppressed responses of oriented ridge filters, and is followed by the CCA parameter estimation. Parameter optimization is carried out by minimizing a novel cost-function. The performance was tested on gas dispersion images of pulp suspension and oil dispersion images. The mean error of gas/oil volume estimation was used as a performance criterion due to the fact that the main goal of the applications driving the research was the bubble volume estimation. The method achieved 28 and 13 % of gas and oil volume estimation errors correspondingly outperforming the OpenCV Circular Hough Transform in both cases and the WaldBoost detector in gas volume estimation. Nataliya Strokina, Jiri Matas, Tuomas Eerola, Lasse Lensu, Heikki Kälviäinen |
Mach. Vis. Appl. | 5 |
| 2015 | Segmentation of Overlapping Elliptical Objects in Silhouette ImagesabstractSegmentation of partially overlapping objects with a known shape is needed in an increasing amount of various machine vision applications. This paper presents a method for segmentation of clustered partially overlapping objects with a shape that can be approximated using an ellipse. The method utilizes silhouette images, which means that it requires only that the foreground (objects) and background can be distinguished from each other. The method starts with seedpoint extraction using bounded erosion and fast radial symmetry transform. Extracted seedpoints are then utilized to associate edge points to objects in order to create contour evidence. Finally, contours of the objects are estimated by fitting ellipses to the contour evidence. The experiments on one synthetic and two different real data sets showed that the proposed method outperforms two current state-of-art approaches in overlapping objects segmentation. Sahar F. Zafar, Tuomas Eerola, Jouni Sampo, Heikki Kälviäinen, Heikki Haario |
IEEE Trans. Image Process. | 4 |
| 2014 | Estimation of Bubble Size Distribution Based on Power Spectrum
Jarmo Ilonen, Tuomas Eerola, Heikki Mutikainen, Lasse Lensu, Jari Käyhkö, Heikki Kälviäinen |
CIARP | 6 |
| 2014 | Comparison of General Object Trackers for Hand Tracking in High-Speed VideosabstractThe problem of tracking a hand in video has gained a lot of attention due to its numerous applications in human computer interfaces. So far, the work has been limited to the use of standard speed videos, but the recent developments in imaging technology and computing hardware have made it attractive to exploit high-speed imaging for tracking the hand more accurately both in space and time. To produce videos of good quality, the high-speed imaging requires more light when compared to imaging with conventional frame rates. Therefore, grey-scale high-speed imaging is in common use and this makes the use of hand tracking methods relying specifically on color information unsuitable. In this work, we provide the first solid comparison of state-of-the-art general object trackers on hand tracking with a primary focus on grey-scale high-speed videos. Novel annotated high-speed video data were collected and made publicly available for evaluation purposes. The algorithms were tested with both finger and hand targets, and with grey-scale and color videos. In addition to tracking accuracies, the stability, sensitivity, and the processing speeds of the algorithms were evaluated. The experiments show that the results vary significantly in all aspects, but certain methods such as Compressive Tracking and Hough Track methods performed better overall. Ville Hiltunen, Tuomas Eerola, Lasse Lensu, Heikki Kälviäinen |
ICPR | 4 |
| 2013 | Framework for developing image-based dirt particle classifiers for dry pulp sheetsabstractOne important aspect of assessing the quality in pulp and papermaking is dirt particle counting and classification. Knowing the number and types of dirt particles present in pulp is useful for detecting problems in the production process as early as possible and for fixing them. Since manual quality control is a time-consuming and laborious task, the problem calls for an automated solution using machine vision techniques. However, the ground truth required to train an automated system is difficult to ascertain, since all of the dirt particles should be manually segmented and classified based on image information. This paper proposes a framework for developing and tuning dirt particle detection and classification systems. To avoid manual annotation, dry pulp sheets with a single dirt type in each were exploited to generate semisynthetic images with the ground truth information. To classify the dirt particles, a set of features were computed for each image segment. Sequential feature selection was employed to determine a close-to-optimal set of features to be used in classification. The framework was tested both with semisynthetically generated images based on real pulp sheets and with independent original real pulp sheets without any generation. The results of the experiments show that the semisynthetic procedure does not significantly change the properties of images and has little effect on the particle segmentation. The feature selection proved to be important when the number of dirt classes changes since it allows to improve the classification results. Using the standard classification methods, it is possible to obtain satisfactory results, although the methods modeling the data, such as the Bayesian classifier using the Gaussian Mixture Model, show better performance. Nataliya Strokina, Aki Mankki, Tuomas Eerola, Lasse Lensu, Jari Käyhkö, Heikki Kälviäinen |
Mach. Vis. Appl. | 6 |
| 2012 | Visual saliency and categorisation of abstract images
Mari Laine-Hernandez, Teemu Kinnunen, Joni-Kristian Kämäräinen, Lasse Lensu, Heikki Kälviäinen, Pirkko Oittinen |
ICPR | 5 |
| 2012 | Detection of bubbles as Concentric Circular Arrangements
Nataliya Strokina, Jiri Matas, Tuomas Eerola, Lasse Lensu, Heikki Kälviäinen |
ICPR | 5 |
| 2012 | Unsupervised object discovery via self-organisation
Teemu Kinnunen, Joni-Kristian Kämäräinen, Lasse Lensu, Heikki Kälviäinen |
Pattern Recognit. Lett. | 4 |
| 2011 | Bayesian network model of overall print quality: Construction and structural optimisation
Tuomas Eerola, Lasse Lensu, Joni-Kristian Kämäräinen, Tuomas Leisti, Risto Ritala, Göte Nyman, Heikki Kälviäinen |
Pattern Recognit. Lett. | 7 |
| 2010 | Unsupervised Visual Object Categorisation via Self-organisationabstractVisual object categorisation (VOC) has become one of the most actively investigated topic in computer vision. In the mainstream studies, the topic is considered as a supervised problem, but recently, the ultimate challenge has been posed: Unsupervised visual object categorisation. Hitherto only a few methods have been published, all of them being computationally demanding successors of their supervised counterparts. In this study, we address this problem with a simple and effective method: competitive learning leading to self organisation (self-categorisation). The unsupervised competitive learning approach is implemented using the Kohonen self-organising map algorithm (SOM). The SOM is used to perform the both unsupervised codebook generation and object categorisation. We present our method in detail and compare results to the supervised approach. Teemu Kinnunen, Joni-Kristian Kämäräinen, Lasse Lensu, Heikki Kälviäinen |
ICPR | 4 |
| 2010 | Making Visual Object Categorization More Challenging: Randomized Caltech-101 Data SetabstractVisual object categorization is one of the most active research topics in computer vision, and Caltech-101 data set is one of the standard benchmarks for evaluating the method performance. Despite of its wide use, the data set has certain weaknesses: (i) the objects are practically in a standard pose and scale in the middle of the images and (ii) background varies too little in certain categories making it more discriminative than the foreground objects. In this work, we demonstrate how these weaknesses bias the evaluation results in an undesired manner. In addition, we reduce the bias effect by replacing the backgrounds with random landscape images from Google and by applying random Euclidean transformations to the foreground objects. We demonstrate how the proposed randomization process makes visual object categorization more challenging improving the relative results of methods which categorize objects by their visual appearance and are invariant to pose changes. The new data set is made publicly available for other researchers. Teemu Kinnunen, Joni-Kristian Kämäräinen, Lasse Lensu, Jukka Lankinen, Heikki Kälviäinen |
ICPR | 5 |
| 2009 | Visual measurement and tracking in laser hybrid welding
Henri Fennander, Ville Kyrki, Anna Fellman, Antti Salminen, Heikki Kälviäinen |
Mach. Vis. Appl. | 5 |
| 2008 | Simple and Robust Optic Disc Localisation Using Colour Decorrelated Templates
Tomi Kauppi, Heikki Kälviäinen |
ACIVS | 2 |
| 2008 | Is there hope for predicting human visual quality experience?abstractOne of the most important research goals in media science is a computational model for the human perception of visual quality, that is, how to predict the subjective visual quality experience. This research area has converged to developing new and investigating existing lower-level measurable quantities, physical, visual or computational, which could explain the high level experience. A principal research question, whether the prediction of the visual quality experience based on any lower-level objective measurements is possible at all, has received much less attention. This question is investigated in this study. First, we describe a large psychological experiment where true factors of the human quality experience are pair-wise resolved for dedicatedly selected samples. Second, we describe a ranking measure which reveals the relationship between selected measurable quantities and the human evaluation. Finally, the presented ranking method is used to provide quantitative evidence that visual quality experience can be predicted using lower-level measurable quantities. This result is novel and by simultaneously revealing the underlying lower-level factors it should re-direct the future research towards the true model. Tuomas Eerola, Joni-Kristian Kämäräinen, Tuomas Leisti, Raisa Halonen, Lasse Lensu, Heikki Kälviäinen, Göte Nyman, Pirkko Oittinen |
SMC | 6 |
| 2008 | Finding best measurable quantities for predicting human visual quality experienceabstractThe literature of visual quality is mainly concentrated on devising new physical, visual, or computational quality features which could indirectly reflect ldquotrue visual qualityrdquo. The problem is that the true visual quality is always a subjective and context sensitive judgement of a single individual or a group of individuals. Therefore, the developed methods are only loosely connected to this ultimate objective, and the existing de facto and official standards have been designed by forming a consensus among experts of a specific field (e.g., in the printing industry). In this study, we describe a large psychological experiment where true factors of the human quality experience are pair-wise resolved for dedicatedly selected samples. Then we describe a ranking measure which reveals the relationship between selected measurable quantities and the human evaluation trial. Finally by using the above framework, we devise the best combinations from a set of well-known measurable quantities. The devised combinations can be considered as optimal when agreement with the human visual quality experience is desired, and therefore, they also reveal completely novel information about measuring visual quality. Tuomas Eerola, Joni-Kristian Kämäräinen, Tuomas Leisti, Raisa Halonen, Lasse Lensu, Heikki Kälviäinen, Pirkko Oittinen, Göte Nyman |
SMC | 6 |
| 2008 | Detection of irregularities in regular patterns
Jarkko Vartiainen, Albert Sadovnikov, Joni-Kristian Kämäräinen, Lasse Lensu, Heikki Kälviäinen |
Mach. Vis. Appl. | 5 |
| 2008 | Image Feature Localization by Multiple Hypothesis Testing of Gabor FeaturesabstractSeveral novel and particularly successful object and object category detection and recognition methods based on image features, local descriptions of object appearance, have recently been proposed. The methods are based on a localization of image features and a spatial constellation search over the localized features. The accuracy and reliability of the methods depend on the success of both tasks: image feature localization and spatial constellation model search. In this paper, we present an improved algorithm for image feature localization. The method is based on complex-valued multi resolution Gabor features and their ranking using multiple hypothesis testing. The algorithm provides very accurate local image features over arbitrary scale and rotation. We discuss in detail issues such as selection of filter parameters, confidence measure, and the magnitude versus complex representation, and show on a large test sample how these influence the performance. The versatility and accuracy of the method is demonstrated on two profoundly different challenging problems (faces and license plates). Jarmo Ilonen, Joni-Kristian Kämäräinen, Pekka Paalanen, Miroslav Hamouz, Josef Kittler, Heikki Kälviäinen |
IEEE Trans. Image Process. | 6 |
| 2007 | The DIARETDB1 Diabetic Retinopathy Database and Evaluation ProtocolabstractAutomatic diagnosis of diabetic retinopathy from digital fundus images has been an active research topic in the medical image processing community. The research interest is justified by the excellent potential for new products in the medical industry and significant reductions in health care costs. However, the maturity of proposed algorithms cannot be judged due to the lack of commonly accepted and representative image database with a verified ground truth and strict evaluation protocol. In this study, an evaluation methodology is proposed and an image database with ground truth is described. The database is publicly available for benchmarking diagnosis algorithms. With the proposed database and protocol, it is possible to compare different algorithms, and correspondingly, analyse their maturity for technology transfer from the research laboratories to the medical practice. Tomi Kauppi, Valentina Kalesnykiene, Joni-Kristian Kämäräinen, Lasse Lensu, Iiris Sorri, A. Raninen, R. Voutilainen, Hannu Uusitalo, Heikki Kälviäinen, Juhani Pietilä |
BMVC | 9 |
| 2006 | Smooth Transition from Motion to Force Control in Robotic Manipulation Using VisionabstractSensor-based robotic manipulation is becoming more and more popular as it promises increases in productivity, flexibility, and robustness of manipulation. Combining visual and force sensing is currently one of the most promising approaches for sensor-based manipulation, as vision and force are two complementary sensing modalities. One approach for multi-sensor use is the traded control where the robot is at each time controlled using one sensing modality, and the controllers are switched based on sensory input. One of the major problems with such systems is the transition between visual and force controllers. In this paper, we present a smooth transition method from motion to force control. The velocity of the end-effector is controlled by estimating the distance to the target by vision and determining an optimal velocity profile giving rapid approach and minimal force overshoot. Experiments show that the proposed control scheme is superior to earlier approaches Olli Alkkiomäki, Ville Kyrki, Heikki Kälviäinen, Heikki Handroos |
ICARCV | 3 |
| 2006 | Feature representation and discrimination based on Gaussian mixture model probability densities - Practices and algorithms
Pekka Paalanen, Joni-Kristian Kämäräinen, Jarmo Ilonen, Heikki Kälviäinen |
Pattern Recognit. | 4 |
| 2006 | Invariance properties of Gabor filter-based features-overview and applicationsabstractFor almost three decades the use of features based on Gabor filters has been promoted for their useful properties in image processing. The most important properties are related to invariance to illumination, rotation, scale, and translation. These properties are based on the fact that they are all parameters of Gabor filters themselves. This is especially useful in feature extraction, where Gabor filters have succeeded in many applications, from texture analysis to iris and face recognition. This study provides an overview of Gabor filters in image processing, a short literature survey of the most significant results, and establishes invariance properties and restrictions to the use of Gabor filters in feature extraction. Results are demonstrated by application examples. Joni-Kristian Kämäräinen, Ville Kyrki, Heikki Kälviäinen |
IEEE Trans. Image Process. | 3 |
| 2005 | Quantified and Perceived Unevenness of Solid Printed Areas
Albert Sadovnikov, Lasse Lensu, Joni-Kristian Kämäräinen, Heikki Kälviäinen |
CIARP | 4 |
| 2005 | Finding and Ranking Research Directions for Software Testing
Ossi Taipale, Kari Smolander, Heikki Kälviäinen |
EuroSPI | 3 |
| 2005 | Feature-Based Affine-Invariant Localization of FacesabstractWe present a novel method for localizing faces in person identification scenarios. Such scenarios involve high resolution images of frontal faces. The proposed algorithm does not require color, copes well in cluttered backgrounds, and accurately localizes faces including eye centers. An extensive analysis and a performance evaluation on the XM2VTS database and on the realistic BioID and BANCA face databases is presented. We show that the algorithm has precision superior to reference methods. Miroslav Hamouz, Josef Kittler, Joni-Kristian Kämäräinen, Pekka Paalanen, Heikki Kälviäinen, Jiri Matas |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2004 | Classification of Computerized Learning Tools for Introductory Programming Courses: Learning ApproachabstractLearning programming is a difficult task since programming requires new concepts in thinking and creative skills in problem solving. A number of learning tools and environments have been built to assist both teachers and students in introductory programming courses. In this study, we have established a classification for these tools. Tools are divided into four categories: A) integrated development interface; B) visualization; C) virtual learning environments; and D) systems for submitting, managing, and testing of exercises. The classification is based on a review of existing tools, both commercial and freely available. Guidelines for the selection of a suitable tool are discussed. Timo Rongas, Arto Kaarna, Heikki Kälviäinen |
ICALT | 3 |
| 2004 | Simple Gabor feature space for invariant object recognition
Ville Kyrki, Joni-Kristian Kämäräinen, Heikki Kälviäinen |
Pattern Recognit. Lett. | 3 |
| 2003 | Geometric correction and classification of images in change detection of water plants in SoinilansalmiabstractAutomation of change detection greatly enhances environmental monitoring. This study describes possibilities for automation of change detection through image processing actions that include geometric correction, classification and segmentation. The water area was photographed twice and the changes in water vegetation were under consideration. The results from image processing monitoring were compared to the on-site measurements. The correspondence between these two was satisfactory. The automation of the change detection would require both on-site actions and controlled photographing together with well defined image processing. Arto Kaarna, Heikki Kälviäinen, Alexey Anufriev, Jukka Mankki, Terhi Malkavaara, Matti Jantunen |
IGARSS | 2 |
| 2003 | Intermediate-level feature extraction in novel parallel environments
Ville Kyrki, Jani Peusaari, Heikki Kälviäinen |
Mach. Vis. Appl. | 3 |
| 2003 | Improving similarity measures of histograms using smoothing projections
Joni-Kristian Kämäräinen, Ville Kyrki, Jarmo Ilonen, Heikki Kälviäinen |
Pattern Recognit. Lett. | 4 |
| 2001 | Positioning of Flexible Boom Structure Using Neural Networks
Jarno Mielikäinen, Ilkka Koskinen, Heikki Handroos, Pekka J. Toivanen, Heikki Kälviäinen |
CAIP | 5 |
| 2000 | Hough Transform for Rotation Invariant Matching of Line-Drawing ImagesabstractHough transform can be used for indexing of line-drawing images for content-based image retrieval. Angular information is used for generating the feature vector (index) as it gives global description of the image, allows compact indexing, fast retrieval and scale, translation and rotation invariant matching. In the case of very large images, however, the angular information is not always sufficient to differentiate images from each other. To alleviate this problem, we extend the idea by including also positional information of the lines in the feature vector. This gives more representative description of the images and therefore allows more accurate image matching. The main problems of this approach are: (1) to keep the feature vector compact, and (2) to preserve the property of the matching being translation and rotation invariant. We give solutions to both of these problems and introduce a new indexing scheme, which has better matching accuracy but at the cost of slower retrieval time. Pasi Fränti, Alexey Mednonogov, Heikki Kälviäinen |
ICPR | 3 |
| 2000 | Multispectral VideoabstractThis paper describes a new approach to compressing temporal sequences of multispectral images by encoding and presenting them as multispectral video. The compression methods under study include motion compensation and principal component analysis. Motion compensation is utilized to reduce repeated information between consecutive pictures in a sequence. Using the principal component analysis, an attempt is made to retain only necessary parts of the information in a picture. Panu Koponen, Heikki Kälviäinen, Jussi Parkkinen |
ICPR | 2 |
| 2000 | High Precision 2-D Geometrical InspectionabstractAutomated visual inspection has become important for modern industry mainly because production rates and the level of automation have continually increased. The paper presents a system for automated visual inspection of large two-dimensional parts. The system is capable of inspecting sheet metal parts using CAD data. The inspection is performed based on a CAD model. There are existing systems that perform this function but they are not particularly well suitable for in-place inspection. In the proposed system, the high precision inspection of individual features is performed by firstly estimating the global position of a part. Next, each local feature is measured using subpixel techniques. Finally, the measurements are compared with a CAD model. Experiments are presented to evaluate the precision of system components. According to the results, the calibration procedure seems to be the factor that has the greatest effect on the final precision. Last, a comparison to similar systems is made and some suggestions are given how the precision could be further improved. Ville Kyrki, Heikki Kälviäinen |
ICPR | 2 |
| 2000 | Multispectral Image Color EncodingabstractAdvances in colour image sensors and computer technology have opened up new opportunities for colour-based image analysis by allowing it to use multispectral images - images that have tens or even hundreds of spectral colour channels. As multispectral images usually occupy large amounts of memory, some suitable way is needed to compress such images and to represent them efficiently. Image compression has been one of the mainstream research topics for a long time; however, the research usually focuses on compressing images that are intended to be finally seen by humans. While many traditional methods can also be reused for multispectral images, some features of the multispectral images can be addressed differently than in traditional images. This paper describes a way of how to represent the multispectral colour data as a linear combination of colours contained in the clusters of colours that most frequently occur in an image. Pavel Zemcík, Jan Vorácek, Michael Frydrych, Heikki Kälviäinen, Pekka J. Toivanen |
ICPR | 4 |
| 2000 | Content-based matching of line-drawing images using the Hough transform
Pasi Fränti, Alexey Mednonogov, Ville Kyrki, Heikki Kälviäinen |
Int. J. Document Anal. Recognit. | 4 |
| 2000 | Randomized or probabilistic Hough transform: unified performance evaluation
Nahum Kiryati, Heikki Kälviäinen, Satu Alaoutinen |
Pattern Recognit. Lett. | 2 |
| 2000 | Compression of multispectral remote sensing images using clustering and spectral reductionabstractImage compression has been one of the main research topics in the field of image processing for a long time. The research usually focuses on compressing images that are visible to humans. The images being compressed are usually gray-level images or RGB color images. Recent advances in technology, however, enable the authors to make the detailed processing of spectral features in the images. Therefore, the compression of images with many spectral channels, called multispectral images, is required. Many methods used in traditional lossy image compression can be reused also in the compression of multispectral images. In this paper, a new combination of clustering spectra, manipulating spectral vectors, and encoding and decoding for multispectral images is presented. In the manipulation of the spectral vectors PCA, ICA, and wavelets are used. The approach is based on extracting relevant spectral information. Furthermore, some quantitative quality measures for multispectral images are presented. Arto Kaarna, Pavel Zemcík, Heikki Kälviäinen, Jussi Parkkinen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 1998 | Multispectral image compressionabstractImage compression has been one of the mainstream research topics in image processing. The research usually focuses on compressing images that are visible to humans. Images are usually gray-level images or RGB color images. Advances in technology enable one to make the detailed processing of spectral color features in the images. Therefore, compression of images with many spectral color channels, called multispectral images, is required. Many methods used in traditional lossy image compression can be reused also in the compression of multispectral images. In this paper a new combination of clustering of colors, manipulating spectral color encoding and decoding for multispectral images is presented. The approach is based on extracting relevant color information. Furthermore, some quantitative quality measures for multispectral images are presented. Arto Kaarna, Pavel Zemcík, Heikki Kälviäinen, Jussi Parkkinen |
ICPR | 3 |
| 1997 | Robust unmixing of large sets of mixed pixels
Panagiota Bosdogianni, Heikki Kälviäinen, Maria Petrou, Josef Kittler |
Pattern Recognit. Lett. | 2 |
| 1997 | An extension to the randomized hough transform exploiting connectivity
Heikki Kälviäinen, Petri Hirvonen |
Pattern Recognit. Lett. | 1 |
| 1996 | Mixed pixel classification with the randomized Hough transformabstractWe propose the use of the randomized Hough transform algorithm for the determination of the proportions of pure classes present in sets of mixed pixels, for large datasets (for which the deterministic Hough is prohibitively slow) and in the presence of outliers (i.e. in cases that the classical least square error method cannot cope). We demonstrate our results both with simulated and laboratory real data. Heikki Kälviäinen, Panagiota Bosdogianni, Maria Petrou, Josef Kittler |
ICPR | 1 |
| 1996 | Houghtool -- A software package for the use of the Hough transform
Heikki Kälviäinen, Petri Hirvonen, Erkki Oja |
Pattern Recognit. Lett. | 1 |
| 1995 | Probabilistic and non-probabilistic Hough transforms: overview and comparisons
Heikki Kälviäinen, Petri Hirvonen, Lei Xu 0001, Erkki Oja |
Image Vis. Comput. | 1 |
| 1994 | Comparisons of Probabilistic and Non-probabilistic Hough Transforms
Heikki Kälviäinen, Petri Hirvonen, Lei Xu 0001, Erkki Oja |
ECCV (2) | 1 |
| 1993 | Motion Estimation and the Randomized Hough Transform (RHT): New Methods with Gradient Information
Heikki Kälviäinen |
CAIP | 1 |
| 1992 | Randomized Hough transform applied to translational and rotational motion analysisabstractA method has been developed to calculate 2-D motion in a sequence of time-varying images. The method, called motion detection using randomized Hough transform (MDRHT), is based on the randomized Hough transform (RHT). The RHT decreases considerably the time consumption and memory requirements of the Hough transform. The idea of the MDRHT is to pick randomly point pairs from two images and calculate the translation with them. The points can be e.g. edge points of the original images. This approach can avoid difficulties of standard segmentation methods like overlapping and covering, and has the advantages provided by the RHT. The method can be generalized by picking more than two points. After a brief review of the RHT applied to motion detection, the extended algorithm to calculate both translation and rotation is represented in this paper.> Heikki Kälviäinen, Erkki Oja, Lei Xu 0001 |
ICPR (1) | 1 |