VLDB 2026 Research / reviewers in the wild / expert
Jorma Laaksonen
dblp:71/4708
· DBLP profile ↗
117ranked-venue papers
11as first author
33since 2021 · last 2026
0000-0001-7218-3131ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 80 · 11 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 45 · 1 first-author · 18 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 7 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowledge-Integrated Reasoning: A Novel Approach for External Knowledge Based Visual Question Answering
Pyry Satama, Abduljalil Radman, Jorma Laaksonen |
ICPR (7) | 3 |
| 2026 | ARB: A Comprehensive Arabic Multimodal Reasoning BenchmarkabstractAs Large Multimodal Models (LMMs) become more capable, there is growing interest in evaluating their reasoning processes alongside their final outputs. However, most benchmarks remain focused on English, overlooking languages with rich linguistic and cultural contexts, such as Arabic. To address this gap, we introduce the Comprehensive Arabic Multimodal Reasoning Benchmark (ARB), the first benchmark designed to evaluate step-by-step reasoning in Arabic across both textual and visual modalities. ARB spans 11 diverse domains, including visual reasoning, document understanding, OCR, scientific analysis, and cultural interpretation. It comprises 1,356 multimodal samples paired with 5,119 human-curated reasoning steps and corresponding actions. We evaluated 12 state-of-the-art open- and closed-source LMMs and found persistent challenges in coherence, faithfulness, and cultural grounding. ARB offers a structured framework for diagnosing multimodal reasoning in underrepresented languages and marks a critical step toward inclusive, transparent, and culturally aware AI systems. We release the benchmark, rubric, and evaluation suit to support future research and reproducibility. Code available at: https://github.com/mbzuai-oryx/ARB Sara Ghaboura, Shubham Patle, Ketan More, Wafa Hamad Mohamed Alghallabi, Omkar Thawakar, Jorma Laaksonen, Hisham Cholakkal, Salman Khan 0001, Rao Muhammad Anwer |
LREC | 6 |
| 2025 | TSAM: Temporal SAM Augmented with Multimodal Prompts for Referring Audio-Visual SegmentationabstractReferring audio-visual segmentation (Ref-AVS) aims to segment objects within audio-visual scenes using multimodal cues embedded in text expressions. While the Segment Anything Model (SAM) has revolutionized visual segmentation, its applicability to Ref-AVS, where multimodal cues act as novel prompts, remains unexplored. SAM’s limitation to single-frame segmentation also hinders its ability to capture essential temporal context needed for multi-frame audio-visual segmentation. To address this gap, we propose TSAM, a novel extension of SAM designed to leverage multimodal cues for precise segmentation in dynamic audio-visual scenes. TSAM enhances SAM’s image encoder with a temporal modeling branch, enabling spatio-temporal learning and deep multimodal fusion across video frames, while retaining SAM’s pre-trained knowledge. Additionally, TSAM replaces SAM’s user-interactive prompting mechanism with sparse and dense data-driven prompts, enabling more effective integration of audio-visual inputs and reference text expressions. Extensive experiments on the Ref-AVS dataset demonstrate TSAM’s superiority over state-of-the-art methods. The results illustrate its effectiveness in segmenting objects in dynamic audio-visual scenes using text-based multimodal cues and its strong generalization to unseen objects. Project webpage: https://abdurad.github.io/TSAM/. Abduljalil Radman, Jorma Laaksonen |
CVPR | 2 |
| 2025 | All Languages Matter: Evaluating LMMs on Culturally Diverse 100 LanguagesabstractExisting Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corresponding visual cues. In pursuit of culturally diverse global multimodal models, our proposed All Languages Matter Benchmark (ALM-bench) represents the largest and most comprehensive effort to date for evaluating LMMs across 100 languages. ALM-bench challenges existing models by testing their ability to understand and reason about culturally diverse images paired with text in various languages, including many low-resource languages traditionally underrepresented in LMM research. The benchmark offers a robust and nuanced evaluation framework featuring various question formats, including true/false, multiple choice, and open-ended questions, which are further divided into short and long-answer categories. ALM-bench design ensures a comprehensive assessment of a model’s ability to handle varied levels of difficulty in visual and linguistic reasoning. To capture the rich tapestry of global cultures, ALM-bench carefully curates content from 13 distinct cultural aspects, ranging from traditions and rituals to famous personalities and celebrations. Through this, ALM-bench not only provides a rigorous testing ground for state-of-the-art open and closed-source LMMs but also highlights the importance of cultural and linguistic inclusivity, encouraging the development of models that can serve diverse global populations effectively. Our benchmark is publicly available at https://mbzuai-oryx.github.io/ALM-Bench/. Ashmal Vayani, Dinura Dissanayake, Hasindri Watawana, Noor Ahsan, Nevasini Sasikumar, Omkar Thawakar, Henok Biadglign Ademtew, Yahya Hmaiti, Amandeep Kumar, Kartik Kuckreja, Mykola Maslych, Wafa Al Ghallabi, Mihail Minkov Mihaylov, Abdelrahman M. Shaker, Mike Zhang, Mahardika Krisna Ihsani, Amiel Esplana, Monil Gokani, Shachar Mirkin, Harsh Singh, Ashay Srivastava, Endre Hamerlik, Fathinah Asma Izzati, Fadillah A. Maani, Sebastian Cavada, Jenny Chim, Rohit Gupta 0012, Sanjay Manjunath, Kamila Zhumakhanova, Feno Heriniaina Rabevohitra, Azril Hafizi Amirudin, Muhammad Ridzuan, Daniya Najiha Abdul Kareem, Ketan More, Pramesh Shakya, Amirpouya Ghasemaghaei, Amirbek Djanibekov, Dilshod Azizov, Branislava Jankovic, Naman Bhatia, Alvaro Cabrera, Johan S. Obando-Ceron, Olympiah Otieno, Fabian Farestam, Muztoba Rabbani, Sanoojan Baliah, Santosh Sanjeev, Abduragim Shtanchaev, Maheen Fatima, Amrin Kareem, Toluwani Aremu, Nathan A. Z. Xavier, Amit Bhatkal, Hawau Olamide Toyin, Aman Chadha, Hisham Cholakkal, Rao Muhammad Anwer, Michael Felsberg, Jorma Laaksonen, Thamar Solorio, Monojit Choudhury, Ivan Laptev, Mubarak Shah, Salman Khan 0001, Fahad Shahbaz Khan |
CVPR | 63 |
| 2025 | Valor32k-AVQA v2.0: Open-Ended Audio-Visual Question Answering Dataset and BenchmarkabstractDespite growing interest in Audio-Visual Question Answering (AVQA), existing datasets often suffer from limited diversity, rigid formats, and insufficient integration of audio and visual modalities. To address these limitations, we introduce Valor32k-AVQA v2.0, a large-scale dataset containing 28,863 real-world videos and over 225,000 QA pairs, designed to support diverse and realistic multimodal understanding. The dataset features both open-ended and multiple-choice questions, each annotated with the required modality ( visual, audio, or audio-visual ) and question category ( description, action, count, temporal, location, or relative position ). All annotations-including questions, answers, and metadata-are generated through a fully automated prompting pipeline using GPT-4o, with human validation performed on a representative sample to ensure quality. We benchmark a few state-of-the-art models, with additional evaluations available on the project page, and observe that incorporating audio consistently improves performance during fine-tuning without compromising visual reasoning capabilities. These findings highlight that the audio signals in our dataset are not only well integrated, but also informative and complementary, establishing Valor32k-AVQA v2.0 as a valuable resource for developing and evaluating robust audio-visual question answering systems. Ines Riahi, Abduljalil Radman, Zixin Guo, Rachid Hedjam, Jorma Laaksonen |
ACM Multimedia | 5 |
| 2025 | MIRA: A Novel Framework for Fusing Modalities in Medical RAGabstractPublisher Copyright: © 2025 Copyright held by the owner/author(s). Tajamul Ashraf, Zongyan Han, Jorma Laaksonen, Rao Muhammad Anwer |
ACM Multimedia | 4 |
| 2025 | Video Instance Segmentation in an Open-World
Omkar Thawakar, Sanath Narayan, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan 0001, Jorma Laaksonen, Mubarak Shah, Fahad Shahbaz Khan |
Int. J. Comput. Vis. | 6 |
| 2025 | Hybrid Deep Learning for Hyperspectral Single-Image Super-ResolutionabstractHyperspectral single image super-resolution (SISR) remains a challenging task due to the difficulty of restoring fine spatial details and preserving spectral fidelity across a wide range of wavelengths, which inherently limits the performance of conventional deep learning models. To effectively address this challenge, we introduce a novel module called Spectral-Spatial Unmixing Fusion (SSUF), which can be seamlessly integrated into existing 2D convolutional architectures to enhance both spatial resolution and spectral integrity. Specifically, the SSUF combines spectral unmixing and spectral–spatial feature extraction to subsequently guide a ResNet‑based convolutional neural network. Additionally, we employ a custom Spatial-Spectral Gradient Loss function, which integrates Mean Squared Error (MSE) with spatial and spectral gradient components, encouraging the model to accurately reconstruct features across both spatial and spectral dimensions. Experiments on three public remote sensing hyperspectral datasets demonstrate that our proposed hybrid deep learning (HDL) achieves competitive performance while reducing model complexity. The source codes are publicly available at: https://github.com/Usman1021/hsi-super-resolution. Jorma Laaksonen |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Prompt-based Weakly-supervised Vision-language Pre-trainingabstractWeakly-supervised Vision-Language Pre-training (W-VLP) explores methods leveraging weak cross-modal supervision, typically relying on object tags generated by a pre-trained object detector (OD) from images. However, training such an OD necessitates dense cross-modal information, including images paired with numerous object-level annotations. To alleviate that requirement, this paper addresses W-VLP in two stages: (1) creating data with weaker cross-modal supervision and (2) pre-training a vision-language (VL) model with the created data. The data creation process involves collecting knowledge from large language models (LLMs) to describe images. Given a category label of an image, its descriptions generated by an LLM are used as the language counterpart. This knowledge supplements what can be obtained using an OD, such as spatial relationships among objects most likely appearing in a scene. To mitigate the noise in the LLM-generated descriptions that destabilizes the training process and may lead to overfitting, we incorporate knowledge distillation and external retrieval-augmented knowledge during pre-training. Furthermore, we present an effective VL model pre-trained with the created data. Empirically, despite its weaker cross-modal supervision, our pre-trained VL model notably outperforms other W-VLP works in image and text retrieval tasks, e.g., VLMixer by 17.7% on MSCOCO and RELIT by 11.25% on Flickr30K relatively in Recall@1 in text-to-image retrieval task. It also shows superior performance on other VL downstream tasks, making a big stride towards matching the performances of strongly supervised VLP models. The results reveal the effectiveness of the proposed W-VLP methodology. Zixin Guo, Julius Wang, Selen Pehlivan, Abduljalil Radman, Min Cao 0005, Jorma Laaksonen |
Pattern Recognit. Lett. | 6 |
| 2024 | Diffusion-Based Multimodal Video Captioning
Jaakko Kainulainen, Zixin Guo, Jorma Laaksonen |
ACCV (3) | 3 |
| 2024 | Text-to-Multimodal Retrieval with Bimodal Input Fusion in Shared Cross-Modal TransformerabstractThe rapid proliferation of multimedia content has necessitated the development of effective multimodal video retrieval systems. Multimodal video retrieval is a non-trivial task involving retrieval of relevant information across different modalities, such as text, audio, and visual. This work aims to improve multimodal retrieval by guiding the creation of a shared embedding space with task-specific contrastive loss functions. An important aspect of our work is to propose a model that learns retrieval cues for the textual query from multiple modalities both separately and jointly within a hierarchical architecture that can be flexibly extended and fine-tuned for any number of modalities. To this end, the loss functions and the architectural design of the model are developed with a strong focus on increasing the mutual information between the textual and cross-modal representations. The proposed approach is quantitatively evaluated on the MSR-VTT and YouCook2 text-to-video retrieval benchmark datasets. The results showcase that the approach not only holds its own against state-of-the-art methods, but also outperforms them in a number of scenarios, with a notable relative improvements from baseline in R@1, R@5 and R@10 metrics. Pranav Arora, Selen Pehlivan, Jorma Laaksonen |
LREC/COLING | 3 |
| 2024 | Size-Modulated Deformable Attention in Spatio-Temporal Video Grounding Pipelines
Hans Tiwari, Selen Pehlivan, Jorma Laaksonen |
ICPR (18) | 3 |
| 2024 | A Comparison of Hyperspectral Super-Resolution Techniques for Boreal Forest ImageryabstractDespite the widespread use of deep learning models for super-resolution image enhancement, their use for hyper-spectral imagery has not yet been researched thoroughly. This study reviews a number of recent hyperspectral image super-resolution techniques and explores also other single-image super-resolution methods. Our work targets to forestry images, highlighting the main methodologies, contributions, advantages, and limitations of the studied methods. The state-of-the-art methods are categorized into three distinct groups, those based on the Convolutional Neural Network (CNN), the Transformer, and the Generative Adversarial Network (GAN). Subsequently, the selected methods are compared in terms of six different performance measures on an airborne hyperspectral image dataset of a boreal forest. Our findings conclude that Transformer-based methods consistently outperform other current hyperspectral super-resolution techniques, while the GAN approach is the most promising one among the studied non-hyperspectral models. Yuvrajsinh Chudasama, Ville Mayra, Florent Guiotte, Jorma Laaksonen |
IGARSS | 5 |
| 2024 | Mesh Surface And Morphological Hierarchies For Individual Tree Detection And Segmentation From LiDAR DataabstractThis paper presents a novel and efficient individual tree detection and segmentation method for LiDAR point clouds. We rely on a surface model of the forest to find tree tops in the canopy. Efficient connected component filtering is used to filter the surface model, detect and segment individual trees by tuning a single physically interpretable parameter. We validate our method on a genuine LiDAR point cloud and tree inventory dataset and show on-par results with a recent state-of-the-art individual tree detection study. Our method is original because, unlike the previous methods based on connected components, we do not depend on an intermediate raster to carry out the morphological filtering. Instead, our method relies on a graph that directly connects the points of the LiDAR data. This original approach not only opens direct improvements for tree detection in surface models, but also provides a broader and more efficient way to process LiDAR point clouds beyond individual tree detection and segmentation. Florent Guiotte, Joel Kostensalo, Jorma Laaksonen |
IGARSS | 3 |
| 2024 | Domain Generalization via Ensemble Stacking for Face Presentation Attack DetectionabstractAbstract Face presentation attack detection (PAD) plays a pivotal role in securing face recognition systems against spoofing attacks. Although great progress has been made in designing face PAD methods, developing a model that can generalize well to unseen test domains remains a significant challenge. Moreover, due to the different types of spoofing attacks, creating a dataset with a sufficient number of samples for training deep neural networks is a laborious task. This work proposes a comprehensive solution that combines synthetic data generation and deep ensemble learning to enhance the generalization capabilities of face PAD. Specifically, synthetic data is generated by blending a static image with spatiotemporal-encoded images using alpha composition and video distillation. In this way, we simulate motion blur with varying alpha values, thereby generating diverse subsets of synthetic data that contribute to a more enriched training set. Furthermore, multiple base models are trained on each subset of synthetic data using stacked ensemble learning. This allows the models to learn complementary features and representations from different synthetic subsets. The meta-features generated by the base models are used as input for a new model called the meta-model. The latter combines the predictions from the base models, leveraging their complementary information to better handle unseen target domains and enhance overall performance. Experimental results from seven datasets—WMCA, CASIA-SURF, OULU-NPU, CASIA-MFSD, Replay-Attack, MSU-MFSD, and SiW-Mv2—highlight the potential to enhance presentation attack detection by using large-scale synthetic data and a stacking-based ensemble approach. Jorma Laaksonen, Djamila Romaissa Beddiar, Mourad Oussalah 0002 |
Int. J. Comput. Vis. | 2 |
| 2024 | AS-Net: active speaker detection using deep audio-visual attentionabstractAbstract Active Speaker Detection (ASD) aims at identifying the active speaker among multiple speakers in a video scene. Previous ASD models often seek audio and visual features from long video clips with a complex 3D Convolutional Neural Network (CNN) architecture. However, models based on 3D CNNs can generate discriminative spatial-temporal features, but this comes at the expense of computational complexity, and they frequently face challenges in detecting active speakers in short video clips. This work proposes the Active Speaker Network (AS-Net) model, a simple yet effective ASD method tailored for detecting active speakers in relatively short video clips without relying on 3D CNNs. Instead, it incorporates the Temporal Shift Module (TSM) into 2D CNNs, facilitating the extraction of dense temporal visual features without the need for additional computations. Moreover, self-attention and cross-attention schemes are introduced to enhance long-term temporal audio-visual synchronization, thereby improving ASD performance. Experimental results demonstrate that AS-Net outperforms state-of-the-art 2D CNN-based methods on the AVA-ActiveSpeaker dataset and remains competitive with the methods utilizing more complex architectures. Abduljalil Radman, Jorma Laaksonen |
Multim. Tools Appl. | 2 |
| 2024 | Temporal teacher with masked transformers for semi-supervised action proposal generationabstractAbstract By conditioning on unit-level predictions, anchor-free models for action proposal generation have displayed impressive capabilities, such as having a lightweight architecture. However, task performance depends significantly on the quality of data used in training, and most effective models have relied on human-annotated data. Semi-supervised learning, i.e., jointly training deep neural networks with a labeled dataset as well as an unlabeled dataset, has made significant progress recently. Existing works have either primarily focused on classification tasks, which may require less annotation effort, or considered anchor-based detection models. Inspired by recent advances in semi-supervised methods on anchor-free object detectors, we propose a teacher-student framework for a two-stage action detection pipeline, named Temporal Teacher with Masked Transformers (TTMT), to generate high-quality action proposals based on an anchor-free transformer model. Leveraging consistency learning as one self-training technique, the model jointly trains an anchor-free student model and a gradually progressing teacher counterpart in a mutually beneficial manner. As the core model, we design a Transformer-based anchor-free model to improve effectiveness for temporal evaluation. We integrate bi-directional masks and devise encoder-only Masked Transformers for sequences. Jointly training on boundary locations and various local snippet-based features, our model predicts via the proposed scoring function for generating proposal candidates. Experiments on the THUMOS14 and ActivityNet-1.3 benchmarks demonstrate the effectiveness of our model for temporal proposal generation task. Selen Pehlivan, Jorma Laaksonen |
Mach. Vis. Appl. | 2 |
| 2024 | Saliency-based video summarization for face anti-spoofing
Mourad Oussalah 0002, Jorma Laaksonen |
Pattern Recognit. Lett. | 3 |
| 2023 | Person Image Synthesis via Denoising Diffusion ModelabstractThe pose-guided person image generation task requires synthesizing photorealistic images of humans in arbitrary poses. The existing approaches use generative adversarial networks that do not necessarily maintain realistic textures or need dense correspondences that struggle to handle complex deformations and severe occlusions. In this work, we show how denoising diffusion models can be applied for high-fidelity person image synthesis with strong sample diversity and enhanced mode coverage of the learnt data distribution. Our proposed Person Image Diffusion Model (PIDM) disintegrates the complex transfer problem into a series of simpler forward-backward denoising steps. This helps in learning plausible source-to-target transformation trajectories that result in faithful textures and undistorted appearance details. We introduce a ‘texture diffusion module’ based on cross-attention to accurately model the correspondences between appearance and pose information available in source and target images. Further, we propose ‘disentangled classifier-free guidance’ to ensure close resemblance between the conditional inputs and the synthesized output in terms of both pose and appearance information. Our extensive results on two large-scale benchmarks and a user study demonstrate the photorealism of our proposed approach under challenging scenarios. We also show how our generated images can help in downstream tasks. Code is available at https://github.com/ankanbhunia/PIDM. Ankan Bhunia, Salman Khan 0001, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Mubarak Shah, Fahad Shahbaz Khan |
CVPR | 5 |
| 2023 | Anchor-Free Action Proposal Network with Uncertainty EstimationabstractProposal generation is a fundamental yet challenging task for two-stage temporal action detection pipelines. The task aims at predicting starting and ending boundaries of segments in realistic video sequences and action recognition methods cannot be directly applied to such videos due to their untrimmed nature. Most state-of-the-art models rely on temporal convolutional neural networks with pre-defined anchor segments. By eliminating anchors, we propose a lighter end-to-end trainable Anchor-Free Multiscale Transformer-based Generator (AMTG) model using local clues via video snippets. To improve effectiveness for temporal evaluation, we apply multiscale Transformer encoders to sequences with a bi-directional mask extension that simultaneously predicts boundary distances with uncertainties and various snippet-based local scores. Later, our model integrates local predictions to generate proposal candidates using the proposed scoring function. Experiments on the THUMOS14 and ActivityNet-1.3 benchmarks demonstrate the effectiveness of AMTG for the temporal proposal generation task. Selen Pehlivan, Jorma Laaksonen |
ICME | 2 |
| 2023 | Cross-Modulated Few-Shot Image Generation for Colorectal Tissue Classification
Amandeep Kumar, Ankan Bhunia, Sanath Narayan, Hisham Cholakkal, Rao Muhammad Anwer, Jorma Laaksonen, Fahad Shahbaz Khan |
MICCAI (3) | 6 |
| 2023 | 3D Mitochondria Instance Segmentation with Spatio-Temporal Transformers
Omkar Thawakar, Rao Muhammad Anwer, Jorma Laaksonen, Orly Reiner, Mubarak Shah, Fahad Shahbaz Khan |
MICCAI (8) | 3 |
| 2023 | PiTL: Cross-modal Retrieval with Weakly-supervised Vision-language Pre-training via PromptingabstractVision-language (VL) Pre-training (VLP) has shown to well generalize VL models over a wide range of VL downstream tasks, especially for cross-modal retrieval. However, it hinges on a huge amount of image-text pairs, which requires tedious and costly curation. On the contrary,weakly-supervised VLP (W-VLP) explores means with object tags generated by a pre-trained object detector (OD) from images. Yet, they still require paired information, i.e. images and object-level annotations, as supervision to train an OD. Zixin Guo, Julius Wang, Selen Pehlivan, Abduljalil Radman, Jorma Laaksonen |
SIGIR | 5 |
| 2023 | Learning by Hallucinating: Vision-Language Pre-training with Weak SupervisionabstractWeakly-supervised vision-language (V-L) pre-training (W-VLP) aims at learning cross-modal alignment with little or no paired data, such as aligned images and captions. Recent W-VLP methods, which pair visual features with object tags, help achieve performances comparable with some VLP models trained with aligned pairs in various V-L downstream tasks. This, however, is not the case in cross-modal retrieval (XMR). We argue that the learning of such a W-VLP model is curbed and biased by the object tags of limited semantics.We address the lack of paired V-L data for model supervision with a novel Visual Vocabulary based Feature Hallucinator (WFH), which is trained via weak supervision as a W-VLP model, not requiring images paired with captions. WFH generates visual hallucinations from texts, which are then paired with the originally unpaired texts, allowing more diverse interactions across modalities.Empirically, WFH consistently boosts the prior W-VLP works, e.g. U-VisualBERT (U-VB), over a variety of V-L tasks, i.e. XMR, Visual Question Answering, etc. Notably, benchmarked with recall@{1,5,10}, it consistently U-VB on image-to-text and improves text-to-image retrieval on two popular datasets Flickr30K and MSCOCO. Meanwhile, it gains by at least 14.5% in cross-dataset generalization tests on these XMR tasks. Moreover, in other V-L downstream tasks considered, our WFH models are on par with models trained with paired V-L data, revealing the utility of unpaired data. These results demonstrate greater generalization of the proposed W-VLP model with WFH. Julius Wang, Jorma Laaksonen, Tomas Langer, Heikki Arponen, Tom E. Bishop |
WACV | 2 |
| 2023 | Improved action proposals using fine-grained proposal features with recurrent attention modelsabstractRecent models for the temporal action proposal task show that local properties can be an alternative to the region proposal network (RPN) for generating good proposal candidates on untrimmed videos. In this study, we devise an RPN model with a new two-stage pipeline and a new joint scoring function for temporal proposals. The evaluation of local properties is integrated into our RPN model to search for the best proposal candidates that can be distinguished mainly in fine details of proposal regions. Our network models proposals in multiple scales using two recurrent neural network layers with attention mechanisms. We observe that joint training of the RPN with local clues and multi-scale modeling of proposals with recurrent attention mechanisms improve the performance of the proposal generation task. Our model yields state-of-the-art results on the THUMOS-14 and comparable results on the ActivityNet-1.3 datasets. Selen Pehlivan, Jorma Laaksonen |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | When to Laugh and How Hard? A Multimodal Approach to Detecting Humor and Its IntensityabstractPrerecorded laughter accompanying dialog in comedy TV shows encourages the audience to laugh by clearly marking humorous moments in the show. We present an approach for automatically detecting humor in the Friends TV show using multimodal data. Our model is capable of recognizing whether an utterance is humorous or not and assess the intensity of it. We use the prerecorded laughter in the show as annotation as it marks humor and the length of the audience’s laughter tells us how funny a given joke is. We evaluate the model on episodes the model has not been exposed to during the training phase. Our results show that the model is capable of correctly detecting whether an utterance is humorous 78% of the time and how long the audience’s laughter reaction should last with a mean absolute error of 600 milliseconds. Khalid Al-Najjar, Mika Hämäläinen, Jörg Tiedemann, Jorma Laaksonen, Mikko Kurimo |
COLING | 4 |
| 2022 | DoodleFormer: Creative Sketch Drawing with Transformers
Ankan Bhunia, Salman Khan 0001, Hisham Cholakkal, Rao Muhammad Anwer, Fahad Shahbaz Khan, Jorma Laaksonen, Michael Felsberg |
ECCV (17) | 6 |
| 2022 | Post-Attention Modulator for Dense Video CaptioningabstractDense video captioning (VC) aims at generating a paragraph-long description for events in video segments. Borrowing from the success in language modeling, Transformer-based models for VC have been shown effective also in modeling cross-domain video-text representations with cross-attention (Xatt). Despite Xatt’s effectiveness, the queries and outputs of attention, which are from different domains, tend to be weakly related. In this paper, we argue that the weak relatedness, or domain discrepancy, could impede a model from learning meaningful cross-domain representations. Hence, we propose a simple yet effective Post-Attention Modulator (PAM) that post-processes Xatt’s outputs to narrow the discrepancy. Specifically, PAM modulates and enhances the average similarity over Xatt’s queries and outputs. The modulated similarities are then utilized as a weighting basis to interpolate PAM’s outputs. In our experiments, PAM was applied to two strong VC baselines, VTransformer and MART, with two different video features on the well-known VC benchmark datasets ActivityNet Captions and YouCookII. According to the results, the proposed PAM brings consistent improvements in, e.g., CIDEr-D at most to 14.5%, as well as other metrics, BLEU and METEOR, considered. Zixin Guo, Julius Wang, Jorma Laaksonen |
ICPR | 3 |
| 2022 | Learning a Dynamic Cross-Modal Network for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection that enables continuous (day and night) localization of pedestrians has numerous applications. Existing approaches typically aggregate multispectral features by a simple element-wise operation. However, such a local feature aggregation scheme ignores the rich non-local contextual information. Further, we argue that a local tight correspondence across modalities is desired for multi-modal feature aggregation. To address these issues, we introduce a multispectral pedestrian detection framework that comprises a novel dynamic cross-modal network (DCMNet), which strives to adaptively utilize the local and non-local complementary information between multi-modal features. The proposed DCMNet consists of a local and a non-local feature aggregation module. The local module employs dynamically learned convolutions to capture local relevant information across modalities. On the other hand, the non-local module captures non-local cross-modal information by first projecting features from both modalities into the latent space and then obtaining dynamic latent feature nodes for feature aggregation. Comprehensive experiments are performed on two challenging benchmarks: KAIST and LLVIP. Experiments reveal the benefits of the proposed DCMNet, leading to consistently improved detection performance on diverse detection paradigms and backbones. When using the same backbone, our proposed detector achieves absolute gains of 1.74% and 1.90% over the baseline Cascade RCNN on the KAIST and LLVIP datasets. Jin Xie 0005, Rao Muhammad Anwer, Hisham Cholakkal, Jing Nie 0001, Jiale Cao, Jorma Laaksonen, Fahad Shahbaz Khan |
ACM Multimedia | 6 |
| 2022 | Understanding videos with face recognition: a complete pipeline and applications
Pasquale Lisena, Jorma Laaksonen, Raphaël Troncy |
Multim. Syst. | 2 |
| 2022 | TAIGA: A Novel Dataset for Multitask Learning of Continuous and Categorical Forest Variables From Hyperspectral ImageryabstractThe spectral and spatial resolutions of modern optical Earth observation data are continuously increasing. To fully utilize the data, integrate them with other information sources, and create applications relevant to real-world problems, extensive training data are required. We present TAIGA, an open dataset including continuous and categorical forestry data, accompanied by airborne hyperspectral imagery with a pixel size of 0.7 m. The dataset contains over 70 million labeled pixels belonging to more than 600 forest stands. To establish a baseline on TAIGA dataset for multitask learning, we trained and validated a convolutional neural network to simultaneously retrieve 13 forest variables. Due to the size of the imagery, the training and testing sets were independent, with strictly no overlap for patches up to$45\times 45$pixels. Our retrieval results show that including both spectral and textural information improves the accuracy of mapping key boreal forest structural characteristics, compared with an earlier study including only spectral information from the same image. TAIGA responds to the increased availability of hyperspectral and very high resolution imagery, and includes the forestry variables relevant for forestry and environmental applications. We propose the dataset as a new benchmark for spatial–spectral methods that overcomes the limitations of widely used small-scale hyperspectral datasets. Matti Mottus, Phu Pham, Eelis Halme, Matthieu Molinier, Hai Cu, Jorma Laaksonen |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2021 | Patch Size Selection for Analysis of Sub-Meter Resolution Hyperspectral Imagery of ForestsabstractVery high resolution remote sensing data of forests, where individual tree crowns are separable, contains structural information on tree size and density. Such information is complementary to the spectral signatures currently used in forestry applications. Advanced machine learning methods, e.g. convolutional neural networks (CNNs), offer an automated and standardized way of retrieving both spectral and structural information from imagery. A key characteristic in CNNs is patch size, which should be large enough to include dominant structural scale, yet as small as possible to avoid unnecessary averaging. Our results show that the patch should be larger than one tree, but increasing it excessively reduces retrieval accuracy. Furthermore, large patch sizes can cause loss of independence between training and validation data, leading to overestimating model performance. Matti Mottus, Matthieu Molinier, Eelis Halme, Hai Cu, Jorma Laaksonen |
IGARSS | 5 |
| 2021 | Compact Deep Color Features for Remote Sensing Scene ClassificationabstractAbstract Aerial scene classification is a challenging problem in understanding high-resolution remote sensing images. Most recent aerial scene classification approaches are based on Convolutional Neural Networks (CNNs). These CNN models are trained on a large amount of labeled data and the de facto practice is to use RGB patches as input to the networks. However, the importance of color within the deep learning framework is yet to be investigated for aerial scene classification. In this work, we investigate the fusion of several deep color models, trained using color representations, for aerial scene classification. We show that combining several deep color models significantly improves the recognition performance compared to using the RGB network alone. This improvement in classification performance is, however, achieved at the cost of a high-dimensional final image representation. We propose to use an information theoretic compression approach to counter this issue, leading to a compact deep color feature set without any significant loss in accuracy. Comprehensive experiments are performed on five remote sensing scene classification benchmarks: UC-Merced with 21 scene classes, WHU-RS19 with 19 scene types, RSSCN7 with 7 categories, AID with 30 aerial scene classes, and NWPU-RESISC45 with 45 categories. Our results clearly demonstrate that the fusion of deep color features always improves the overall classification performance compared to the standard RGB deep features. On the large-scale NWPU-RESISC45 dataset, our deep color features provide a significant absolute gain of 4.3% over the standard RGB deep features. Rao Muhammad Anwer, Fahad Shahbaz Khan, Jorma Laaksonen |
Neural Process. Lett. | 3 |
| 2020 | Tackling the Unannotated: Scene Graph Generation with Bias-Reduced Models
Julius Wang, Selen Pehlivan, Jorma Laaksonen |
BMVC | 3 |
| 2020 | AI4TV 2020: 2nd International Workshop on AI for Smart TV Content Production, Access and DeliveryabstractTechnological developments in comprehensive video understanding - detecting and identifying visual elements of a scene, combined with audio understanding (music, speech), as well as aligned with textual information such as captions, subtitles, etc. and background knowledge - have been undergoing a significant revolution during recent years. The workshop brings together experts from academia and industry in order to discuss the latest progress in artificial intelligence research in topics related to multimodal information analysis, and in particular, semantic analysis of video, audio, and textual information for smart digital TV content production, access and delivery. Raphaël Troncy, Jorma Laaksonen, Hamed Rezazadegan Tavakoli, Lyndon J. B. Nixon, Vasileios Mezaris, Mohammad Hosseini 0002 |
ACM Multimedia | 2 |
| 2020 | Film Mood and Its Quantitative Determinants in Different Types of ScenesabstractFilms elicit emotions in viewers by infusing the story they tell with an affective character or tone-in a word, a mood. Considerable effort has been made recently to develop computational methods to estimate affective content in film. However, these efforts have focused almost exclusively on style-based features while neglecting to consider different scene types separately. In this study, we investigated the quantitative determinants of film mood across scenes classified by their setting and use of sounds. We examined whether viewers could assess film mood directly in terms of hedonic tone, energetic arousal, and tense arousal; whether their mood ratings differed by scene type; and how various narrative and stylistic film attributes as well as low- and high-level computational features related to the ratings. We found that the viewers were adept at assessing film mood, that sound-based scene classification brought out differences in the mood ratings, and that the low- and high-level features related to different mood dimensions. The study showed that computational film mood estimation can benefit from scene type classification and the use of both low- and high-level features. We have made our clip assessment and annotation data as well as the extracted computational features publicly available. Jussi Tarvainen, Jorma Laaksonen, Tapio Takala |
IEEE Trans. Affect. Comput. | 2 |
| 2019 | Multi-stream Convolutional Networks for Indoor Scene Recognition
Rao Muhammad Anwer, Fahad Shahbaz Khan, Jorma Laaksonen, Nazar Zaki |
CAIP (1) | 3 |
| 2019 | Deep Contextual Attention for Human-Object Interaction DetectionabstractHuman-object interaction detection is an important and relatively new class of visual relationship detection tasks, essential for deeper scene understanding. Most existing approaches decompose the problem into object localization and interaction recognition. Despite showing progress, these approaches only rely on the appearances of humans and objects and overlook the available context information, crucial for capturing subtle interactions between them. We propose a contextual attention framework for human-object interaction detection. Our approach leverages context by learning contextually-aware appearance features for human and object instances. The proposed attention module then adaptively selects relevant instance-centric context information to highlight image regions likely to contain human-object interactions. Experiments are performed on three benchmarks: V-COCO, HICO-DET and HCVRD. Our approach outperforms the state-of-the-art on all datasets. On the V-COCO dataset, our method achieves a relative gain of 4.4% in terms of role mean average precision (mAP role ), compared to the existing best approach. Tiancai Wang, Rao Muhammad Anwer, Muhammad Haris Khan, Fahad Shahbaz Khan, Yanwei Pang, Ling Shao 0001, Jorma Laaksonen |
ICCV | 7 |
| 2019 | A Multi-Task Bayesian Deep Neural Net for Detecting Life-Threatening Infant Incidents From Head ImagesabstractThe notorious incident of sudden infant death syndrome (SIDS) can easily happen to a newborn due to many environmental factors. To prevent such tragic incidents from happening, we propose a multi-task deep learning framework that detects different facial traits and two life-threatening indicators, i.e. which facial parts are occluded or covered, by analyzing the infant head image. Furthermore, we extend and adapt the recently developed models that capture data-dependent uncertainty from noisy observations for our application. The experimental results show significant improvements on YunInfants dataset across most of the tasks over the models that simply adopt the regular cross-entropy losses without addressing the effect of the underlying uncertainties. Julius Wang, Jorma Laaksonen, Yi-Ping Liao, Bo-Zong Wu, Shih-Yun Shen |
ICIP | 2 |
| 2019 | AI4TV 2019: 1st International Workshop on AI for Smart TV Content Production, Access and DeliveryabstractTechnological developments in comprehensive video understanding - detecting and identifying visual elements of a scene, combined with audio understanding (music, speech), as well as aligned with textual information such as captions, subtitles, etc. and background knowledge - have been undergoing a significant revolution during recent years. The workshop brings together experts from academia and industry in order to discuss the latest progress in artificial intelligence research in topics related to multimodal information analysis, and in particular, semantic analysis of video, audio, and textual information for smart digital TV content production, access and delivery. Raphaël Troncy, Jorma Laaksonen, Hamed Rezazadegan Tavakoli, Lyndon J. B. Nixon, Vasileios Mezaris |
ACM Multimedia | 2 |
| 2018 | Scale coding bag of deep features for human attribute and action recognitionabstractMost approaches to human attribute and action recognition in still images are based on image representation in which multi-scale local features are pooled across scale into a single, scale-invariant encoding. Both in bag-of-words and the recently popular representations based on convolutional neural networks, local features are computed at multiple scales. However, these multi-scale convolutional features are pooled into a single scale-invariant representation. We argue that entirely scale-invariant image representations are sub-optimal and investigate approaches to scale coding within a bag of deep features framework. Our approach encodes multi-scale information explicitly during the image encoding stage. We propose two strategies to encode multi-scale information explicitly in the final image representation. We validate our two scale coding techniques on five datasets: Willow, PASCAL VOC 2010, PASCAL VOC 2012, Stanford-40 and Human Attributes (HAT-27). On all datasets, the proposed scale coding approaches outperform both the scale-invariant method and the standard deep features of the same network. Further, combining our scale coding approaches with standard deep features leads to consistent improvement over the state of the art. Fahad Shahbaz Khan, Joost van de Weijer 0001, Rao Muhammad Anwer, Andrew D. Bagdanov, Michael Felsberg, Jorma Laaksonen |
Mach. Vis. Appl. | 6 |
| 2017 | Saliency Revisited: Analysis of Mouse Movements Versus FixationsabstractThis paper revisits visual saliency prediction by evaluating the recent advancements in this field such as crowd-sourced mouse tracking-based databases and contextual annotations. We pursue a critical and quantitative approach towards some of the new challenges including the quality of mouse tracking versus eye tracking for model training and evaluation. We extend quantitative evaluation of models in order to incorporate contextual information by proposing an evaluation methodology that allows accounting for contextual factors such as text, faces, and object attributes. The proposed contextual evaluation scheme facilitates detailed analysis of models and helps identify their pros and cons. Through several experiments, we find that (1) mouse tracking data has lower inter-participant visual congruency and higher dispersion, compared to the eye tracking data, (2) mouse tracking data does not totally agree with eye tracking in general and in terms of different contextual regions in specific, and (3) mouse tracking data leads to acceptable results in training current existing models, and (4) mouse tracking data is less reliable for model selection and evaluation. The contextual evaluation also reveals that, among the studied models, there is no single model that performs best on all the tested annotations. Hamed Rezazadegan Tavakoli, Fawad Ahmed, Ali Borji, Jorma Laaksonen |
CVPR | 4 |
| 2017 | VisualLabel: An Integrated Multimedia Content Management and Access FrameworkabstractWith the rapid growth of image and video data as well as the fast spread of user-generated content in social media and cloud services, it has become increasingly difficult for users to have efficient access and effective management of their digital content. In this paper we present a novel integrated open source multimedia content management and access framework, called VisualLabel, that enables smart photo services based on automated visual content analysis, annotation, search and retrieval using state of the art analysis back ends for services such as Facebook and Flickr. This paper includes detailed descriptions of the high-level architecture used in the VisualLabel framework and proof-of-concept implementations of a front-end service, along with three analysis back ends and a web client, all of which demonstrate the basic functionality provided by the framework. Iftikhar Ahmad 0001, Petri Rantanen, Pekka Sillberg, Jorma Laaksonen, Thomas Forss, Aqdas Malik, Marko Nieminen, Rakshith Shetty, Satoru Ishikawa, Jarno Kallio, Jukka Saarinen, Moncef Gabbouj, Jari Soini |
EJC | 4 |
| 2017 | Paying Attention to Descriptions Generated by Image Captioning ModelsabstractTo bridge the gap between humans and machines in image understanding and describing, we need further insight into how people describe a perceived scene. In this paper, we study the agreement between bottom-up saliency-based visual attention and object referrals in scene description constructs. We investigate the properties of human-written descriptions and machine-generated ones. We then propose a saliency-boosted image captioning model in order to investigate benefits from low-level cues in language models. We learn that (1) humans mention more salient objects earlier than less salient ones in their descriptions, (2) the better a captioning model performs, the better attention agreement it has with human descriptions, (3) the proposed saliencyboosted model, compared to its baseline form, does not improve significantly on the MS COCO database, indicating explicit bottom-up boosting does not help when the task is well learnt and tuned on a data, (4) a better generalization is, however, observed for the saliency-boosted model on unseen data. Hamed Rezazadegan Tavakoli, Rakshith Shetty, Ali Borji, Jorma Laaksonen |
ICCV | 4 |
| 2017 | Image pseudo tag generation with Deep Boltzmann machine anc topic-concept similarity mapabstractGeneral purpose search engines are used for searching not only plain text but also multimedia information. In multimodal search, it is common to use multiple queries to find the demanded information in the different media modalities. In most cases, however, it is hard to prepare such multimodal search queries. In addition, the semantic connection between the individual modalities is often weak or totally lacking in such multimodal search. Hence, single modality searching makes it hard to find the searched for information in the multimodal domain. In this paper we improve the Deep Boltzmann Machine applied to multimodal search by using GoogLeNet deep convolutional neural network and semantic concept features. We also propose a supervised method to produce a similarity map between hidden topics in text documents and the visual concepts in corresponding images, and an unsupervised method which uses the hidden topics in the documents as pseudo labels. The model can be used to infer and generate pseudo tags for untagged input query images in order to complement an image-only query to a multimodal one. The classification results with pseudo tag inputs show in our experiments improvement compared to the original tag inputs. Satoru Ishikawa, Jorma Laaksonen, Juha Karhunen |
IJCNN | 2 |
| 2017 | Computational and Perceptual Determinants of Film Mood in Different Types of ScenesabstractFilms seek to elicit emotions in viewers by infusing the story they tell with an affective character or tone - in a word, a mood. In content-based multimedia analysis, considerable effort has been made to develop methods to estimate film affect computationally. However, results have been hampered by a tendency to classify film scenes either by genre or not at all, while other potentially helpful classification methods have been neglected. In this study, we investigated the quantitative determinants of film mood across different types of scenes. We first collected style and mood ratings for 50 film scenes, which we classified by their location, time of day, and their use of dialogue and music. We then investigated whether the viewers rated the mood (in terms of hedonic tone, energetic arousal, and tense arousal) of various scene types differently, and how well perceptual stylistic attributes as well as low- and high-level computational features correlated with the mood ratings. We found that the mood ratings and their quantitative determinants differed across the scene types. We also found that the energetic arousal ratings were associated with the stylistic attributes and their corresponding low-level features, while hedonic tone and tense arousal were associated with high-level features related to the emotional expression in faces, dialogue, and music. The study contributes to ongoing efforts to estimate film affect computationally in showing that results can be improved by utilizing both low- and high-level features and by considering different scene types separately. Jussi Tarvainen, Jorma Laaksonen, Tapio Takala |
ISM | 2 |
| 2017 | TEX-Nets: Binary Patterns Encoded Convolutional Neural Networks for Texture RecognitionabstractRecognizing materials and textures in realistic imaging conditions is a challenging computer vision problem. For many years, local features based orderless representations were a dominant approach for texture recognition. Recently deep local features, extracted from the intermediate layers of a Convolutional Neural Network (CNN), are used as filter banks. These dense local descriptors from a deep model, when encoded with Fisher Vectors, have shown to provide excellent results for texture recognition. The CNN models, employed in such approaches, take RGB patches as input and train on a large amount of labeled images. We show that CNN models, which we call TEX-Nets, trained using mapped coded images with explicit texture information provide complementary information to the standard deep models trained on RGB patches. We further investigate two deep architectures, namely early and late fusion, to combine the texture and color information. Experiments on benchmark texture datasets clearly demonstrate that TEX-Nets provide complementary information to standard RGB deep network. Our approach provides a large gain of 4.8%, 3.5%, 2.6% and 4.1% respectively in accuracy on the DTD, KTH-TIPS-2a, KTH-TIPS-2b and Texture-10 datasets, compared to the standard RGB network of the same architecture. Further, our final combination leads to consistent improvements over the state-of-the-art on all four datasets. Rao Muhammad Anwer, Fahad Shahbaz Khan, Joost van de Weijer 0001, Jorma Laaksonen |
ICMR | 4 |
| 2017 | Exploiting inter-image similarity and ensemble of extreme learners for fixation prediction using deep features
Hamed Rezazadegan Tavakoli, Ali Borji, Jorma Laaksonen, Esa Rahtu |
Neurocomputing | 3 |
| 2017 | Uni- and multimodal methods for single- and multi-label recognition
Satoru Ishikawa, Jorma Laaksonen |
Multim. Tools Appl. | 2 |
| 2016 | Combining Holistic and Part-based Deep Representations for Computational Painting CategorizationabstractAutomatic analysis of visual art, such as paintings, is a challenging inter-disciplinary research problem. Conventional approaches only rely on global scene characteristics by encoding holistic information for computational painting categorization. We argue that such approaches are sub-optimal and that discriminative common visual structures provide complementary information for painting classification. Rao Muhammad Anwer, Fahad Shahbaz Khan, Joost van de Weijer 0001, Jorma Laaksonen |
ICMR | 4 |
| 2016 | Frame- and Segment-Level Features and Candidate Pool Evaluation for Video Caption GenerationabstractWe present our submission to the Microsoft Video to Language Challenge of generating short captions describing videos in the challenge dataset. Our model is based on the encoder--decoder pipeline, popular in image and video captioning systems. We propose to utilize two different kinds of video features, one to capture the video content in terms of objects and attributes, and the other to capture the motion and action information. Using these diverse features we train models specializing in two separate input sub-domains. We then train an evaluator model which is used to pick the best caption from the pool of candidates generated by these domain expert models. We argue that this approach is better suited for the current video captioning task, compared to using a single model, due to the diversity in the dataset. Rakshith Shetty, Jorma Laaksonen |
ACM Multimedia | 2 |
| 2015 | Compact color-texture description for texture classification
Fahad Shahbaz Khan, Rao Muhammad Anwer, Joost van de Weijer 0001, Michael Felsberg, Jorma Laaksonen |
Pattern Recognit. Lett. | 5 |
| 2014 | Unsupervised feature extraction for multimedia event detection and ranking using audio contentabstractIn this paper, we propose a new approach to classify and rank multimedia events based purely on audio content using video data from TRECVID-2013 multimedia event detection (MED) challenge. We perform several layers of nonlinear mappings to extract a set of unsupervised features from an initial set of temporal and spectral features to obtain a superior presentation of the atomic audio units. Additionally, we propose a novel weighted divergence measure for kernel based classifiers. The extensive set of experiments confirms that augmentation of the proposed steps results in an improved accuracy for most of the event classes. Ehsan Amid, Annamaria Mesaros, Kalle J. Palomäki, Jorma Laaksonen, Mikko Kurimo |
ICASSP | 4 |
| 2014 | Experiments on Recognising the Handshape in Blobs Extracted from Sign Language VideosabstractHandshape has an important role in sign languages. It would be inconceivable to try to understand sign language without recognising the handshapes. Over the years, numerous different approaches have been proposed for extracting the hand configuration information. The existing approaches for hand-shape recognition have problems especially with the huge sizes of modern linguistic corpora. Computationally expensive methods become easily infeasible with such large amounts of data. In this paper we examine the straightforward and efficient approach of recognising handshapes by our existing image category detection methodology, involving state-of-the-art local image descriptors. In the experiments the approach produces promising results. On the image feature side, we find that surprisingly complex hierarchical descriptors of shape primitive statistics provide the best overall performance in hand shape recognition. The accuracy of feature-wise detections can be improved by fusing together several features. Considering the temporal succession of the hand blobs markedly improves the accuracy over detecting the hand shape in each video frame in isolation. Ville Viitaniemi, Matti Karppa, Jorma Laaksonen |
ICPR | 3 |
| 2014 | SLMotion - An extensible sign language oriented video analysis tool
Matti Karppa, Ville Viitaniemi, Marcos Luzardo, Jorma Laaksonen, Tommi Jantunen |
LREC | 4 |
| 2014 | S-pot - a benchmark in spotting signs within continuous signing
Ville Viitaniemi, Tommi Jantunen, Leena Savolainen, Matti Karppa, Jorma Laaksonen |
LREC | 5 |
| 2014 | Convolutional Network Features for Scene RecognitionabstractConvolutional neural networks have recently been used to obtain record-breaking results in many vision benchmarks. In addition, the intermediate layer activations of a trained network when exposed to new data sources have been shown to perform very well as generic image features, even when there are substantial differences between the original training data of the network and the new domain. In this paper, we focus on scene recognition and show that convolutional networks trained on mostly object recognition data can successfully be used for feature extraction in this task as well. We train a total of four networks with different training data and architectures, and show that the proposed method combining multiple scales and multiple features obtains state-of-the-art performance on four standard scene datasets. Markus Koskela, Jorma Laaksonen |
ACM Multimedia | 2 |
| 2014 | Content-Based Prediction of Movie Style, Aesthetics, and Affect: Data Set and Baseline ExperimentsabstractThe affective content of a movie is often considered to be largely determined by its style and aesthetics. Recently, studies have attempted to estimate affective movie content with computational features, but results have been mixed, one of the main reasons being a lack of data on perceptual stylistic and aesthetic attributes of film, which would provide a ground truth for the features. The distinctions between energetic and tense arousal as well as perceived and felt affect are also often neglected. In this study, we present a data set of ratings by 73 viewers of 83 stylistic, aesthetic, and affective attributes for a selection of movie clips containing complete scenes taken from mainstream movies. The affective attributes include the temporal progression of perceived and felt valence and arousal within the clips. The data set is aimed to be used to train algorithms that predict viewer assessments based on low-level computational features. With this data set, we performed a baseline study modeling the relation between a large selection of low-level computational features (i.e., visual, auditory, and temporal) and perceptual stylistic, aesthetic, and affective attributes of movie clips. Two algorithms were compared in a realistic prediction scenario: linear regression and the neural-network-based Extreme Learning Machine (ELM). Felt and perceived affect as well as stylistic attributes were shown to be equally easy to predict, whereas the prediction of aesthetic attributes failed. The performance of the ELM predictor was overall found to be slightly better than the linear regression. A feature selection experiment illustrated that features from all low-level computational modalities, visual, auditory and temporal, contribute to the prediction of the affect assessments. We have made our assessment data and extracted computational features publicly available. Jussi Tarvainen, Mats Sjöberg, Stina Westman, Jorma Laaksonen, Pirkko Oittinen |
IEEE Trans. Multim. | 4 |
| 2013 | Adaptive timeline interface to personal history dataabstractAs the growth of stored personal digital information, such as photographs and emails, is continuously increasing, new tools for browsing and searching are needed. We introduce an intelligent mobile information access tool for personal data. The data are presented in an adaptive timeline where the displayed items function as search cues. The novelty is that the visualization is dynamically changed to emphasize relevant items, which makes them easier to recognize and select. The relevance is inferred during usage of the system from user feedback. In a user study, the dynamic timeline-based interface on a mobile device was shown to require less effort than conventional textual search. Antti Ajanki, Markus Koskela, Jorma Laaksonen, Samuel Kaski |
ICMI | 3 |
| 2013 | Affective Abstract Image Classification and Retrieval Using Multiple Kernel Learning
He Zhang 0009, Zhirong Yang, Mehmet Gönen, Markus Koskela, Jorma Laaksonen, Timo Honkela, Erkki Oja |
ICONIP (3) | 5 |
| 2013 | Large-scale visual concept detection with explicit kernel maps and power mean SVMabstractMany emerging application areas in video and image processing require large-scale visual concept detection. Examples include content-based indexing of online user-generated videos and 24/7 archival of TV broadcasts. The current state of the art in concept detection uses bag-of-visual-words features with computationally heavy exponential kernel classifiers. We argue that this classifier approach is not feasible for large-scale real-time applications, and propose instead to use combinations of approximate additive kernel classifiers. By using explicit kernel maps and the power mean SVM, followed by fusion of classifiers trained on different features, we achieve high retrieval precision while retaining real-time performance for large sets of concepts. This paper presents a series of experiments with the large-scale TRECVID 2012 video database and the commonly used Fifteen Scene Categories image database. We show significantly improved retrieval performance over standard linear classifiers, and by late fusion over several visual features, the approximative additive kernels outperform any single exponential kernel in only a fraction of the detection time. Mats Sjöberg, Markus Koskela, Satoru Ishikawa, Jorma Laaksonen |
ICMR | 4 |
| 2012 | Real-time large-scale visual concept detection with linear classifiers
Mats Sjöberg, Markus Koskela, Satoru Ishikawa, Jorma Laaksonen |
ICPR | 4 |
| 2012 | Comparing computer vision analysis of signed language video with motion capture recordings
Matti Karppa, Tommi Jantunen, Ville Viitaniemi, Jorma Laaksonen, Birgitta Burger, Danny De Weerdt |
LREC | 4 |
| 2011 | Gaze- and Speech-Enhanced Content-Based Image Retrieval in Image Tagging
He Zhang 0009, Teemu Ruokolainen, Jorma Laaksonen, Christina Hochleitner, Rudolf Traunmüller |
ICANN (2) | 3 |
| 2011 | A Multimodal Information Collector for Content-Based Image Retrieval System
He Zhang 0009, Mats Sjöberg, Jorma Laaksonen, Erkki Oja |
ICONIP (3) | 3 |
| 2011 | Analyzing Emotional Semantics of Abstract Art Using Low-Level Image Features
He Zhang 0009, Eimontas Augilius, Timo Honkela, Jorma Laaksonen, Hannes Gamper, Henok Alene |
IDA | 4 |
| 2010 | Region Matching Techniques for Spatial Bag of Visual Words Based Image Category Recognition
Ville Viitaniemi, Jorma Laaksonen |
ICANN (1) | 2 |
| 2009 | Image Theft Detection with Self-Organising Maps
Philip Prentis, Mats Sjöberg, Markus Koskela, Jorma Laaksonen |
ICANN (1) | 4 |
| 2009 | Representing Images with chi2 Distance Based Histograms of SIFT Descriptors
Ville Viitaniemi, Jorma Laaksonen |
ICANN (2) | 2 |
| 2008 | Image Classification by Histogram Features Created with Learning Vector Quantization
Marcin Blachnik, Jorma Laaksonen |
ICANN (1) | 2 |
| 2008 | Classification of Fundus Images for Diagnosing Glaucoma by Self-Organizing Map and Learning Vector Quantization
Nobuo Matsuda, Jorma Laaksonen, Fumiaki Tajima, Hideaki Sato |
ICONIP (2) | 2 |
| 2008 | Inferring semantics from textual information in multimedia retrieval
Mats Sjöberg, Jorma Laaksonen, Timo Honkela, Matti Pöllä |
Neurocomputing | 2 |
| 2008 | Principal whitened gradient for information geometry
Zhirong Yang, Jorma Laaksonen |
Neural Networks | 2 |
| 2007 | How Marginal Likelihood Inference Unifies Entropy, Correlation and SNR-Based Stopping in Nonlinear Diffusion Scale-Spaces
Ramunas Girdziusas, Jorma Laaksonen |
ACCV (1) | 2 |
| 2007 | Face Recognition Using Parzenfaces
Zhirong Yang, Jorma Laaksonen |
ICANN (2) | 2 |
| 2007 | When is a Discrete Diffusion a Scale-Space?abstractNecessary and sufficient conditions are discussed which state when the Euler-inspired forward diffusion in a discrete space-time is a scale-space in the sense of both the total and sign variation diminishing. We emphasize that the problem is algebraic and reduces to characterization of the elements of the generalized Laplacian so that the diffusion propagators are positive definite. As a key-product, explicit analytical expressions are found for the principal minors of the frequently-applied class of tridiagonal (Jacobi) matrices. Further generalizations are outlined by introducing novel techniques of evaluating matrix determinants. Ramunas Girdziusas, Jorma Laaksonen |
ICCV | 2 |
| 2007 | Detecting changes in polarimetric SAR data with content-based image retrievalabstractIn this study, we extended the potential of a Content- Based Image Retrieval (CBIR) system based on Self-Organizing Maps (SOMs), for the analysis of remote sensing data. A database was artificially created by splitting each image to be analyzed into small images (orimagelets). Content-based image retrieval was applied to fully polarimetric airborne SAR data, using a selection of polarimetric features. After training the system on this imagelet database, automatic queries could detect changes. Results were encouraging on airborne SAR data and may be more useful for spaceborne polarimetric data. Matthieu Molinier, Jorma Laaksonen, Yrjö Rauste, Tuomas Häme |
IGARSS | 2 |
| 2007 | Approximated Geodesic Updates with Principal Natural GradientsabstractWe propose a novel optimization algorithm which overcomes two drawbacks of Amari's natural gradient updates for information geometry. First, prewhitening the tangent vectors locally converts a Riemannian manifold to an Euclidean space so that the additive parameter update sequence approximates geodesics. Second, we prove that dimensionality reduction of natural gradients is necessary for learning multidimensional linear transformations. Removal of minor components also leads to noise reduction and better computational efficiency. The proposed method demonstrates faster and more robust convergence in the simulations on recovering a Gaussian mixture of artificial data and on discriminative learning of ionosphere data. Zhirong Yang, Jorma Laaksonen |
IJCNN | 2 |
| 2007 | Multiplicative updates for non-negative projections
Zhirong Yang, Jorma Laaksonen |
Neurocomputing | 2 |
| 2007 | Projective Non-Negative Matrix Factorization with Applications to Facial Image ProcessingabstractWe propose a new variant of Non-negative Matrix Factorization (NMF), including its model and two optimization rules. Our method is based on positively constrained projections and is related to the conventional SVD or PCA decomposition. The new model can potentially be applied to image compression and feature extraction problems. Of the latter, we consider processing of facial images, where each image consists of several parts and for each part the observations with different lighting mainly distribute along a straight line through the origin. No regularization terms are required in the objective functions and both suggested optimization rules can easily be implemented by matrix manipulations. The experiments show that the derived base vectors are spatially more localized than those of NMF. In turn, the better part-based representations improve the recognition rate of semantic classes such as the gender or existence of mustache in the facial images. Zhirong Yang, Zhijian Yuan, Jorma Laaksonen |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2007 | Adaptive combination of adaptive classifiers for handwritten character recognition
Matti Aksela, Jorma Laaksonen |
Pattern Recognit. Lett. | 2 |
| 2007 | Evaluating the performance in automatic image annotation: Example case by adaptive fusion of global image features
Ville Viitaniemi, Jorma Laaksonen |
Signal Process. Image Commun. | 2 |
| 2007 | Detecting Man-Made Structures and Changes in Satellite Imagery With a Content-Based Information Retrieval System Built on Self-Organizing MapsabstractThe increasing amount and resolution of satellite sensors demand new techniques for browsing remote sensing image archives. Content-based querying allows an efficient retrieval of images based on the information they contain, rather than their acquisition date or geographical extent. Self-organizing maps (SOMs) have been successfully applied in the PicSOM system to content-based image retrieval in databases of conventional images. In this paper, we investigate and extend the potential of PicSOM for the analysis of remote sensing data. We propose methods for detecting man-made structures, as well as supervised and unsupervised change detection, based on the same framework. In this paper, a database was artificially created by splitting each satellite image to be analyzed into small images. After training the PicSOM on this imagelet database, both interactive and off-line queries were made to detect man-made structures, as well as changes between two very high resolution images from different years. Experimental results were both evaluated quantitatively and discussed qualitatively, and suggest that this new approach is suitable for analyzing very high resolution optical satellite imagery. Possible applications of this work include interactive detection of man-made structures or supervised monitoring of sensitive sites Matthieu Molinier, Jorma Laaksonen, Tuomas Häme |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2007 | Measuring Concept Similarities in Multimedia Ontologies: Analysis and EvaluationsabstractThe recent development of large-scale multimedia concept ontologies has provided a new momentum for research in the semantic analysis of multimedia repositories. Different methods for generic concept detection have been extensively studied, but the question of how to exploit the structure of a multimedia ontology and existing inter-concept relations has not received similar attention. In this paper, we present a clustering-based method for modeling semantic concepts on low-level feature spaces and study the evaluation of the quality of such models with entropy-based methods. We cover a variety of methods for assessing the similarity of different concepts in a multimedia ontology. We study three ontologies and apply the proposed techniques in experiments involving the visual and semantic similarities, manual annotation of video, and concept detection. The results show that modeling inter-concept relations can provide a promising resource for many different application areas in semantic multimedia processing. Markus Koskela, Alan F. Smeaton, Jorma Laaksonen |
IEEE Trans. Multim. | 3 |
| 2006 | Retrieval of Multimedia Objects by Combining Semantic Information from Visual and Textual Descriptors
Mats Sjöberg, Jorma Laaksonen, Matti Pöllä, Timo Honkela |
ICANN (2) | 2 |
| 2006 | Techniques for Still Image Scene Classification and Object Detection
Ville Viitaniemi, Jorma Laaksonen |
ICANN (2) | 2 |
| 2006 | A Fast Fixed-Point Algorithm for Two-Class Discriminative Feature Extraction
Zhirong Yang, Jorma Laaksonen |
ICANN (2) | 2 |
| 2006 | A Self-Organizing Map Framework for Detection of Man-Made Structures and Changes in Satellite ImageryabstractContent-based querying allows efficient retrieval of images based on the information they contain, rather than acquisition date or geographical extent. We extend the potential of a content-based image retrieval (CBIR) system based on Self- Organizing Maps (SOMs), to the analysis of remote sensing data. A database was artificially created by splitting each satellite image to be analyzed into small images. After training the CBIR system on this imagelet database, both interactive and off-line queries were made to detect man-made structures, as well as changes. Experimental results suggest that this new approach is suitable for analyzing very high-resolution optical satellite imagery. Possible applications include interactive detection of man-made structures and supervised monitoring of sensitive sites. Matthieu Molinier, Jorma Laaksonen, Tuomas Häme |
IGARSS | 2 |
| 2006 | Using diversity of errors for selecting members of a committee classifier
Matti Aksela, Jorma Laaksonen |
Pattern Recognit. | 2 |
| 2005 | Optimal Stopping and Constraints for Diffusion Models of Signals with Discontinuities
Ramunas Girdziusas, Jorma Laaksonen |
ECML | 2 |
| 2005 | Jacobi Alternative to Bayesian Evidence Maximization in Diffusion Filtering
Ramunas Girdziusas, Jorma Laaksonen |
ICANN (2) | 2 |
| 2005 | Content-Based Retrieval of Web Pages and Other Hierarchical Objects with Self-organizing Maps
Mats Sjöberg, Jorma Laaksonen |
ICANN (2) | 2 |
| 2005 | Gaussian processes of nonlinear diffusion filteringabstractNonlinear diffusion filtering can be improved if viewed as Bayesian Gaussian process regression. We relate the covariance functions of the diffusion process outcome to the spatial diffusion operator and show how Bayesian evidence criterion can he utilized to determine the parameters of the nonlinear diffusivity and the optimal diffusion stopping time. Computational example is given where the nonlinear diffusion filtering outperforms typical Gaussian process regression. Ramunas Girdziusas, Jorma Laaksonen |
IJCNN | 2 |
| 2004 | Gaussian Process Regression with Fluid Hyperpriors
Ramunas Girdziusas, Jorma Laaksonen |
ICONIP | 2 |
| 2004 | Class distributions on SOM surfaces for feature extraction and object retrieval
Jorma Laaksonen, Markus Koskela, Erkki Oja |
Neural Networks | 1 |
| 2003 | Methods for adaptive combination of classifiers with application to recognition of handwritten characters
Matti Aksela, Ramunas Girdziusas, Jorma Laaksonen, Erkki Oja, Jari Kangas 0001 |
Int. J. Document Anal. Recognit. | 3 |
| 2002 | Implementing Relevance Feedback as Convolutions of Local Neighborhoods on Self-Organizing Maps
Markus Koskela, Jorma Laaksonen, Erkki Oja |
ICANN | 2 |
| 2002 | Influence of erroneous learning samples on adaptation in on-line handwriting recognition
Vuokko Vuori, Jorma Laaksonen, Jari Kangas 0001 |
Pattern Recognit. | 2 |
| 2002 | PicSOM-self-organizing image retrieval with MPEG-7 content descriptorsabstractDevelopment of content-based image retrieval (CBIR) techniques has suffered from the lack of standardized ways for describing visual image content. Luckily, the MPEG-7 international standard is now emerging as both a general framework for content description and a collection of specific agreed-upon content descriptors. We have developed a neural, self-organizing technique for CBIR. Our system is named PicSOM and it is based on pictorial examples and relevance feedback (RF). The name stems from "picture" and the self-organizing map (SOM). The PicSOM system is implemented by using tree structured SOMs. In this paper, we apply the visual content descriptors provided by MPEG-7 in the PicSOM system and compare our own image indexing technique with a reference system based on vector quantization (VQ). The results of our experiments show that the MPEG-7-defined content descriptors can be used as such in the PicSOM system even though Euclidean distance calculation, inherently used in the PicSOM system, is not optimal for all of them. Also, the results indicate that the PicSOM technique is a bit slower than the reference system in starting to find relevant images. However, when the strong RF mechanism of PicSOM begins to function, its retrieval precision exceeds that of the reference system. Jorma Laaksonen, Markus Koskela, Erkki Oja |
IEEE Trans. Neural Networks | 1 |
| 2001 | Rejection Methods for an Adaptive Committee ClassifierabstractAdaptation is an effective method for improving classification accuracy and a committee structure can in general improve on its members' performance. Therefore an adaptive committee structure is a tempting approach. Rejection may be used in handwriting recognition to improve performance through either directing the problematic character to a special classifier that handles such hard cases or discarding it. The experiments in this paper compare several fundamentally different approaches to implementing rejection in an adaptive committee classifier. A dynamically expanding context (DEC) - based committee is used for evaluating these approaches. The results show that if the rejected classes are handled with a 50% error rate, the performance is improved. A scheme in which there is an adjustable threshold for distance-based rejection is an effective method for implementing rejection in this setting. Matti Aksela, Jorma Laaksonen, Erkki Oja, Jari Kangas 0001 |
ICDAR | 2 |
| 2001 | Speeding Up On-line Recognition of Handwritten Characters by Pruning the Prototype SetabstractThis work describes a prototype-based online handwritten character recognition system and a two-phase recognition scheme aimed to speed up the recognition. In the first phase, the prototype set is pruned and ordered on the basis of preclassification performed with heavily down-sampled characters and prototypes. In the second phase, the final classification is performed without down-sampling by using the reduced set of prototypes. Two down-sampling methods, a linear and nonlinear one, have been analyzed to see their properties regarding the recognition time and accuracy. Vuokko Vuori, Jorma Laaksonen, Erkki Oja, Jari Kangas 0001 |
ICDAR | 2 |
| 2001 | Experiments with adaptation strategies for a prototype-based recognition system for isolated handwritten characters
Vuokko Vuori, Jorma Laaksonen, Erkki Oja, Jari Kangas 0001 |
Int. J. Document Anal. Recognit. | 2 |
| 2001 | Self-Organising Maps as a Relevance Feedback Technique in Content-Based Image Retrieval
Jorma Laaksonen, Markus Koskela, Sami Laakso, Erkki Oja |
Pattern Anal. Appl. | 1 |
| 2000 | Statistical Shape Features in Content-Based Image RetrievalabstractIn this article the use of shape features in content-based image retrieval is studied. The emphasis is on techniques which do not demand object segmentation. PicSOM, the image retrieval system used in the experiments, requires that features are represented by constant-sized feature vectors for which the Euclidean distance can be used as a similarity measure. The shape features suggested here are edge histograms and Fourier transform based features computed for an edge image in Cartesian and polar coordinate planes. The results show that both local and global shape features are important clues of shapes in an image. Sami S. Brandt, Jorma Laaksonen, Erkki Oja |
ICPR | 2 |
| 2000 | Controlling On-Line Adaptation of a Prototype-Based Classifier for Handwritten CharactersabstractMethods for controlling the adaptation process of an online handwritten character recognizer are studied. The classifier is based on the k-nearest neighbor rule and it is adapted to a new writing style by adding new prototypes, deactivating confusing prototypes, and reshaping existing prototypes in a self-supervised fashion. The dissimilarity measure used for the comparison of characters is a nonlinear curve matching method base on dynamic time warping algorithm. Time needed for the evaluation of the dissimilarity measure for a single character depends linearly on the size of the prototype set. The purpose of the control methods is to increase the classifier's tolerance to malformed or mislabelled learning samples and to limit the growth of the prototype set. The control methods either set an upper limit for the number of prototypes per class or switch the adaptation of a particular character class on or off depending on the earlier performance of the classifier. Vuokko Vuori, Jorma Laaksonen, Erkki Oja, Jari Kangas 0001 |
ICPR | 2 |
| 2000 | PicSOM - content-based image retrieval with self-organizing maps
Jorma Laaksonen, Markus Koskela, Sami Laakso, Erkki Oja |
Pattern Recognit. Lett. | 1 |
| 1999 | Dynamically Expanding Context as Committee Adaptation Method in On-Line Recognition of Handwritten Latin CharactersabstractWe have developed an adaptive handwriting recognizer for isolated Latin characters in which the adaptive behavior is based on the dynamically expanding context (DEC) algorithm. In our current system, the outputs of a set of static classifiers are combined in a committee machine, whose rules are adapted. Every misclassified character gives rise to adding a new DEC rule to the rule set of the committee. When the existing rules fail to produce a correct recognition output, more and more context information is utilized in forming the new DEC rules. Not only the first-ranking outputs from the member classifiers but also the second-ranking ones can be taken into account when forming the DEC rules. In the experiments described in this paper, various options in the implementation of the DEC committee classifier are evaluated. The results of the experiments show that the system is capable of fast adaptation to the user's handwriting and lead to lowered recognition error rates. Jorma Laaksonen, Matti Aksela, Erkki Oja, Jari Kangas 0001 |
ICDAR | 1 |
| 1999 | On-line Adaptation in Recognition of Handwritten Alphanumeric CharactersabstractWe have developed an adaptive online recognizer that is suitable for recognizing isolated alphanumeric characters. It is based on the k nearest neighbor rule. Various dissimilarity measures, all based on dynamic time warping (DTW), have been studied. The main focus of this work is on online adaptation. The adaptation is performed by modifying the prototype set of the classifier according to its recognition performance and the user's writing style. These adaptations include: (1) adding new prototypes, (2) inactivating confusing prototypes, and (3) reshaping existing prototypes. The reshaping algorithm is based on learning vector quantization (LVQ). The writers are allowed to use their own natural style of writing, and the adaptation is carried out during normal use in a self-supervised fashion and thus remains otherwise unnoticed by the user. Vuokko Vuori, Jorma Laaksonen, Erkki Oja, Jari Kangas 0001 |
ICDAR | 2 |
| 1999 | Adaptive local subspace classifier in on-line recognition of handwritten charactersabstractSubsystems for online recognition of handwriting are needed in personal digital assistants (PDA) and other portable handheld devices. We have developed a recognition system which enhances its accuracy by applying continuous adaptation to the user's writing style. The forms of adaptation we have experimented with take place simultaneously with the normal operation of the system and therefore, there is no need for separate training period of the device. The present implementation uses dynamic time warping (DTW) in matching the input characters with the stored prototypes. The DTW algorithm implemented with dynamic programming (DP) is, however both time and memory consuming. In our current research we have experimented with methods that transform the elastic templates to pixel images which can then be recognized by using statistical or neural classification. The particular neural classifier we have used is the local subspace classifier (LSC) of which we have developed an adaptive version. Jorma Laaksonen, Matti Aksela, Erkki Oja, Jari Kangas 0001 |
IJCNN | 1 |
| 1999 | PicSOM: self-organizing maps for content-based image retrievalabstractContent-based image retrieval is an important approach to the problem of processing the increasing amount of visual data. It is based on automatically extracted features from the content of the images, such as color, texture, shape and structure. We have started a project to study methods for content-based image retrieval using the self-organizing map (SOM) as the image similarity scoring method. Our image retrieval system, named PicSOM, can be seen as a SOM-based approach to relevance feedback which is a form of supervised learning to adjust the subsequent queries based on the user's responses during the information retrieval session. In PicSOM, a separate tree structured SOM (TS-SOM) is trained for each feature vector type in use. The system then adapts to the user's preferences by returning her more images from those SOMs where her responses have been most densely mapped. Jorma Laaksonen, Markus Koskela, Erkki Oja |
IJCNN | 1 |
| 1998 | Learning Subspace Classifiers and Error-Corrective Feature ExtractionabstractSubspace methods are a powerful class of statistical pattern classification algorithms. The subspaces form semiparametric representations of the pattern classes in the form of principal components. In this sense, subspace classification methods are an application of classical optimal data compression techniques. Additionally, the subspace formalism can be given a neural network interpretation. There are learning versions of the subspace classification methods, in which error-driven learning procedures are applied to the subspaces in order to reduce the number of misclassified vectors. An algorithm for iterative selection of the subspace dimensions is presented in this paper. Likewise, a modified formula for calculating the projection lengths in the subspaces is investigated. The principle of adaptive learning in subspace methods can further be applied to feature extraction. In our work, we have studied two adaptive feature extraction schemes. The adaptation process is directed by errors occurring in the classifier. Unlike most traditional classifier models which take the preceding feature extraction stage as given, this scheme allows for reducing the loss of information in the feature extraction stage. The enhanced overall classification performance resulting from the added adaptivity is demonstrated with experiments in which recognition of handwritten digits has been used as an exemplary application. Jorma Laaksonen, Erkki Oja |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1997 | Local Subspace Classifier
Jorma Laaksonen |
ICANN | 1 |
| 1997 | Neural and statistical classifiers-taxonomy and two case studiesabstractPattern classification using neural networks and statistical methods is discussed. We give a tutorial overview in which popular classifiers are grouped into distinct categories according to their underlying mathematical principles; also, we assess what makes a classifier neural. The overview is complemented by two case studies using handwritten digit and phoneme data that test the performance of a number of most typical neural-network and statistical classifiers. Four methods of our own are included: reduced kernel discriminant analysis, the learning k-nearest neighbors classifier, the averaged learning subspace method, and a version of kernel discriminant analysis. Lasse Holmström, Petri Koistinen, Jorma Laaksonen, Erkki Oja |
IEEE Trans. Neural Networks | 3 |
| 1996 | Subspace Dimension Selection and Averaged Learning Subspace Method in Handwritten Digit Classification
Jorma Laaksonen, Erkki Oja |
ICANN | 1 |
| 1996 | Neural network and statistical perspectives of classificationabstractPattern classification using neural networks and statistical methods is discussed and a taxonomy based on their underlying mathematical principles is presented. Typical neural network and statistical classifiers are then compared in a case study using handwritten digit data. Lasse Holmström, Petri Koistinen, Jorma Laaksonen, Erkki Oja |
ICPR | 3 |
| 1991 | A new reliability-based phoneme segmentation method for the "neural" phonetic typewriter
Jorma Laaksonen |
EUROSPEECH | 1 |
| 1990 | Variants of self-organizing mapsabstractSelf-organizing maps have a bearing on traditional vector quantization. A characteristic that makes them more closely resemble certain biological brain maps, however, is the spatial order of their responses, which is formed in the learning process. A discussion is presented of the basic algorithms and two innovations: dynamic weighting of the input signals at each input of each cell, which improves the ordering when very different input signals are used, and definition of neighborhoods in the learning algorithm by the minimal spanning tree, which provides a far better and faster approximation of prominently structured density functions. It is cautioned that if the maps are used for pattern recognition and decision process, it is necessary to fine tune the reference vectors so that they directly define the decision borders. Jari Kangas 0001, Teuvo Kohonen, Jorma Laaksonen |
IEEE Trans. Neural Networks | 3 |