Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Tapio Seppänen

dblp:48/6023 · DBLP profile ↗
← Back
53ranked-venue papers
1as first author
10since 2021 · last 2025
0000-0002-3963-0750ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 since 2021Human-computer interaction and ubiquitous computing · 5Security and privacy · 4Computer networks · 3 · 2 since 2021Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Digital forensics and information hiding · 77% Privacy and data protection · 23%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 75% Geometric modeling and processing · 25%
Human-computer interaction and pervasive computing
1 paper
Usability and user experience research · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Digital forensics and information hiding
steganography
0.212016
A data hiding approach for sensitive smartphone data · UbiComp 2016
Privacy and data protection
sensitive data protection
0.112016
A data hiding approach for sensitive smartphone data · UbiComp 2016
Multimedia analysis and retrieval › video retrieval
interactive video search
0.112006
On the significance of cluster-temporal browsing for generic video retrieval: a statistical analysis · ACM Multimedia 2006
Multimedia analysis and retrieval
video retrieval
0.112006
On the significance of cluster-temporal browsing for generic video retrieval: a statistical analysis · ACM Multimedia 2006
Usability and user experience research
user study
0.012006
On the significance of cluster-temporal browsing for generic video retrieval: a statistical analysis · ACM Multimedia 2006
Geometric modeling and processing › shape descriptor
fourier descriptors
0.011995
An Experimental Comparison of Autoregressive and Fourier-Based Descriptors in 2D Shape Classification · IEEE Trans. Pattern Anal. Mach. Intell. 1995
Geometric modeling and processing › shape analysis
shape classification
0.011995
An Experimental Comparison of Autoregressive and Fourier-Based Descriptors in 2D Shape Classification · IEEE Trans. Pattern Anal. Mach. Intell. 1995
Geometric modeling and processing
shape representation
0.011995
An Experimental Comparison of Autoregressive and Fourier-Based Descriptors in 2D Shape Classification · IEEE Trans. Pattern Anal. Mach. Intell. 1995

Methods — techniques the papers use, named apart from their topics

encryption · 0.2data embedding · 0.2statistical analysis · 0.1relevance feedback · 0.1fourier descriptors · 0.0autoregressive modeling · 0.0
YearPublicationVenuePosition
2025 FedMLC: White-Box Model Watermarking for Copyright Protection in Federated Learning for IoT Environment
abstract
With the widespread application of the Internet of Things (IoT), data processing has gradually migrated to edge devices that are closer to the data source. This shift has significantly improved the ability of real-time data analysis while effectively reducing bandwidth requirements and latency. Furthermore, Federated Learning (FL) has been introduced as a decentralized training method to achieve collaborative training of multiple devices while ensuring local data privacy. However, malicious clients in FL may theft trained models for unauthorized use, which causes model misuse or copyright challenges. To address these issues, this paper proposes FedMLC (Malicious client detection, Leakage tracing, and Copyright verification), a server-side white-box watermarking scheme. FedMLC utilizes the embedded watermark at different stages to achieve both traceability and copyright verification, simplifying the watermarking process. Additionally, the watermarking can also detect malicious clients in FL. Specifically, FedMLC uses the regularization term to guide the parameter signs of the normalization layer to be consistent with the watermark sign, thereby achieving watermark embedding. Experimental results show that our FL model watermarking scheme excels in malicious client detection, leakage tracing, and copyright verification, with minimal impact on model performance, able to resist various attacks such as fine-tuning, pruning, and quantization.
Weitong Chen 0002, Wei Zhang 0098, Di Wu 0050, Anja Keskinarkaus, Tapio Seppänen, Jiale Zhang 0001, Longxiang Gao, Tom H. Luan
IEEE Internet Things J.5
2025 CodePhys: Robust Video-Based Remote Physiological Measurement Through Latent Codebook Querying
abstract
Remote photoplethysmography (rPPG) aims to measure non-contact physiological signals from facial videos, which has shown great potential in many applications. Most existing methods directly extract video-based rPPG features by designing neural networks for heart rate estimation. Although they can achieve acceptable results, the recovery of rPPG signal faces intractable challenges when interference from real-world scenarios takes place on facial video. Specifically, facial videos are inevitably affected by non-physiological factors (e.g., camera device noise, defocus, and motion blur), leading to the distortion of extracted rPPG signals. Recent rPPG extraction methods are easily affected by interference and degradation, resulting in noisy rPPG signals. In this paper, we propose a novel method named CodePhys, which innovatively treats rPPG measurement as a code query task in a noise-free proxy space (i.e., codebook) constructed by ground-truth PPG signals. We consider noisy rPPG features as queries and generate high-fidelity rPPG features by matching them with noise-free PPG features from the codebook. Our approach also incorporates a spatial-aware encoder network with a spatial attention mechanism to highlight physiologically active areas and uses a distillation loss to reduce the influence of non-periodic visual interference. Experimental results on four benchmark datasets demonstrate that CodePhys outperforms state-of-the-art methods in both intra-dataset and cross-dataset settings.
Shuyang Chu, Menghan Xia, Mengyao Yuan, Xin Liu 0012, Tapio Seppänen, Guoying Zhao 0001, Jingang Shi
IEEE J. Biomed. Health Informatics5
2024 Real-time and screen-cam robust screen watermarking
Weitong Chen 0002, Zhenhao Niu, Yanyan Xu 0003, Anja Keskinarkaus, Tapio Seppänen, Xiaobing Sun 0001
Knowl. Based Syst.6
2023 Retrieved Generative Captioning for Medical Images
abstract
Understanding the content of medical images and mapping it into text is a very trending topic in intersection of two main domains; computer vision and natural language processing. This is known as medical image captioning, which plays a vital role in developing automatic systems for diagnosis purposes. Recent research in the medical field provided promising results for both deep-learning based and retrieval-based models for image captioning. However, each one of them has its own drawbacks, that can be overcome if combined. In addition, existing diagnosis systems are still not able to provide enough explanation about the findings, which might be similar to what a physician can deliver. In this regard, we present in this paper a combination of a generative deep-learning based method and a retrieval-based model for medical image captioning. First, we train an attention-based encoder-decoder model to generate new captions for given medical images. Then, we fit the generated caption from the generative model to the retrieval-based model, which retrieves the most similar caption from the training database. This multi-stage approach allows us to generate most important words of the caption (with the generative model) and then search for the most close caption that includes such words (with the retrieval-based model). Another way of combining both models is by selecting at each time the caption with highest score among generated and retrieved captions. We evaluate our proposed model on the medical ROCO dataset for which we achieved a BLEU-4 score of 07.89 for the radiology class and 03.19 for the out-of-class data, for the multi-stage model. Similarly, best results were achieved for the fused model (predicted caption is the best among generated and retrieved) where we obtain a BLEU-4 values of 18.61 for the radiology class and 13.28 for the out-of-class data. Even though our results seem to be low, they outperformed the state-of-the-art results on the same dataset and could be further improved.
Djamila Romaissa Beddiar, Mourad Oussalah 0002, Tapio Seppänen
CBMI3
2023 A Deep learning based data augmentation method to improve COVID-19 detection from medical imaging
abstract
The worldwide spread of the Coronavirus pandemic and its huge impact challenged medical and research communities to explore novel approaches for medical diagnosis from medical imaging. However, the availability of training samples makes it difficult to implement efficient deep learning AI based solutions. In this regard, we propose new data augmentation strategies to compensate for this limitation. Our approach uses noise estimation to preserve the noise/signal ratio of the original images, while performing data augmentation. A pre-trained image Denoising Deep Neural Network DnCNN is used to calculate various sets of augmented images. First, original images are denoised. A Gaussian noise is then applied on the original images with the estimated variance computed for each class to create noisy images which are denoised again with the same DnCNN. Created subsets are fused with the original one and introduced to a Temporal Convolutional Network TCN for classification into COVID-19 and no-COVID-19 classes. We evaluated the performance of our proposal using some pre-trained networks and a convolutional neural network on three popular Covid-19 imaging datasets and the results were compared to several state-of-the-art models, demonstrating the feasibility and technical soundness of our proposal. The outcomes are also investigated to provide some explainability cues.
Djamila Romaissa Beddiar, Mourad Oussalah 0002, Tapio Seppänen
Knowl. Based Syst.4
2022 SAM: Self-augmentation mechanism for COVID-19 detection using chest X-ray images
abstract
COVID-19 is a rapidly spreading viral disease and has affected over 100 countries worldwide. The numbers of casualties and cases of infection have escalated particularly in countries with weakened healthcare systems. Recently, reverse transcription-polymerase chain reaction (RT-PCR) is the test of choice for diagnosing COVID-19. However, current evidence suggests that COVID-19 infected patients are mostly stimulated from a lung infection after coming in contact with this virus. Therefore, chest X-ray (i.e., radiography) and chest CT can be a surrogate in some countries where PCR is not readily available. This has forced the scientific community to detect COVID-19 infection from X-ray images and recently proposed machine learning methods offer great promise for fast and accurate detection. Deep learning with convolutional neural networks (CNNs) has been successfully applied to radiological imaging for improving the accuracy of diagnosis. However, the performance remains limited due to the lack of representative X-ray images available in public benchmark datasets. To alleviate this issue, we propose a self-augmentation mechanism for data augmentation in the feature space rather than in the data space using reconstruction independent component analysis (RICA). Specifically, a unified architecture is proposed which contains a deep convolutional neural network (CNN), a feature augmentation mechanism, and a bidirectional LSTM (BiLSTM). The CNN provides the high-level features extracted at the pooling layer where the augmentation mechanism chooses the most relevant features and generates low-dimensional augmented features. Finally, BiLSTM is used to classify the processed sequential information. We conducted experiments on three publicly available databases to show that the proposed approach achieves the state-of-the-art results with accuracy of 97%, 84% and 98%. Explainability analysis has been carried out using feature visualization through PCA projection and t-SNE plots.
Md. Ziaul Hoque, Mourad Oussalah 0002, Anja Keskinarkaus, Tapio Seppänen, Pinaki Sarder
Knowl. Based Syst.5
2022 Physical Violence Detection Based on Distributed Surveillance Cameras
Susu Yan, Jialing Zhen, Tian Han 0007, Hany Ferdinando, Tapio Seppänen, Esko Alasaarela
Mob. Networks Appl.6
2022 Pain fingerprinting using multimodal sensing: pilot study
abstract
Abstract Pain is a complex phenomenon, the experience of which varies widely across individuals. At worst, chronic pain can lead to anxiety and depression. Cost-effective strategies are urgently needed to improve the treatment of pain, and thus we propose a novel home-based pain measurement system for the longitudinal monitoring of pain experience and variation in different patients with chronic low back pain. The autonomous nervous system and audio-visual features are analyzed from heart rate signals, voice characteristics and facial expressions using a unique measurement protocol. Self-reporting is utilized for the follow-up of changes in pain intensity, induced by well-designed physical maneuvers, and for studying the consecutive trends in pain. We describe the study protocol, including hospital measurements and questionnaires and the implementation of the home measurement devices. We also present different methods for analyzing the multimodal data: electroencephalography, audio, video and heart rate. Our intention is to provide new insights using technical methodologies that will be beneficial in the future not only for patients with low back pain but also patients suffering from any chronic pain.
Anja Keskinarkaus, Ruijing Yang, Angelos Fylakis, Md. Surat-E.-Mostafa, Arto J. Hautala, Yong Hu 0003, Jinye Peng 0001, Guoying Zhao 0001, Tapio Seppänen, Jaro Karppinen
Multim. Tools Appl.9
2022 Non-Contact Atrial Fibrillation Detection From Face Videos by Learning Systolic Peaks
abstract
OBJECTIVE: We propose a non-contact approach for atrial fibrillation (AF) detection from face videos. METHODS: Face videos, electrocardiography (ECG), and contact photoplethysmography (PPG) from 100 healthy subjects and 100 AF patients are recorded. Data recordings from healthy subjects are all labeled as healthy. Two cardiologists evaluated ECG recordings of patients and labeled each recording as AF, sinus rhythm (SR), or atrial flutter (AFL). We use the 3D convolutional neural network for remote PPG monitoring and propose a novel loss function (Wasserstein distance) to use the timing of systolic peaks from contact PPG as the label for our model training. Then a set of heart rate variability (HRV) features are calculated from the inter-beat intervals, and a support vector machine (SVM) classifier is trained with HRV features. RESULTS: Our proposed method can accurately extract systolic peaks from face videos for AF detection. The proposed method is trained with subject-independent 10-fold cross-validation with 30 s video clips and tested on two tasks. 1) Classification of healthy versus AF: the accuracy, sensitivity, and specificity are 96.00%, 95.36%, and 96.12%. 2) Classification of SR versus AF: the accuracy, sensitivity, and specificity are 95.23%, 98.53%, and 91.12%. In addition, we also demonstrate the feasibility of non-contact AFL detection. CONCLUSION: We achieve good performance of non-contact AF detection by learning systolic peaks. SIGNIFICANCE: non-contact AF detection can be used for self-screening of AF symptoms for suspectable populations at home or self-monitoring of AF recurrence after treatment for chronic patients.
Zhaodong Sun, Juhani Junttila, Mikko Tulppo, Tapio Seppänen
IEEE J. Biomed. Health Informatics4
2021 Verbalization has regulatory influences on autonomic activity during recall of unpleasant experience
Antti Rantanen, Seppo J. Laukka, Antti Siipo, Suvi Tiinanen, Mika P. Tarvainen, Jukka Kortelainen, Matti Lehtihalmes, Tapio Seppänen
Speech Commun.8
2020 A Multi-sensor School Violence Detecting Method Based on Improved Relief-F and D-S Algorithms
Jifu Shi, Hany Ferdinando, Tapio Seppänen, Esko Alasaarela
Mob. Networks Appl.4
2020 Atrial Fibrillation Detection From Face Videos by Fusing Subtle Variations
abstract
Atrial fibrillation (AF) is one of the most common cardiac arrhythmias, which particularly occurs in the elderly individuals with heart disease. Though AF is often asymptomatic during normal activities, it has huge potential risks for stroke and other severe diseases. Thus, early detection of AF has great importance in the field of public health. Currently, electrocardiography (ECG) is the commonly used measure for the diagnosis of AF, which presents the irregular rhythm of waveform for AF patients. However, the measurement of the ECG signal requires special medical acquisition devices, which are not comfortable for practical monitoring in daily life. In this paper, we explore a very promising algorithm to detect AF from remote face videos by analyzing the color variations of face skin. The main challenge is that the current remote photoplethysmography (rPPG) technique is rather immature, which causes difficulty in extracting accurate pulse signals for describing the cardiac rhythm. To solve this problem, we first utilize various rPPG algorithms to capture pulse rhythms from different regions on the face video. We then investigate biomedical statistical methods to extract suitable features from each pulse signal. Due to the imprecision of video-extracted pulse signals, some traditional physiological features may lose their utility since they were originally proposed for ECG signals. Furthermore, some of them are very susceptible to the influence of noise. Thus, we propose a feature fusion algorithm to select and combine reasonable information from multiple physiological features, which aims to preserve the discriminability of detecting AF in the presence of the noise and outlier disturbances. The experimental results on a real-world database demonstrate the effectiveness of the proposed method in providing useful information for AF detection.
Jingang Shi, Iman Alikhani, Zitong Yu, Tapio Seppänen, Guoying Zhao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2019 3D Multi-Resolution Optical Flow Analysis of Cardiovascular Pulse Propagation in Human Brain
abstract
The brain is cleaned from waste by glymphatic clearance serving a similar purpose as the lymphatic system in the rest of the body. Impairment of the glymphatic brain clearance precedes protein accumulation and reduced cognitive function in Alzheimer's disease (AD). Cardiovascular pulsations are a primary driving force of the glymphatic brain clearance. We developed a method to quantify cardiovascular pulse propagation in the human brain with magnetic resonance encephalography (MREG). We extended a standard optical flow estimation method to three spatial dimensions, with a multi-resolution processing scheme. We added application-specific criteria for discarding inaccurate results. With the proposed method, it is now possible to estimate the propagation of cardiovascular pulse wavefronts from the whole brain MREG data sampled at 10 Hz. The results show that on average the cardiovascular pulse propagates from major arteries via cerebral spinal fluid spaces into all tissue compartments in the brain. We present an example, that cardiovascular pulsations are significantly altered in AD: coefficient of variation and sample entropy of the pulse propagation speed in the lateral ventricles change in AD. These changes are in line with the theory of glymphatic clearance impairment in AD. The proposed non-invasive method can assess a performance indicator related to the glymphatic clearance in the human brain.
Zalan Rajna, Lauri Raitamaa, Timo Tuovinen, Janne Heikkilä, Vesa Kiviniemi, Tapio Seppänen
IEEE Trans. Medical Imaging6
2018 The OBF Database: A Large Face Video Database for Remote Physiological Signal Measurement and Atrial Fibrillation Detection
abstract
Physiological signals, including heart rate (HR), heart rate variability (HRV), and respiratory frequency (RF) are important indicators of our health, which are usually measured in clinical examinations. Traditional physiological signal measurement often involves contact sensors, which may be inconvenient or cause discomfort in long-term monitoring sessions. Recently, there were studies exploring remote HR measurement from facial videos, and several methods have been proposed. However, previous methods cannot be fairly compared, since they mostly used private, self-collected small datasets as there has been no public benchmark database for the evaluation. Besides, we haven't found any study that validates such methods for clinical applications yet, e.g., diagnosing cardiac arrhythmias/disease, which could be one major goal of this technology. In this paper, we introduce the Oulu Bio-Face (OBF) database as a benchmark set to fill in the blank. The OBF database includes large number of facial videos with simultaneously recorded reference physiological signals. The data were recorded both from healthy subjects and from patients with atrial fibrillation (AF), which is the most common sustained and widespread cardiac arrhythmia encountered in clinical practice. Accuracy of HR, HRV and RF measured from OBF videos are provided as the baseline results for future evaluation. We also demonstrated that the video-extracted HRV features can achieve promising performance for AF detection, which has never been studied before. From a wider outlook, the remote technology may lead to convenient self-examination in mobile condition for earlier diagnosis of the arrhythmia.
Iman Alikhani, Jingang Shi, Tapio Seppänen, Juhani Junttila, Kirsi Majamaa-Voltti, Mikko Tulppo, Guoying Zhao 0001
FG4
2018 A Combined Motion-Audio School Bullying Detection Algorithm
abstract
School bullying is a common social problem, which affects children both mentally and physically, making the prevention of bullying a timeless topic all over the world. This paper proposes a method for detecting bullying in school based on activity recognition and speech emotion recognition. In this method, motion and voice data are gathered by movement sensors and a microphone, followed by extraction of a set of motion and audio features to distinguish bullying incidents from daily life events. Among extracted motion features are both time-domain and frequency-domain features, while audio features are computed with classical MFCCs. Feature selection is implemented using the wrapper approach. At the next stage, these motion and audio features are merged to form combined feature vectors for classification, and LDA is used for further dimension reduction. A BPNN is trained to recognize bullying activities and distinguish them from normal daily life activities. The authors also propose an action transition detection method to reduce computational complexity for practical use. Thus, the bullying detection algorithm will only run, when an action transition event has been detected. Simulation results show that the combined motion-audio feature vector outperforms separate motion features and acoustic features, achieving an accuracy of 82.4% and a precision of 92.2%. Moreover, with the action transition method, the computation cost can be reduced by half.
Hany Ferdinando, Tapio Seppänen, Esko Alasaarela
Int. J. Pattern Recognit. Artif. Intell.5
2018 Increasing the capturing angle in print-cam robust watermarking
Anu Pramila, Anja Keskinarkaus, Tapio Seppänen
J. Syst. Softw.3
2017 Enhancing Emotion Recognition from ECG Signals using Supervised Dimensionality Reduction
Hany Ferdinando, Tapio Seppänen, Esko Alasaarela
ICPRAM2
2017 Extracting watermarks from printouts captured with wide angles using computational photography
Anu Pramila, Anja Keskinarkaus, Valtteri Takala, Tapio Seppänen
Multim. Tools Appl.4
2016 Comparing features from ECG pattern and HRV analysis for emotion recognition system
abstract
We propose new features for emotion recognition from short ECG signals. The features represent the statistical distribution of dominant frequencies, calculated using spectrogram analysis of intrinsic mode function after applying the bivariate empirical mode decomposition to ECG. KNN was used to classify emotions in valence and arousal for a 3-class problem (low-medium-high). Using ECG from the Mahnob-HCI database, the average accuracies for valence and arousal were 55.8% and 59.7% respectively with 10-fold cross validation. The accuracies using features from standard Heart Rate Variability analysis were 42.6% and 47.7% for valence and arousal respectively for the 3-class problem. These features were also tested using subject-independent validation, achieving an accuracy of 59.2% for valence and 58.7% for arousal. The proposed features also showed better performance compared to features based on statistical distribution of instantaneous frequency, calculated using Hilbert transform of intrinsic mode function after applying standard empirical mode decomposition and bivariate empirical mode decomposition to ECG. We conclude that the proposed features offer a promising approach to emotion recognition based on short ECG signals. The proposed features could be potentially used also in applications in which it is important to detect quickly any changes in emotional state.
Hany Ferdinando, Tapio Seppänen, Esko Alasaarela
CIBCB2
2016 A data hiding approach for sensitive smartphone data
abstract
We develop and evaluate a data hiding method that enables smartphones to encrypt and embed sensitive information into carrier streams of sensor data. Our evaluation considers multiple handsets and a variety of data types, and we demonstrate that our method has a computational cost that allows real-time data hiding on smartphones with negligible distortion of the carrier stream. These characteristics make it suitable for smartphone applications involving privacy-sensitive data such as medical monitoring systems and digital forensics tools.
Chu Luo, Angelos Fylakis, Juha Partala, Simon Klakegg, Jorge Gonçalves 0001, Kaitai Liang, Tapio Seppänen, Vassilis Kostakos
UbiComp7
2016 Multi-modal emotion analysis from facial expressions and electroencephalogram
Xiaohua Huang 0003, Jukka Kortelainen, Guoying Zhao 0001, Antti Moilanen, Tapio Seppänen, Matti Pietikäinen
Comput. Vis. Image Underst.6
2016 MORE - a multimodal observation and analysis system for social interaction research
Anja Keskinarkaus, Sami Huttunen, Antti Siipo, Jukka Holappa, Magda Laszlo, Ilkka Juuso, Eero Väyrynen, Janne Heikkilä, Matti Lehtihalmes, Tapio Seppänen, Seppo J. Laukka
Multim. Tools Appl.10
2015 An instance-based physical violence detection algorithm for school bullying prevention
abstract
School bullying is a common social problem around the world which affects teenagers, and physical violence is considered to be the most harmful. This paper proposed an automatic physical bullying detection method with movement sensors to protect teenagers. Four features were extracted from acceleration and gyro data, and an Instance-Based classifier was applied upon them. Altogether eight kinds of activities, including three bullying kinds and five daily-life kinds, were acted by role playing. Simulations were performed on these data, and the results showed that the proposed algorithm could recognize physical bullying events and distinguish them from daily-life ones at an average accuracy of 80%. This showed a promise in automatic school bullying prevention with activity recognition techniques.
Hany Ferdinando, Tapio Seppänen, Tuija Huuki, Esko Alasaarela
IWCMC3
2013 Security threats against the transmission chain of a medical health monitoring system
abstract
One of the most important aspects of a wireless health monitoring system is the security of data. In this paper, security attacks against the complete transmission chain of a medical health monitoring system are enumerated and classified based on their threat to three security principles: confidentiality, integrity and availability. The communication chain is divided in a standard way into three tiers and relevant threats are identified for each tier. Security requirements corresponding to these threats are presented. It is noted that end-to-end security is not feasible due to distributed computing and the incompatibility of the data standards of different tiers.
Juha Partala, Niina Keränen, Mariella Särestöniemi, Matti Hämäläinen 0001, Jari H. Iinatti, Timo Jämsä, Jarmo Reponen, Tapio Seppänen
Healthcom8
2013 Classifier-Based Learning of Nonlinear Feature Manifold for Visualization of Emotional Speech Prosody
abstract
Visualization of emotional speech data is an important tool for speech researchers who seek means to gain a deeper insight into the structure of complex multidimensional data. A visualization method is presented that utilizes feature selection and classifier optimization for learning Isomap manifolds of emotional speech data. The resulting manifold is based on those features that best discriminate between given emotional classes in the target space of specified embedding dimension. A nonlinear mapping function based on generalized regression neural networks (GRNNs) provides generalization for new data. A low-dimensional manifold of emotional speech data consisting of neutral, sad, angry, and happy expressions was constructed using prosodic and acoustic features of speech. Experimental results indicate that a 3D embedding provides the best classification performance. The manifold structure can be readily visualized and matches the circumplex and conical shapes predicted by dimensional models of emotion. Listening tests show excellent correlation between the organization of the data on the manifold and the listeners' judgment of emotional intensity.
Eero Väyrynen, Jukka Kortelainen, Tapio Seppänen
IEEE Trans. Affect. Comput.3
2012 Optimal anisotropic lead scaling of multichannel ECG to reduce magnitude signal variability
abstract
A method for selecting the best functional to nonlinearly project multilead electrocardiogram (ECG) measurements into a specific type of single channel signal is presented. The functional is restricted to a family of timeinvariant quadratic functionals parameterized with lead-wise weights. This way, the projected signals are useful in multilead ECG delineation. The method determines the optimal weights in the sense of least beat-to-beat variability, eliminating much of the extra-cardiac influence, which in its turn results in a stable signal. According to the results obtained, the multilead approach is better than using any single lead alone as signal variability is reduced in 80 % of the cases even when using a suboptimal uniform weighting scheme. With the presented optimal lead scaling method, the variability is further reduced in all cases compared to individual leads, and in 92 % of cases, compared to the uniform weighting scheme. The results also show that there is no single set of weights suitable for all situations due to notable variation between the test cases.
Kai Noponen, Tapio Seppänen
BIBE2
2012 Image watermarking with feature point based synchronization robust to print-scan attack
Anja Keskinarkaus, Anu Pramila, Tapio Seppänen
J. Vis. Commun. Image Represent.3
2011 Classification of emotion in spoken Finnish using vowel-length segments: Increasing reliability with a fusion technique
Eero Väyrynen, Juhani Toivanen, Tapio Seppänen
Speech Commun.3
2010 Image watermarking with a directed periodic pattern to embed multibit messages resilient to print-scan and compound attacks
Anja Keskinarkaus, Anu Pramila, Tapio Seppänen
J. Syst. Softw.3
2009 Reading Watermarks from Printed Binary Images with a Camera Phone
Anu Pramila, Anja Keskinarkaus, Tapio Seppänen
IWDW3
2009 Invariant trajectory classification of dynamical systems with a case study on ECG
Kai Noponen, Jukka Kortelainen, Tapio Seppänen
Pattern Recognit.3
2007 Multiple Domain Watermarking for Print-Scan and JPEG Resilient Data Hiding
Anu Pramila, Anja Keskinarkaus, Tapio Seppänen
IWDW3
2007 Analysing performance in a word prediction system with multiple prediction methods
Pertti Alvar Väyrynen, Kai Noponen, Tapio Seppänen
Comput. Speech Lang.3
2006 Advancing Content-Based Retrieval Effectiveness with Cluster-Temporal Browsing in Multilingual Video Databases
abstract
Interactive experiments on video retrieval systems need to address the problem of internal validity, i.e. how much the test users' experience affects the retrieval effectiveness. This paper compares the semantic retrieval performance of novice users and expert system developers. The test system utilizes cluster-temporal browsing, which combines chronological video structure and computation of similarities into single interface. Interactive experiments with eight test users were carried out in a database of ~80 hours of multilingual news video from TRECVID 2005 benchmark. A cluster-temporal browser was found to improve the retrieval effectiveness by 12% with novice system users. Expert users were able to achieve 18% better performance than the novice users. Additionally, manual search experiments demonstrated that search performance can be improved by 19-25% when a plain text search is supplemented with content-based features
Mika Rautiainen, Tapio Seppänen, Timo Ojala
ICME2
2006 Mobile DRM-Enabled Multimedia Platform for Peer-to-Peer Applications
abstract
A mobile digital rights management platform is proposed that follows the DRM reference architecture while including several important functions lacking in the OMA DRM. The platform provides strong protection methods utilizing encryption and digital watermarking, and supports content super distribution with peer-to-peer networking and Bluetooth. Licenses for defining content usage rights are acquired from license servers with a license negotiation subsystem of the client. The DRM Player software has been implemented into a smart phone and tested with many applications
Mikko Löytynoja, Timo Koskela 0001, Marko Brockman, Tapio Seppänen
ISM4
2006 Wavelet Domain Print-Scan and JPEG Resilient Data Hiding Method
Anja Keskinarkaus, Anu Pramila, Tapio Seppänen, Jaakko J. Sauvola
IWDW3
2006 On the significance of cluster-temporal browsing for generic video retrieval: a statistical analysis
abstract
In this paper, we test statistically the effect of content-based browsing in generic video retrieval. Using TRECVID 2004 and 2005 experiments, we demonstrate that content-based browsing improves retrieval over sequential queries and relevance feedback. Two user groups, novices and system developers have been used in the experiments on large and multilingual video collections. Novice users were found to achieve improvement in search effectiveness with cluster-temporal browsing by statistically significant amount. System developers did not have statistically significant difference between the different system configurations.
Mika Rautiainen, Tapio Seppänen, Timo Ojala
ACM Multimedia2
2005 Hash-based Counter Scheme for Digital Rights Management
abstract
This paper describes a counter scheme that uses hash functions to count how many times the user is allowed to play protected content in a DRM-enabled player. The proposed basic scheme can be used in scenarios where the user cannot be assumed to have online connection. We discuss the weaknesses of the proposed scheme and present alternative to the basic scheme, which increases the security of the counter
Mikko Löytynoja, Tapio Seppänen
ICME2
2005 Comparison of Visual Features and Fusion Techniques in Automatic Detection of Concepts from News Video
abstract
This study describes experiments on automatic detection of semantic concepts, which are textual descriptions about the digital video content. The concepts can be further used in content-based categorization and access of digital video repositories. Temporal gradient correlograms, temporal color correlograms and motion activity low-level features are extracted from the dynamic visual content of a video shot. Semantic concepts are detected with an expeditious method that is based on the selection of small positive example sets and computational low-level feature similarities between video shots. Detectors using several feature and fusion operator configurations are tested in 60-hour news video database from TRECVID 2003 benchmark. Results show that the feature fusion based on ranked lists gives better detection performance than fusion of normalized low-level feature spaces distances. Best performance was obtained by pre-validating the configurations of features and rank fusion operators. Results also show that minimum rank fusion of temporal color and structure provides comparable performance
Mika Rautiainen, Tapio Seppänen
ICME2
2004 Cluster-temporal browsing of large news video databases
abstract
The paper describes cluster-temporal browsing of news video databases. Cluster-temporal browsing combines content similarities and temporal adjacency into a single representation. Visual, conceptual and lexical features are used to organize and view similar shot content. Interactive experiments with eight test users have been carried out using a database of roughly 60 hours of news video. Results indicate improvements in browsing efficiency when automatic speech recognition transcripts are incorporated into browsing by visual similarity. The cluster-temporal browsing application received positive comments from the test users and performed well in overall comparison with interactive video retrieval systems in TRECVID 2003 evaluation.
Mika Rautiainen, Timo Ojala, Tapio Seppänen
ICME3
2004 A novel scheme for merging digital audio watermarking and authentication
abstract
We present a novel scheme that is able to combine digital watermarking and content authentication of digital audio. The embedding of additional data is performed in discrete wavelet domain. Watermark embedding is done by frequency hopping method, while the additional authentication data is hidden using the LSB modulation. The perceptual transparency is achieved using the frequency masking property of the HAS. The scheme obtains high robustness against standard watermark attacks and localizes the accurately tampered parts of the audio clip.
Nedeljko Cvejic, Tapio Seppänen
MMSP2
2004 Spread spectrum audio watermarking using frequency hopping and attack characterization
Nedeljko Cvejic, Tapio Seppänen
Signal Process.2
2003 Increasing robustness of an audio watermark using turbo codes
abstract
Standard spread spectrum audio watermarking algorithms offer BER unacceptable for reliable transmission of data. Causes of unreliable watermark detection are often attacks that distort the watermarked audio in fading-like manner, disabling the correlation-based detectors to extract data. In this paper, consideration of capacity of the audio watermark channel in the presence of fading-like distortion is performed. It is shown that turbo codes offer convenient trade-off between capacity of the watermark channel and BER due to large coding gain in fading environment. Test results proved a large advantage of the proposed algorithm over standard detection as robustness is significantly increased for a given watermark data rate.
Nedeljko Cvejic, Djordje Tujkovic, Tapio Seppänen
ICME3
2003 Prosody-based classification of emotions in spoken finnish
abstract
An emotional speech corpus of Finnish was collected that includes utterances of four emotional states of speakers. More than 40 prosodic features were derived and automatically computed for the speech samples. Statistical classification experiments with kNN classifier and human listening tests indicate that emotion recognition performance comparable to human listeners can be achieved.
Tapio Seppänen, Eero Väyrynen, Juhani Toivanen
INTERSPEECH1
2003 Increasing Robustness of an Improved Spread Spectrum Audio Watermarking Method Using Attack Characterization
Nedeljko Cvejic, Tapio Seppänen
IWDW2
2003 Adapting applications in handheld devices using fuzzy context information
abstract
Context-aware devices are able to take advantage of fusing sensory and application specific information to provide proper information on a situation, for more flexible services, and adaptive user interfaces (UI). It is characteristic for handheld devices and their users that they are continuously moving in several simultaneous fuzzy contexts. The dynamic environment sets special requirements for usability and acceptance of context-aware applications. Context-aware applications must be able to operate sensibly even if the context recognition is not 100% reliable and there are multiple contexts present at the same time. We present an approach for controlling context-aware applications in the case of multiple fuzzy contexts. This work has several potential applications in the area of adaptive UI application control. Our study is focused on the adaptation of applications representing information in handheld devices. The design of controllers and experiments with real context data from user scenarios are presented. Experimental results show that the proposed approach enhances the capability of adapting information representation in a handheld device. User reactions indicate that they accept application adaptation in many situations while insisting on retaining the most control over their device. Moreover, user feedback indicates that abrupt adaptations and instability should be avoided in the application control.
Jani Mäntyjärvi, Tapio Seppänen
Interact. Comput.2
2003 Bayesian approach to sensor-based context awareness
Panu Korpipää, Miika Koskinen, Johannes Peltola, Satu-Marja Mäkelä, Tapio Seppänen
Pers. Ubiquitous Comput.5
2002 Adapting Applications in Mobile Terminals Using Fuzzy Context Information
Jani Mäntyjärvi, Tapio Seppänen
Mobile HCI2
2001 Recognizing human motion with multiple acceleration sensors
abstract
In this paper experiments with acceleration sensors are described for human activity recognition of a wearable device user. The use of principal component analysis and independent component analysis with a wavelet transform is tested for feature generation. Recognition of human activity is examined with a multilayer perceptron classifier. Best classification results for recognition of different human motion were 83-90%, and they were achieved by utilizing independent component analysis and principal component analysis. The difference between these methods turned out to be negligible.
Jani Mäntyjärvi, Johan Himberg, Tapio Seppänen
SMC3
1997 A distributed management system for testing document image analysis algorithms
abstract
We describe a new approach to manage the testing of document analysis and understanding applications. We propose and present a collection of document images, a set of techniques to prepare the test cases interactively and means to control the testing process. The systems architecture is designed to be distributed, scalable and platform independent utilizing Java, C++ and object-oriented databases. The main features of this system are a basic document categorization and ground truth, degradation models, custom test case creation facilities, a test management module (pipelining, test history), the ability to embed document analysis algorithms into the system, remote usage facilities and robust graphical user interfaces.
Jaakko J. Sauvola, Sami Haapakoski, Hannu Kauniskangas, Tapio Seppänen, Matti Pietikäinen, David S. Doermann
ICDAR4
1997 Adaptive Document Binarization
abstract
A new method is presented for adaptive document image binarization, where the page is considered as a collection of subcomponents such as text, background and picture. The problems caused by noise, illumination and many source type related degradations are addressed. The algorithm uses document characteristics to determine (surface) attributes, often used in document segmentation. Using characteristic analysis, two new algorithms are applied to determine a local threshold for each pixel. An algorithm based on soft decision control is used for thresholding the background and picture regions. An approach utilizing local mean and variance of gray values is applied to textual regions. Tests were performed with images including different types of document components and degradations. The results show that the method adapts and performs well in each case.
Jaakko J. Sauvola, Tapio Seppänen, Sami Haapakoski, Matti Pietikäinen
ICDAR2
1995 An Experimental Comparison of Autoregressive and Fourier-Based Descriptors in 2D Shape Classification
abstract
An experimental comparison of shape classification methods based on autoregressive modeling and Fourier descriptors of closed contours is carried out. The performance is evaluated using two independent sets of data: images of letters and airplanes. Silhouette contours are extracted from non-occluded 2D objects rotated, scaled, and translated in 3D space. Several versions of both types of methods are implemented and tested systematically. The comparison clearly shows better performance of Fourier-based methods, especially for images containing noise.>
Hannu Kauppinen, Tapio Seppänen, Matti Pietikäinen
IEEE Trans. Pattern Anal. Mach. Intell.2
1990 A hybrid computer architecture for machine vision
abstract
A hybrid computer architecture for machine vision which combines the useful properties of different types of architectures is introduced. HYBRID, an experimental hybrid system consisting of specialized Datacube-compatible processors and a transputer network, has been developed in a Sun-3 environment. The VLSI implementation of an edge-preserving smoothing operator for the low-level vision system is described, and the performance of transputer-based systems for higher-level vision is evaluated. Methods for analyzing and optimizing the performance of a hybrid architecture are discussed.>
Matti Pietikäinen, Tapio Seppänen, Pertti Alapuranen
ICPR (2)2