VLDB 2026 Research / reviewers in the wild / expert
Jukka-Pekka Onnela
dblp:11/5392
· DBLP profile ↗
14ranked-venue papers
0as first author
9since 2021 · last 2026
0000-0001-6613-8668ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition With Text Description of the EnvironmentabstractSpeech emotion recognition(SER) systems often struggle in real-world environments, where ambient noise severely degrades their performance. This paper explores a novel approach that exploits prior knowledge of testing environments to maximize SER performance under noisy conditions. To address this task, we propose a text-guided, environment-aware training where an SER model is trained with contaminated speech samples and their paired noise description. We use a pre-trained text encoder to extract the text-based environment embedding and then fuse it to a transformer-based SER model during training and inference. We demonstrate the effectiveness of our approach through our experiment with the MSP-Podcast corpus and real-world additive noise samples collected from the Freesound and DEMAND repositories. Our experiment indicates that the text-based environment descriptions processed by alarge language model(LLM) produce representations that improve the noise-robustness of the SER system. With acontrastive learning(CL)-based representation, our proposed method can be improved by jointly fine-tuning the text encoder with the emotion recognition model. Under the -5dBsignal-to-noise ratio(SNR) level, fine-tuning the text encoder improves our CL-based representation method by 76.4% (arousal), 100.0% (dominance), and 27.7% (valence). Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Connecting mass-action models and network models for infectious diseasesabstractInfectious disease modeling is used to forecast epidemics and assess the effectiveness of intervention strategies. Although the core assumption of mass-action models of homogeneously mixed population is often implausible, they are nevertheless routinely used in studying epidemics and provide useful insights. Network models can account for the heterogeneous mixing of populations, which is especially important for studying sexually transmitted diseases. Despite the abundance of research on mass-action and network models, the relationship between them is not well understood. Here, we attempt to bridge the gap by first identifying a spreading rule that results in an exact match between disease spreading on a fully connected network and the classic mass-action models. We then propose a method for mapping epidemic spread on arbitrary networks to a form similar to that of mass-action models. We also provide a theoretical justification for the procedure. Finally, we demonstrate the application of the proposed method in the theoretical analysis of reproduction numbers and the estimation of model parameters using synthetic data based on an empirical network. The method proves advantageous in explicitly handling both finite and infinite networks, significantly reducing the computation time required to estimate model parameters for spreading processes on networks. These findings help us understand when mass-action models and network models are expected to provide similar results and identify reasons when they do not. Thien-Minh Le, Jukka-Pekka Onnela |
PLoS Comput. Biol. | 2 |
| 2024 | Keep, Delete, or Substitute: Frame Selection Strategy for Noise-Robust Speech Emotion Recognition
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso |
INTERSPEECH | 3 |
| 2024 | Selective Acoustic Feature Enhancement for Speech Emotion Recognition With Noisy SpeechabstractAspeech emotion recognition(SER) system deployed on a real-world application is highly likely to encounter speech contaminated with unconstrained background noise. To deal with this issue, aspeech enhancement(SE) module can be attached to the SER system to compensate for the environmental difference of an input. Although the SE module can improve the quality and intelligibility of a given speech, there is a risk of affecting discriminative acoustic features for SER that are resilient to environmental differences. Exploring this idea, we propose to enhance only weak features that degrade the emotion recognition performance, while keeping strong features that are resilient to environmental differences. Our model first identifies weak feature sets by using multiple models trained with one acoustic feature at a time using clean speech. After training the single-feature models, we rank each speech feature by measuring three criteria: performance, robustness, and a joint rank ranking that combines performance and robustness. We group the weak features by cumulatively incrementing the features from the bottom to the top of each rank. Once the weak feature set is defined, we only enhance those weak features, keeping the resilient features unchanged. We implement these ideas with thelow-level descriptors(LLDs). We show that extracting LLDs from an enhanced speech signal does not improve the performance of weak features. Instead, directly enhancing the LLDs lead to better performance. Our experiment with clean and noisy versions of the MSP-Podcast corpus shows that the selective feature enhancement approach proposed in this study yields a 17.7% (arousal), 21.2% (dominance), and 3.3% (valence) performance gains over a system that enhances all the LLDs for the 10dBsignal-to-noise ratio(SNR) condition. Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Adapting a Self-Supervised Speech Representation for Noisy Speech Emotion Recognition by Using Contrastive Teacher-Student LearningabstractStudies have shown high performance in the speech emotion recognition (SER) task by fine-tuning a self-supervised speech representation model. Although this model can provide emotionally discriminative embedding in clean conditions, adapting it to a noisy target environment is still required when deployed on real-world applications. For adaptation, it is essential to balance between acquiring new knowledge from noisy speech and keeping the previous knowledge acquired during the pre-training and fine-tuning of the model. Therefore, we propose a contrastive teacher-student learning framework to retrain a self-supervised speech representation model for noisy SER. To keep the knowledge of the original model, we minimize the root mean square error between the clean embeddings from the original SER model and the noisy embeddings from the retrained model. To acquire the discriminative knowledge in the target noisy condition, we also minimize the InfoNCE loss by selecting the corresponding clean embedding as a positive sample and other noisy embeddings with different emotional labels as negative samples. Our experiment with the clean and noisy version of the MSP-Podcast corpus demonstrates that the contrastive teacher-student learning framework can significantly improve the performance of the model only trained with the clean speech in the target noisy condition for all the emotional attributes. Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso |
ICASSP | 3 |
| 2023 | Computation and Memory Efficient Noise Adaptation of Wav2Vec2.0 for Noisy Speech Emotion Recognition with Skip Connection Adapters
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso |
INTERSPEECH | 3 |
| 2022 | Not All Features are Equal: Selection of Robust Features for Speech Emotion Recognition in Noisy EnvironmentsabstractSpeech emotion recognition (SER) system deployed in real-world applications often encounters noisy speech. While most noise compensation techniques consider all acoustic features to have equal impact on the SER model, some acoustic features may be more sensitive to noisy conditions. This paper investigates the noise robustness of each feature in the acoustic feature set. We focus on low-level descriptors (LLDs) commonly used in SER systems. We firstly train SER models with clean speech by only using a single LLD. Then, we rank each LLD with respect to the absolute performance on a development set contaminated with noise, and the relative performance decrease from the results from the models trained with the clean set. Our experiment shows that using all the LLDs leads to worse performance than training the system with a single robust LLD. We propose to select a group of robust features according to their performance and robustness in noisy condition. Without using any compensation method, our feature selection methods improve the performance by 24.4% (arousal), 23.9% (dominance), and 43.2% (valence) in the 10dB noisy condition. Moreover, even though the selection is conducted with the 10dB condition, our selection methods also yield performance improvements in unseen noisy recording conditions. Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso |
ICASSP | 3 |
| 2021 | Separation of Emotional and Reconstruction Embeddings on Ladder Network to Improve Speech Emotion Recognition Robustness in Noisy Conditions
Seong-Gyun Leem, Daniel Fulford, Jukka-Pekka Onnela, David Gard, Carlos Busso |
Interspeech | 3 |
| 2021 | Bidirectional imputation of spatial GPS trajectories with missingness using sparse online Gaussian ProcessabstractOBJECTIVE: We propose a bidirectional GPS imputation method that can recover real-world mobility trajectories even when a substantial proportion of the data are missing. The time complexity of our online method is linear in the sample size, and it provides accurate estimates on daily or hourly summary statistics such as time spent at home and distance traveled. MATERIALS AND METHODS: To preserve a smartphone's battery, GPS may be sampled only for a small portion of time, frequently <10%, which leads to a substantial missing data problem. We developed an algorithm that simulates an individual's trajectory based on observed GPS location traces using sparse online Gaussian Process to addresses the high computational complexity of the existing method. The method also retains the spherical geometry of the problem, and imputes the missing trajectory in a bidirectional fashion with multiple condition checks to improve accuracy. RESULTS: We demonstrated that (1) the imputed trajectories mimic the real-world trajectories, (2) the confidence intervals of summary statistics cover the ground truth in most cases, and (3) our algorithm is much faster than existing methods if we have more than 3 months of observations; (4) we also provide guidelines on optimal sampling strategies. CONCLUSIONS: Our approach outperformed existing methods and was significantly faster. It can be used in settings in which data need to be analyzed and acted on continuously, for example, to detect behavioral anomalies that might affect treatment adherence, or to learn about colocations of individuals during an epidemic. Jukka-Pekka Onnela |
J. Am. Medical Informatics Assoc. | 2 |
| 2020 | Determining sample size and length of follow-up for smartphone-based digital phenotyping studiesabstractOBJECTIVE: Studies that use patient smartphones to collect ecological momentary assessment and sensor data, an approach frequently referred to as digital phenotyping, have increased in popularity in recent years. There is a lack of formal guidelines for the design of new digital phenotyping studies so that they are powered to detect both population-level longitudinal associations as well as individual-level change points in multivariate time series. In particular, determining the appropriate balance of sample size relative to the targeted duration of follow-up is a challenge. MATERIALS AND METHODS: We used data from 2 prior smartphone-based digital phenotyping studies to provide reasonable ranges of effect size and parameters. We considered likelihood ratio tests for generalized linear mixed models as well as for change point detection of individual-level multivariate time series. RESULTS: We propose a joint procedure for sequentially calculating first an appropriate length of follow-up and then a necessary minimum sample size required to provide adequate power. In addition, we developed an accompanying accessible sample size and power calculator. DISCUSSION: The 2-parameter problem of identifying both an appropriate sample size and duration of follow-up for a longitudinal study requires the simultaneous consideration of 2 analysis methods during study design. CONCLUSION: The temporally dense longitudinal data collected by digital phenotyping studies may warrant a variety of applicable analysis choices. Our use of generalized linear mixed models as well as change point detection to guide sample size and study duration calculations provide a tool to effectively power new digital phenotyping studies. Ian Barnett, John B. Torous, Harrison T. Reeder, Justin Baker, Jukka-Pekka Onnela |
J. Am. Medical Informatics Assoc. | 5 |
| 2019 | Use of Beiwe Smartphone App to Identify and Track Speech Decline in Amyotrophic Lateral Sclerosis (ALS)
Kathryn P. Connaghan, Jordan R. Green, Sabrina Paganoni, James Chan, Harli Weber, Ella Collins, Brian Richburg, Marziye Eshghi, Jukka-Pekka Onnela, James D. Berry |
INTERSPEECH | 9 |
| 2018 | Beyond smartphones and sensors: choosing appropriate statistical methods for the analysis of longitudinal dataabstractObjectives: As smartphones and sensors become more prominently used in mobile health, the methods used to analyze the resulting data must also be carefully considered. The advantages of smartphone-based studies, including large quantities of temporally dense longitudinally captured data, must be matched with the appropriate statistical methods in order draw valid conclusions. In this paper, we review and provide recommendations in 3 critical domains of analysis for these types of temporally dense longitudinal data and highlight how misleading results can arise from improper use of these methods. Target Audience: Clinicians, biostatisticians, and data analysts who have digital phenotyping data or are interested in performing a digital phenotyping study or any other type of longitudinal study with frequent measurements taken over an extended period of time. Scope: We cover the following topics: 1) statistical models using longitudinal repeated measures, 2) multiple comparisons of correlated tests, and 3) dimension reduction for correlated behavioral covariates. While these 3 classes of methods are frequently used in digital phenotyping data analysis, we demonstrate via actual clinical studies data that they may sometimes not perform as expected when applied to novel digital data. Ian Barnett, John B. Torous, Patrick C. Staples, Matcheri S. Keshavan, Jukka-Pekka Onnela |
J. Am. Medical Informatics Assoc. | 5 |
| 2017 | Biomarker correlation network in colorectal carcinoma by tumor anatomic locationabstractBACKGROUND: Colorectal carcinoma evolves through a multitude of molecular events including somatic mutations, epigenetic alterations, and aberrant protein expression, influenced by host immune reactions. One way to interrogate the complex carcinogenic process and interactions between aberrant events is to model a biomarker correlation network. Such a network analysis integrates multidimensional tumor biomarker data to identify key molecular events and pathways that are central to an underlying biological process. Due to embryological, physiological, and microbial differences, proximal and distal colorectal cancers have distinct sets of molecular pathological signatures. Given these differences, we hypothesized that a biomarker correlation network might vary by tumor location. RESULTS: We performed network analyses of 54 biomarkers, including major mutational events, microsatellite instability (MSI), epigenetic features, protein expression status, and immune reactions using data from 1380 colorectal cancer cases: 690 cases with proximal colon cancer and 690 cases with distal colorectal cancer matched by age and sex. Edges were defined by statistically significant correlations between biomarkers using Spearman correlation analyses. We found that the proximal colon cancer network formed a denser network (total number of edges, n = 173) than the distal colorectal cancer network (n = 95) (P < 0.0001 in permutation tests). The value of the average clustering coefficient was 0.50 in the proximal colon cancer network and 0.30 in the distal colorectal cancer network, indicating the greater clustering tendency of the proximal colon cancer network. In particular, MSI was a key hub, highly connected with other biomarkers in proximal colon cancer, but not in distal colorectal cancer. Among patients with non-MSI-high cancer, BRAF mutation status emerged as a distinct marker with higher connectivity in the network of proximal colon cancer, but not in distal colorectal cancer. CONCLUSION: In proximal colon cancer, tumor biomarkers tended to be correlated with each other, and MSI and BRAF mutation functioned as key molecular characteristics during the carcinogenesis. Our findings highlight the importance of considering multiple correlated pathways for therapeutic targets especially in proximal colon cancer. Reiko Nishihara, Kimberly Glass, Kosuke Mima, Tsuyoshi Hamada, Jonathan A. Nowak, Zhi Rong Qian, Peter Kraft, Edward L. Giovannucci, Charles S. Fuchs, Andrew T. Chan, John Quackenbush, Shuji Ogino, Jukka-Pekka Onnela |
BMC Bioinform. | 13 |
| 2011 | Understanding the Demographics of Twitter Users
Alan Mislove, Sune Lehmann, Yong-Yeol Ahn, Jukka-Pekka Onnela, J. Niels Rosenquist |
ICWSM | 4 |