Aki Härmä

dblp:21/5537 · DBLP profile ↗
← Back
34ranked-venue papers
15as first author
8since 2021 · last 2025
0000-0002-2966-3305ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 13 first-author · 5 since 2021Artificial intelligence and machine learning · 12 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Fully Autonomous Programming Using Iterative Multi-Agent Debugging with Large Language Models
abstract
Program synthesis with Large Language Models (LLMs) suffers from a “near-miss syndrome”: The generated code closely resembles a correct solution but fails unit tests due to minor errors. We address this with a multi-agent framework called Synthesize, Execute, Instruct, Debug, and Repair (SEIDR). Effectively applying SEIDR to instruction-tuned LLMs requires determining (a) optimal prompts for LLMs, (b) what ranking algorithm selects the best programs in debugging rounds, and (c) balancing the repair of unsuccessful programs with the generation of new ones. We empirically explore these tradeoffs by comparing replace-focused, repair-focused, and hybrid debug strategies. We also evaluate lexicase and tournament selection to rank candidates in each generation. On Program Synthesis Benchmark 2 (PSB2), our framework outperforms both conventional use of OpenAI Codex without a repair phase and traditional genetic programming approaches. SEIDR outperforms the use of an LLM alone, solving 18 problems in C++ and 20 in Python on PSB2 at least once across experiments. To assess generalizability, we employ GPT-3.5 and Llama 3 on the PSB2 and HumanEval-X benchmarks. Although SEIDR with these models does not surpass current state-of-the-art methods on the Python benchmarks, the results on HumanEval-C++ are promising. SEIDR with Llama 3-8B achieves an average pass@100 of 84.2%. Across all SEIDR runs, 163 of 164 problems are solved at least once with GPT-3.5 in HumanEval-C++, and 162 of 164 with the smaller Llama 3-8B. We conclude that SEIDR effectively overcomes the near-miss syndrome in program synthesis with LLMs.
Anastasiia Grishina, Vadim Liventsev, Aki Härmä, Leon Moonen
ACM Trans. Evol. Learn. Optim.3
2023 Fully Autonomous Programming with Large Language Models
abstract
Current approaches to program synthesis with Large Language Models (LLMs) exhibit a "near miss syndrome": they tend to generate programs that semantically resemble the correct answer (as measured by text similarity metrics or human evaluation), but achieve a low or even zero accuracy as measured by unit tests due to small imperfections, such as the wrong input or output format. This calls for an approach known as Synthesize, Execute, Debug (SED), whereby a draft of the solution is generated first, followed by a program repair phase addressing the failed tests. To effectively apply this approach to instruction-driven LLMs, one needs to determine which prompts perform best as instructions for LLMs, as well as strike a balance between repairing unsuccessful programs and replacing them with newly generated ones. We explore these trade-offs empirically, comparing replace-focused, repair-focused, and hybrid debug strategies, as well as different template-based and model-based prompt-generation techniques. We use OpenAI Codex as the LLM and Program Synthesis Benchmark 2 as a database of problem descriptions and tests for evaluation. The resulting framework outperforms both conventional usage of Codex without the repair phase and traditional genetic programming approaches.
Vadim Liventsev, Anastasiia Grishina, Aki Härmä, Leon Moonen
GECCO3
2023 Forecasting of Breathing Events from Speech for Respiratory Support
abstract
When a patient using a breathing support system such as a portable oxygen concentrator (POC) talks, the flow of oxygen to the lungs is disturbed. The ideal moment to administer oxygen-rich air during talking would be the brief inhale moments between utterances. However, the detection of the inhale moment is difficult and the latency of the air transfer from the pump, through the hose, to the nasal cannula may be larger than the inspiratory period in normal speech. The prediction of the next inhale moment could be used to compensate the latency and therefore provide a significantly better support for a patient who needs oxygen-rich air but want to communicate normally. In this paper we provide the first evidence that it is possible to forecasts the next inhale moment from the speech of the talker using deep learning techniques. Breathing forecasting from speech has also other potential applications that we briefly discuss in the paper.
Aki Härmä, Ulf Großekathöfer, Okke Ouweltjes, Venkata Srikanth Nallanthighal
ICASSP1
2022 Detection of COPD Exacerbation from Speech: Comparison of Acoustic Features and Deep Learning Based Speech Breathing Models
abstract
Respiration is a primary process involved in speech production. We can often hear if a person has respiratory difficulty, thus making speech a good pathological indicator for respiratory conditions. This is more relevant to conditions like chronic obstructive pulmonary disease (COPD). Patients with COPD suffer from voice changes with respect to the healthy population. Medical professionals observe that the speech of COPD patients during stable periods differs from the speech during exacerbation. In this paper, we investigate this detection of COPD exacerbation from speech in three approaches: acoustic features identification using a statistical approach, low-level descriptive features with classification, and speech breathing models based on deep learning architectures to estimate the patients’ breathing rate. Our analysis indicates that each of these approaches indeed results in a clear distinction of speech during exacerbation and stable periods of COPD.
Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik
ICASSP2
2022 COVID-19 detection based on respiratory sensing from speech
abstract
COVID-19 affects a person's respiratory health, which is manifested in the form of shortness of breath during speech. Recent work shows that it is possible to use deep learning techniques to sense the speaker's respiratory parameters from a speech signal directly. Thus respiratory parameters like speech breathing rate and tidal volume can be computed and compared using deep learning techniques to detect COVID-19 from speech recordings. In this paper, we compute respiratory parameters using our pre-trained deep learning-based speech breathing models and use them for detecting COVID-19 from speech. Apart from using speech breathing models, we perform acoustic features identification using a statistical approach and classification based on low-level descriptive features. Our analysis investigates the distinction of speech of a healthy person and COVID-19 affected person.
Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik
INTERSPEECH2
2021 On The Relationship Between Speech-Based Breathing Signal Prediction Evaluation Measures and Breathing Parameters Estimation
abstract
The respiratory system is one of the major components of the speech production system. Any alteration in breathing can result in changes in speech. Specific breathing characteristics, such as breathing rate and tidal volume, can indicate a person’s pathological condition. More recently, neural network-based methods have started emerging for predicting the breathing signal from the speech signal. The neural networks are trained and evaluated with different objective measures, such as mean squared error (MSE) and Pearson’s correlation. This paper investigates whether there is a systematic relationship between the different objective measures used for training and evaluating the neural network models and the end-goal, i.e. estimation of breathing parameters such as, breathing rate and tidal volume. Our investigations on two different data sets with two different neural network-based approaches show that there is no clear systematic relationship. In other words, obtaining a high Pearson’s correlation on the evaluation set does not necessarily mean better breathing parameter estimation. Thus, indicating the need for developing other objective evaluation measures.
Zohreh Mostaani, Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik, Mathew Magimai-Doss
ICASSP3
2021 Multi-Task Estimation of Age and Cognitive Decline from Speech
abstract
Speech is a common physiological signal that can be affected by both ageing and cognitive decline. Often the effect can be confounding, as would be the case for people at, e.g., very early stages of cognitive decline due to dementia. Despite this, the automatic predictions of age and cognitive decline based on cues found in the speech signal are generally treated as two separate tasks. In this paper, multi-task learning is applied for the joint estimation of age and the Mini-Mental Status Evaluation criteria (MMSE) commonly used to assess cognitive decline. To explore the relationship between age and MMSE, two neural network architectures are evaluated: a SincNet-based end-to-end architecture, and a system comprising of a feature extractor followed by a shallow neural network. Both are trained with single-task or multi-task targets. To compare, an SVM-based regressor is trained in a single-task setup. i-vector, x-vector and ComParE features are explored. Results are obtained on systems trained on the DementiaBank dataset and tested on an in-house dataset as well as the ADReSS dataset. The results show that both the age and MMSE estimation is improved by applying multitask learning, with state-of-the-art results achieved on the ADReSS dataset acoustic-only task.
Yilin Pan, Venkata Srikanth Nallanthighal, Daniel Blackburn, Heidi Christensen, Aki Härmä
ICASSP5
2021 Deep learning architectures for estimating breathing signal and respiratory parameters from speech recordings
abstract
Respiration is an essential and primary mechanism for speech production. We first inhale and then produce speech while exhaling. When we run out of breath, we stop speaking and inhale. Though this process is involuntary, speech production involves a systematic outflow of air during exhalation characterized by linguistic content and prosodic factors of the utterance. Thus speech and respiration are closely related, and modeling this relationship makes sensing respiratory dynamics directly from the speech plausible, however is not well explored. In this article, we conduct a comprehensive study to explore techniques for sensing breathing signal and breathing parameters from speech using deep learning architectures and address the challenges involved in establishing the practical purpose of this technology. Estimating the breathing pattern from the speech would give us information about the respiratory parameters, thus enabling us to understand the respiratory health using one's speech.
Venkata Srikanth Nallanthighal, Zohreh Mostaani, Aki Härmä, Helmer Strik, Mathew Magimai-Doss
Neural Networks3
2020 Detection of Mild Dyspnea from Pairs of Speech Recordings
abstract
Shortness of breath, or dyspnea is a condition of the cardio-pulmonary system that may be caused by, for example, a heart or lung disease, or physical load. In this paper, we explore techniques of detecting mild dyspnea directly from conversational speech, for example, in a telehealth application. We demonstrate with a collection of speech recordings before and after a light physical exercise that a siamese neural network, when presented examples of the two conditions, can detect the difference between two speech signals. This shows that this signal can be detected using data-pairs, removing the need for ratings of severity or the distinction of separate classes.
Sander M. Boelders, Venkata Srikanth Nallanthighal, Vlado Menkovski, Aki Härmä
ICASSP4
2020 Speech Breathing Estimation Using Deep Learning Methods
abstract
Breathing is the primary mechanism for maintaining the subglottal pressure for speech production. Speech can be seen as a systematic outflow of air during exhalation characterized by linguistic content and prosodic factors. Thus, sensing respiratory dynamics from the speech is plausible. In this paper, we explore techniques for sensing breathing from speech using deep learning architectures including multi-task learning approaches. Estimating the breathing pattern from the speech would give us information about the respiration rate, breathing capacity and thus enable us to understand the pathological condition of a person using one's speech. Training and evaluation of our model on our database of breathing signal and speech for 40 subjects yielded a sensitivity of 0.88 for breath event detection and 5.6 % error for breathing rate estimation.
Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik
ICASSP2
2020 A Comparison of Acoustic and Linguistics Methodologies for Alzheimer's Dementia Recognition
abstract
Contains fulltext : 228158.pdf (Publisher’s version ) (Open Access)
Nicholas Cummins, Yilin Pan, Zhao Ren, Julian Fritsch, Venkata Srikanth Nallanthighal, Heidi Christensen, Daniel Blackburn, Björn W. Schuller, Mathew Magimai-Doss, Helmer Strik, Aki Härmä
INTERSPEECH11
2019 Deep Sensing of Breathing Signal During Conversational Speech
abstract
Contains fulltext : 214126.pdf (Publisher’s version ) (Open Access)
Venkata Srikanth Nallanthighal, Aki Härmä, Helmer Strik
INTERSPEECH2
2018 Interactive health insight miner: an adaptive, semantic-based approach
abstract
E-health applications aim to support the user in adopting healthy habits.An important feature is to provide insights into the user's lifestyle.To actively engage the user in the insight mining process, we propose an ontology-based framework with a Controlled Natural Language interface, which enables the user to ask for specific insights and to customize personal information.
Isabel Funke, Rim Helaoui, Aki Härmä
INLG3
2011 Stereo audio classification for audio enhancement
abstract
Stereo audio enhancement and upmixing techniques require spatial analysis of the mixture in order to work optimally for different types of contents. In this paper a method is proposed which classifies the time-frequency regions in stereo audio data into six different classes. The individual classes represent special cases of a generic stereo signal model which is introduced and characterized in the paper. Finally, the developed classifier is tested using realistic stereo audio data.
Aki Härmä
ICASSP1
2011 Speaker Distance Detection Using a Single Microphone
abstract
A method to detect the distance of a speaker from a single microphone in a room environment is proposed. Several features, related to statistical parameters of speech source excitation signals, are introduced and are shown to depend on the distance between source and receiver. Those features are used to train a pattern recognizer for distance detection. The method is tested using a database of speech recordings in four rooms with different acoustical properties. Performance is shown to be independent of the signal gain and level, but depends on the reverberation time and the characteristics of the room. Overall, the system performs well especially for close distances and for rooms with low reverberation time and it appears to be robust to small distance mismatches. Finally, a listening test is conducted in order to compare the results of the proposed method to the performance of human listeners.
Eleftheria Georganti, Tobias May, Steven van de Par, Aki Härmä, John Mourjopoulos
IEEE Trans. Speech Audio Process.4
2009 Conversation detection in ambient telephony
abstract
In some speech communication applications such as distributed hands-free telephony it is important that the system can detect the conversational state of a call. This cannot be performed by speech activity only because the captured signal may also contain conversation between two local people, or additional speech noise sources such as speech sounds from a radio or television. In this paper we compare known algorithms and introduce a new algorithm for the real-time detection of active conversation between an incoming caller and a local user. The method is based on the mutual information in speech activity, detection of back-channel speech activity, and statistics of overlapping speech. The proposed method gives over 90% accuracy within one minute observation period which is a clear improvement over the performance of earlier techniques.
Aki Härmä
ICASSP1
2007 Ambient telephony: scenarios and research challenges
abstract
Telecommunications at home is changing rapidly. Many people have moved from the traditional PSTN phone to the mobile phone. Now for increasingly many people Voice-over-IP telephony on a PC platform is becoming the primary technology for voice communications. In this tutorial paper we give an overview of some of the current trends and try to characterize the next generation of home telephony, in particular, the concept of ambient telephony. We give an overview of the research challenges in the development of ambient telephone systems and introduce some potential solutions and scenarios.
Aki Härmä
INTERSPEECH1
2006 Parametric Representations of Bird Sounds for Automatic Species Recognition
abstract
This paper is related to the development of signal processing techniques for automatic recognition of bird species. Three different parametric representations are compared. The first representation is based on sinusoidal modeling which has been earlier found useful for highly tonal bird sounds. Mel-cepstrum parameters are used since they have been found very useful in the parallel problem of speech recognition. Finally, a vector of various descriptive features is tested because such models are popular in audio classification applications, and bird song is almost like music. We briefly introduce the methods and evaluate their performance in the classification and recognition of both individual syllables and song fragments of 14 common North-European Passerine bird species
Panu Somervuo, Aki Härmä, Seppo Fagerlund
IEEE Trans. Speech Audio Process.2
2005 Automatic estimation of reverberation time from binaural signals
abstract
An estimate of the reverberation time (RT) at the space of usage can be useful in many communications applications, such as augmented reality audio and intelligent hearing aid devices. This paper presents a method to measure the reverberation time from two microphone signals. The analysis is based on locating suitable sound segments for RT analysis by using short-time energy and inter-channel coherence measures, followed by the Schroeder integration method, line fitting and finally statistical analysis. The line fitting is used to estimate the slope of the decay. In this paper, we propose a method where the slope is estimated in the region that maximizes the correlation coefficient of the least squares method. This makes the estimation results more accurate than if fixed limits, e.g., -5 to -25 dB on the decay curve, were used, thanks to the absence of the systematic error caused by bending of the decay curves. The system performance was evaluated using a real-time version of the algorithm.
Sampo Vesa, Aki Härmä
ICASSP (3)2
2005 Automatic surveillance of the acoustic activity in our living environment
abstract
We report an experiment with an acoustic surveillance system comprised of a computer and microphone situated in a typical office environment. The system continuously analyzes the acoustic activity at the recording site, separates all interesting events, and stores them in a database. All interesting acoustic events over duration of more than two months were recorded. A number of low-level signal features are computed from the audio signal and used to classify and identify sound events. The analysis reveals interesting patterns and activities which would be difficult to find by any other means.
Aki Härmä, Martin F. McKinney, Janto Skowronek
ICME1
2004 Classification of the harmonic structure in bird vocalization
abstract
The article is related to the development of techniques for automatic recognition of bird species by their sounds. It has been demonstrated earlier that a simple model of one time-varying sinusoid is very useful in classification and recognition of typical bird sounds. However, a large class of bird sounds are not pure sinusoids but have a clear harmonic spectrum structure. We introduce a way to classify bird syllables into four classes by their harmonic structure.
Aki Härmä, Panu Somervuo
ICASSP (5)1
2004 Head-tracking and subject positioning using binaural headset microphones and common modulation anchor sources
abstract
A prerequisite in many systems for virtual and augmented reality audio is the tracking of a subject's head position and direction. When the subject is wearing binaural headset microphones, the signals from them can be cross-correlated with known sound sources, called here anchor sources, to obtain an accurate estimate of the subject's position and orientation. In this paper we propose a method where the anchor sources radiate in separate frequency bands but with common modulation signal. After demodulation, distance estimates between anchor sources and binaural microphones are obtained from positions of cross-correlation maxima, which further yield the desired head coordinates. A particularly attractive case is to use high carrier frequencies where disturbing environment noise is lower and sensitivity of hearing to detect annoying anchor sources is also lower.
Matti Karjalainen, Miikka Tikander, Aki Härmä
ICASSP (4)3
2004 Bird song recognition based on syllable pair histograms
abstract
Bird song can be divided into a sequence of syllabic elements. We investigate the possibility of bird species recognition based on the syllable pair histogram of the song. This representation compresses the variable-length syllable sequence into a fixed-dimensional feature vector. The histogram is computed by means of Gaussian syllable prototypes which are automatically found given the song data and the dissimilarity measure of syllables. Our representation captures the use of the syllable alphabet and also some temporal structure of the song. We demonstrate the method in bird species recognition with song patterns obtained from fifty individuals belonging to four common passerine bird species.
Panu Somervuo, Aki Härmä
ICASSP (5)2
2003 Automatic identification of bird species based on sinusoidal modeling of syllables
abstract
Syllables are elementary building blocks of bird song. In the sounds of many songbirds, a large class of syllables can be approximated as amplitude and frequency varying brief sinusoidal pulses. We test how well bird species can be recognized by comparing simple sinusoidal representations of isolated syllables. Results are encouraging and show that, with limited sets of bird species, a recognizer based on this signal model may already be sufficient.
Aki Härmä
ICASSP (5)1
2002 A method for parametrization of time-varying sounds
abstract
This paper presents a framework for parametrization, identification, and synthesis of brief time-varying sounds. The proposed scheme is based on a frequency-warped modification of Grenier's (1983) time-varying lattice algorithm. A new algorithm, which gives the optimal time-varying model for an ensemble of similar sounds, is introduced.
Aki Härmä, Marko Juntunen
IEEE Signal Process. Lett.1
2001 Modeling and equalization of audio systems using Kautz filters
abstract
Frequency warping using all pass structures or Laguerre filters has increasingly found applications in audio signal processing due to having a good match with auditory frequency resolution. Kautz filters are an extension where the frequency warping and related resolution can have more freedom. We discuss the properties of Kautz filters and how they meet typical requirements found in modeling and equalization of audio systems. Case studies include transfer function modeling of the guitar body and loudspeaker response equalization.
Tuomas Paatero, Matti Karjalainen, Aki Härmä
ICASSP3
2001 Linear predictive coding with modified filter structures
abstract
In conventional one-step forward linear prediction, an estimate for the current sample value is formed as a linear combination of previous sample values. In this paper, a generalized form of this scheme is studied. Here, the prediction is not based simply on the previous sample values but on the signal history as seen through an arbitrary filterbank. It is shown in the paper how the coefficients of a modified model can be obtained and how the inverse and synthesis filters can be implemented. Various properties of such systems are derived in this article. As an example, a novel linear predictive system using inherently logarithmic frequency representation is introduced.
Aki Härmä
IEEE Trans. Speech Audio Process.1
2001 A comparison of warped and conventional linear predictive coding
abstract
Frequency-warped signal processing techniques are attractive to many wideband speech and audio applications since they have a clear connection to the frequency resolution of human hearing. A warped version of linear predictive coding (LPC) is studied. The performance of conventional and warped LPC algorithms are compared in a simulated coding system using listening tests and conventional technical measures. The results indicate that the use of warped techniques is beneficial especially in wideband coding and may result in savings of one bit per sample compared to the conventional algorithm while retaining the same subjective quality.
Aki Härmä, Unto K. Laine
IEEE Trans. Speech Audio Process.1
2000 Evaluation of a warped linear predictive coding scheme
abstract
Basically all conventional digital signal processing techniques can be warped by introducing a simple modification to the system. In this paper, the focus is in warped linear predictive coding techniques with application to speech and audio coding. The performance of warped LPC is compared with a conventional LPC in listening tests and in terms of technical measures. This is done at various sampling rates as a function of the order of the LPC model.
Aki Härmä
ICASSP1
2000 Implementation of frequency-warped recursive filters
Aki Härmä
Signal Process.1
1999 On the utilization of overshoot effects in low-delay audio coding
abstract
In low-delay audio coding (coding delay <5 ms) there is no time for detailed spectral modeling in the case of brief percussive sounds, e.g., the castanets, and onsets of music or speech sounds. On the other hand, it is known from psychoacoustic experiments that the ear is not accurate near the onset of a wideband sound. We study the audibility of coding errors near the onsets of musical sounds in a simulated low-delay audio codec based on frequency-warped linear prediction. It is suggested that for many musical transients it is sufficient to reproduce a rough temporal and spectral envelope of the original signal during the first 5-10 ms. Preliminary listening tests support this idea. It is proposed that the overshoot effect of hearing could be utilized efficiently in enhancing the performance of a low-delay audio coding scheme.
Aki Härmä, Unto K. Laine, Matti Karjalainen
ICASSP1
1998 Implementation of recursive filters having delay free loops
abstract
Certain types of recursive filters have been considered as non-realizable because they contain delayless recursive loops. Usually the problem is rather technical than theoretical. A method of implementing such filters is introduced. The general procedure is to split a delay free recursive filter to a non-delay free and a pure delay free structure. As a combination of these, the filter can be implemented directly and efficiently. In addition, following from the same formulation, a generic procedure to convert any such filter to an equivalent directly realizable structure is also given. As an example, a set of frequency warped all-pole filters is considered. The new warped all-pole lattice introduced in this paper completes the family of warped filters.
Aki Härmä
ICASSP1
1997 An experimental audio codec based on warped linear prediction of complex valued signals
abstract
Bark-scale warped linear prediction (WLP) is a very potential core for a monophonic perceptual audio codec. In the current paper the WLP scheme is extended for processing complex valued signals (CWLP). Three different methods of converting a stereo signal to one complex valued signal are introduced. The philosophy behind the coding scheme is to integrate some aspects of modern wideband audio coding (e.g. perceptuality and stereo signal processing) into one computational element in order to find a more holistic and economic way of processing.
Aki Härmä, Unto K. Laine, Matti Karjalainen
ICASSP1
1997 Realizable warped IIR filters and their properties
abstract
Digital filters where unit delays are replaced with frequency dependent delays, such as first order allpass sections, are often called warped filters since they implement filter specifications on a warped non-uniform frequency scale. Warped IIR (WIIR) filters cannot be realized directly due to delay free loops. Specific solutions have been known that make WIIR filters realizable but no general approach has been available so far. In this paper we will explore the generation of such filters, including new filter structures. The robustness and computational efficiency of WIIR filters are studied and most potential applications are discussed.
Matti Karjalainen, Aki Härmä, Unto K. Laine
ICASSP2