VLDB 2026 Research / reviewers in the wild / expert
Gökhan Ince
dblp:32/1223
· DBLP profile ↗
28ranked-venue papers
10as first author
7since 2021 · last 2026
0000-0002-0034-030XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 10 first-author · 1 since 2021Systems, architecture and hardware · 15 · 7 first-authorHuman-computer interaction and ubiquitous computing · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorComputer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Automated Evaluation of Student-Chatbot Interactions in Climate Education Using Large Language Models
Eceay Çeltik, Bora Senceylan, Ibrahim Delen, Gökhan Ince |
AIED (5) | 4 |
| 2026 | Privacy-aware Hand Gesture Recognition using Textile-based Sensing Glove
Omur Fatmanur Erzurumluoglu, Mehmet Ali Sarikaya, Asli Tunçay Atalay, Ozgur Atalay, Gökhan Ince |
IWQoS | 5 |
| 2025 | Reducing dental anxiety in children using robotic companions: A comparative study of behavior management techniques
Mine Yasemin, Elif Bahar Tuna Ince, Gökhan Ince |
Int. J. Hum. Comput. Stud. | 3 |
| 2025 | FogETex: Fog Computing Framework for Electronic Textile ApplicationsabstractTextile products are present in almost every aspect of human life. With the introduction of electronic textiles (e-textiles), textile products have become capable of converting various physiological and environmental stimuli into electrical signals, many of which are of vital importance to humans. Therefore, these products require real-time (low-latency) and robust computing systems. However, due to comfort considerations, they cannot accommodate powerful computing resources. In this study, a novel fog computing-based framework (FogETex) is proposed to meet the needs of e-textile applications. FogETex is a Platform-as-a-Service model that is cross-platform supported, scalable, and operates in real time. This framework encompasses end-to-end integration of the system, including Textile-based Internet of Things (T-IoT) device, fog devices, and the cloud. Fog devices consist of a broker that manages the fog node and a worker that handles incoming computation requests. Sensor data is transmitted to the fog node through a mobile application, and system architecture can be monitored through developed user interfaces. Resource usage from broker devices is monitored in real time to prevent worker devices from experiencing overload. For the system case study, a deep-learning-based gait phase analysis application using textile-based capacitive sensors is employed. FogETex was evaluated in terms of time characteristics, resource usage, and network bandwidth usage using a mock client to determine the ideal system performance and an actual client to conduct real-world tests. The fog devices outperformed the cloud system in these metrics. Besides being developed primarily for e-textile applications, the FogETex framework can accommodate other IoT devices as well. Kadir Ozlem, Asli Tunçay Atalay, Ozgur Atalay, Gökhan Ince |
IEEE Internet Things J. | 4 |
| 2023 | An Augmented Reality Environment for Testing Cockpit Display Systems
Caner Potur, Gökhan Ince |
CHIRA (2) | 2 |
| 2022 | Usability Evaluation of TV Interfaces: Subjective Evaluation Vs. Objective EvaluationabstractThis study aims to design a usability test enhanced with subjective usability evaluation methods (concurrent think-aloud, post-task and posttest questionnaires) and objective usability evaluation methods (eye-tracking, emotion recognition and logging technologies). Furthermore, the relationships between the subjective metrics such as task difficulty level, PSSUQ (Post-Study System Usability Questionnaire) scores and objective metrics such as task completion time, eye tracking metrics (fixation count, fixation duration, average fixation duration, saccade count, saccade duration, scanpath length, blink count), logging metrics (keying and back keying count) and emotion recognition metrics (negative emotions count) are aimed to be analyzed. Therefore, a user testing study with 38 participants was conducted to evaluate the usability level of the TV interface of Digiturk, which is one of the digital TV broadcasting platforms in Turkey. The participants completed ten tasks related to the features of the TV interface, such as VOD (video-on-demand) watching, channel locking, and recording for future watches. Whether there is a significant difference between tasks in terms of subjective and objective metrics is investigated by using one-way ANOVA. The results show that while task completion time and task difficulty increase, values of all objective metrics increase. The relationships between the subjective and objective metrics are measured with Pearson correlation and the results show that there is a significant relationship between every subjective and objective metric except PSSUQ scores, which shows the perceived satisfaction level of the participants. Furthermore, exploratory factor analysis is performed to explore the latent structure of the usability measures. As a result of the factor analysis, a two-factor structure in which objective and subjective metrics load on two separate factors is obtained. Cigdem Altin Gumussoy, Aycan Pekpazar, Mustafa Esengün, A. Elvan Bayraktaroglu, Gökhan Ince |
Int. J. Hum. Comput. Interact. | 5 |
| 2021 | Real time detection of acoustic anomalies in industrial processes using sequential autoencodersabstractAbstract Development of intelligent systems with the pursuit of detecting abnormal events in real world and in real time is challenging due to difficult environmental conditions, hardware limitations, and computational algorithmic restrictions. As a result, degradation of detection performance in dynamically changing environments is often encountered. However, in the next‐generation factories, an anomaly detection system based on acoustic signals is especially required to quickly detect and interfere with the abnormal events during the industrial processes due to the increased cost of complex equipment and facilities. In this study we propose a real time Acoustic Anomaly Detection (AAD) system with the use of sequence‐to‐sequence Autoencoder (AE) models in the industrial environments. The proposed processing pipeline makes use of the audio features extracted from the streaming audio signal captured by a single‐channel microphone. The reconstruction error generated by the AE model is calculated to measure the degree of abnormality of the sound event. The performance of Convolutional Long Short‐Term Memory AE (Conv‐LSTMAE) is evaluated and compared with sequential Convolutional AE (CAE) using sounds captured from various industrial manufacturing processes. In the experiments conducted with the real time AAD system, it is shown that the Conv‐LSTMAE‐based AAD demonstrates better detection performance than CAE model‐based AAD under different signal‐to‐noise ratio conditions of sound events such as explosion, fire and glass breaking. Baris Bayram, Taha Berkay Duman, Gökhan Ince |
Expert Syst. J. Knowl. Eng. | 3 |
| 2019 | A novel interface for generating choreography based on augmented reality
Tafadzwa Joseph Dube, Gökhan Ince |
Int. J. Hum. Comput. Stud. | 2 |
| 2018 | Failure Detection Using Proprioceptive, Auditory and Visual ModalitiesabstractHandling safety is crucial to achieve lifelong autonomy for robots. Unsafe situations might arise during manipulation in unstructured environments due to noises in sensory feedback, improper action parameters, hardware limitations or external factors. In order to assure safety, continuous execution monitoring and failure detection procedures are mandatory. To this end, we present a multimodal failure monitoring and detection system to detect manipulation failures. Rather than relying only on a single sensor modality, we consider integration of different modalities to get better detection performance in different failure cases. In our system, high level proprioceptive, auditory and visual predicates are extracted by processing each modality separately. Then, the extracted predicates are fused. Experiments on a humanoid robot for tabletop manipulation scenarios indicate that the contribution of each modality is different depending on the action in execution and multimodal fusion results in an overall performance increase in detecting failures compared to the performance attained by unimodal processing. Arda Inceoglu, Gökhan Ince, Yusuf Yaslan, Sanem Sariel |
IROS | 2 |
| 2012 | Online audio beat tracking for a dancing robot in the presence of ego-motion noise in a real environmentabstractThis paper presents the design and implementation of a real-time real-world beat tracking system which runs on a dancing robot. The main problem of such a robot is that, while it is moving, ego noise is generated due to its motors, and this directly degrades the quality of the audio signal features used for beat tracking. Therefore, we propose to incorporate ego noise reduction as a pre-processing stage prior to our tempo induction and beat tracking system. The beat tracking algorithm is based on an online strategy of competing agents sequentially processing a continuous musical input, while considering parallel hypotheses regarding tempo and beats. This system is applied to a humanoid robot processing the audio from its embedded microphones on-the-fly, while performing simplistic dancing motions. A detailed and multi-criteria based evaluation of the system across different music genres and varying stationary/non-stationary noise conditions is presented. It shows improved performance and noise robustness, outperforming our conventional beat tracker (i.e., without ego noise suppression) by 15.2 points in tempo estimation and 15.0 points in beat-times prediction. João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai |
ICRA | 2 |
| 2012 | Online learning for template-based multi-channel ego noise estimationabstractThis paper presents a system that gives a robot the ability to diminish its own disturbing noise (i.e., ego noise) by utilizing template-based ego noise estimation, an algorithm previously developed by the authors. In pursuit of an autonomous, online and adaptive template learning system in this work, we specifically focus on eliminating the requirement of an offline training session performed in advance to build the essential templates, which represent the ego noise. The idea of discriminating ego noise from all other sound sources in the environment enables the robot to learn the templates online without requiring any prior information. Based on the directionality/diffuseness of the sound sources, the robot can easily decide whether the template should be discarded because it is corrupted by external noises, or it should be inserted into the database because the template consists of pure ego noise only. Furthermore, we aim to update the template database optimally by introducing an additional time-variant forgetting factor parameter, which provides a balance between adaptivity and stability of the learning process automatically. Moreover, we enhanced the single-channel noise estimation system to be compatible with the multi-channel robot audition framework so that ego noise can be eliminated from all signals stemming from multiple sound sources respectively. We demonstrate that the proposed system allows the robot to have the ability of online template learning as well as a high performance of noise estimation and suppression for multiple sound sources. Gökhan Ince, Kazuhiro Nakadai, Keisuke Nakamura |
IROS | 1 |
| 2012 | Real-time super-resolution Sound Source Localization for robotsabstractSound Source Localization (SSL) is an essential function for robot audition and yields the location and number of sound sources, which are utilized for post-processes such as sound source separation. SSL for a robot in a real environment mainly requires noise-robustness, high resolution and real-time processing. A technique using microphone array processing, that is, Multiple Signal Classification based on Standard EigenValue Decomposition (SEVD-MUSIC) is commonly used for localization. We improved its robustness against noise with high power by incorporating Generalized EigenValue Decomposition (GEVD). However, GEVD-based MUSIC (GEVD-MUSIC) has mainly two issues: 1) the resolution of pre-measured Transfer Functions (TFs) determines the resolution of SSL, 2) its computational cost is expensive for real-time processing. For the first issue, we propose a TF interpolation method integrating time-domain-based and frequency-domain-based interpolation. The interpolation achieves super-resolution SSL, whose resolution is higher than that of the pre-measured TFs. For the second issue, we propose two methods, MUSIC based on Generalized Singular Value Decomposition (GSVD-MUSIC), and Hierarchical SSL (H-SSL). GSVD-MUSIC drastically reduces the computational cost while maintaining noise-robustness in localization. H-SSL also reduces the computational cost by introducing a hierarchical search algorithm instead of using greedy search in localization. These techniques are integrated into an SSL system using a robot embedded microphone array. The experimental result showed: the proposed interpolation achieved approximately 1 degree resolution although we have only TFs at 30 degree intervals, GSVD-MUSIC attained 46.4% and 40.6% of the computational cost compared to SEVD-MUSIC and GEVD-MUSIC, respectively, H-SSL reduces 59.2% computational cost in localization of a single sound source. Keisuke Nakamura, Kazuhiro Nakadai, Gökhan Ince |
IROS | 3 |
| 2012 | Live assessment of beat tracking for robot auditionabstractIn this paper we propose the integration of an online audio beat tracking system into the general framework of robot audition, to enable its application in musically-interactive robotic scenarios. To this purpose, we introduced a staterecovery mechanism into our beat tracking algorithm, for handling continuous musical stimuli, and applied different multi-channel preprocessing algorithms (e.g., beamforming, ego noise suppression) to enhance noisy auditory signals lively captured in a real environment. We assessed and compared the robustness of our audio beat tracker through a set of experimental setups, under different live acoustic conditions of incremental complexity. These included the presence of continuous musical stimuli, built of a set of concatenated musical pieces; the presence of noises of different natures (e.g., robot motion, speech); and the simultaneous processing of different audio sources on-the-fly, for music and speech. We successfully tackled all these challenging acoustic conditions and improved the beat tracking accuracy and reaction time to music transitions while simultaneously achieving robust automatic speech recognition. João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai, Hiroshi G. Okuno, Luís Paulo Reis, Fabien Gouyon |
IROS | 2 |
| 2012 | An active audition framework for auditory-driven HRI: Application to interactive robot dancingabstractIn this paper we propose a general active audition framework for auditory-driven Human-Robot Interaction (HRI). The proposed framework simultaneously processes speech and music on-the-fly, integrates perceptual models for robot audition, and supports verbal and non-verbal interactive communication by means of (pro)active behaviors. To ensure a reliable interaction, on top of the framework a behavior decision mechanism based on active audition policies the robot's actions according to the reliability of the acoustic signals for auditory processing. To validate the framework's application to general auditory-driven HRI, we propose the implementation of an interactive robot dancing system. This system integrates three preprocessing robot audition modules: sound source localization, sound source separation, and ego noise suppression; two modules for auditory perception: live audio beat tracking and automatic speech recognition; and multi-modal behaviors for verbal and non-verbal interaction: music-driven dancing and speech-driven dialoguing. To fully assess the system, we set up experimental and interactive real-world scenarios with highly dynamic acoustic conditions, and defined a set of evaluation criteria. The experimental tests revealed accurate and robust beat tracking and speech recognition, and convincing dance beat-synchrony. The interactive sessions confirmed the fundamental role of the behavior decision mechanism for actively maintaining a robust and natural human-robot interaction. João Lobato Oliveira, Gökhan Ince, Keisuke Nakamura, Kazuhiro Nakadai, Hiroshi G. Okuno, Luís Paulo Reis, Fabien Gouyon |
RO-MAN | 2 |
| 2011 | Correlation matrix interpolation in Sound Source Localization for a robotabstractIn microphone array processing, a Correlation Matrix (CM) between multiple channel input signals is widely utilized for various purposes such as Sound Source Localization (SSL), etc. The CM corresponds to spatial information between a microphone and a sound source, and it dynamically changes as they move in a dynamic environment. Since pre-measured CMs have difficulties in dealing with dynamically changing environments due to their discreteness, this paper addresses a CM interpolation utilizing a few discrete known CMs. The proposed method is based on an integration of eigen-value-scaling and linear interpolation, which requires a few measurements and low computational costs. Apart from conventional methods, this paper deals with an interpolation based on a correlation-matrix not based on a transfer-function, which achieves the integration approach. The evaluation shows better estimation performance compared to conventional transfer-function-based methods. We also applied the interpolation to SSL and confirmed its validity. Keisuke Nakamura, Kazuhiro Nakadai, Hirofumi Nakajima, Gökhan Ince |
ICASSP | 4 |
| 2011 | Assessment of general applicability of ego noise estimationabstractNoise generated due to the motion of a robot deteriorates the quality of the desired sounds recorded by robot-embedded microphones. On top of that, a moving robot is also vulnerable to its loud fan noise that changes its orientation relative to the moving limbs where the microphones are mounted on. To tackle the non-stationary ego-motion noise and the direction changes of fan noise, we propose an estimation method based on instantaneous prediction of ego noise using parameterized templates. We verify the ego noise suppression capability of the proposed estimation method on a humanoid robot by evaluating it on two important applications in the framework of robot audition: (1) automatic speech recognition and (2) sound source localization. We demonstrate that our method improves recognition and localization performance during both head and arm motions considerably. Gökhan Ince, Keisuke Nakamura, Futoshi Asano, Hirofumi Nakajima, Kazuhiro Nakadai |
ICRA | 1 |
| 2011 | Assessment of single-channel ego noise estimation methodsabstractWhile a robot is moving, ego noise is generated due to the fans and motors of the robot. Furthermore, a robot is not only subject to the ego noise, but also to the ambient noise of the environment, both having different short-term signal characteristics. Because ego-motion noise generated by the motors is non-stationary, and the BackGround Noise (BGN) is stationary, one single noise estimation method is unable to track the changes in both noise spectra rapidly and accurately. Therefore, we propose to use the combination of two different noise estimation methods adequate for each one of co-existing noise types in a unified framework: 1) a stationary noise estimation method called Histogram-based Recursive Level Estimation (HRLE) and 2) a non-stationary noise estimation method called Template-based Estimation (TE). In this paper, we evaluate the performance of several single-channel based noise estimation techniques in terms of their prediction accuracy and quality of the speech signals enhanced by spectral subtraction methods. The experimental results show that our system, compared to the conventional single-stage noise estimation methods, achieves better performance in attaining signal quality and improving word correct rates. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Jun-ichi Imura, Keisuke Nakamura, Hirofumi Nakajima |
IROS | 1 |
| 2011 | Incremental learning for ego noise estimation of a robotabstractUsing pre-recorded templates to estimate and suppress the ego noise of a robot is advantageous because this method is able to cope with the non-stationarity of this particular type of noise. However, standard template-based estimation requires human intervention in the offline training sessions, storage of large amounts of data and does not adapt to the dynamical changes in the environmental conditions. In this paper we investigate the feasibility of an incremental template learning system to tackle these drawbacks. Incremental learning enables the system to acquire new templates on the fly and update the older ones appropriately. Whilst allowing the system to continually increase its knowledge and enhancing its estimation performance, this learning scheme also reduces the size of the database. We evaluate the performance of the proposed noise estimation method in terms of its estimation accuracy, quality of speech signals enhanced by spectral subtraction method, and size of database. The experimental results show that our system compared to conventional single-channel noise estimation methods achieves better performance in attaining signal quality and improving word correct rates. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Jun-ichi Imura, Keisuke Nakamura, Hirofumi Nakajima |
IROS | 1 |
| 2011 | Intelligent sound source localization and its application to multimodal human trackingabstractWe have assessed robust tracking of humans based on intelligent Sound Source Localization (SSL) for a robot in a real environment. SSL is fundamental for robot audition, but has three issues in a real environment: robustness against noise with high power, lack of a general framework for selective listening to sound sources, and tracking of inactive and/or noisy sound sources. To address the first issue, we extended Multiple SIgnal Classification by incorporating Generalized EigenValue Decomposition (GEVD-MUSIC) so that it can deal with high power noise and can select target sound sources. To address the second issue, we proposed Sound Source Identification (SSI) based on hierarchical gaussian mixture models and integrated it with GEVD-MUSIC to realize a selective listening function. To address the third issue, we integrated audio-visual human tracking using particle filtering. Integration of these three techniques into an intelligent human tracking system showed: 1) GEVD-MUSIC improved the noise-robustness of SSL by a signal-to-noise ratio of 5-6 dB; 2) SSI performed more than 70% in F-measure even in a noisy environment; and 3) audio-visual integration improved the average tracking error by approximately 50%. Keisuke Nakamura, Kazuhiro Nakadai, Futoshi Asano, Gökhan Ince |
IROS | 4 |
| 2011 | Ego noise cancellation of a robot using missing feature masks
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura |
Appl. Intell. | 1 |
| 2010 | A hybrid framework for ego noise cancellation of a robotabstractNoise generated due to the motion of a robot is not desired, because it deteriorates the quality and intelligibility of the sounds recorded by robot-embedded microphones. It must be reduced or cancelled to achieve automatic speech recognition with a high performance. In this work, we divide ego-motion noise problem into three subdomains of arm, leg and head motion noise, depending on their complexity and intensity levels. We investigate methods that make use of single-channel and multi-channel processing in order to suppress ego noise separately. For this purpose, a framework consisting of a microphone-array-based geometric source separation, a consequent post filtering process and a parallel module for template subtraction is used. Furthermore, a control mechanism is proposed, which is based on signal-to-noise ratio and instantaneously detected motions, to switch to the most suitable method to deal with the current type of noise. We evaluate the proposed techniques on a humanoid robot using automatic speech recognition (ASR). The preliminary results of isolated word recognition show the effectiveness of our methods by increasing the word correct rates up to 50% compared to the single channel recognition in arm and leg motion noises and up to 25% in very strong head motion noises. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Yuji Hasegawa, Hiroshi Tsujino, Jun-ichi Imura |
ICRA | 1 |
| 2010 | Robust Ego Noise Suppression of a Robot
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura |
IEA/AIE (1) | 1 |
| 2010 | A robust speech recognition system against the ego noise of a robot
Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura |
INTERSPEECH | 1 |
| 2010 | Multi-talker speech recognition under ego-motion noise using Missing Feature TheoryabstractThis paper presents a system that gives a mobile robot the ability to recognize target speaker's speech, even if the robot performs an action and there are multiple speakers talking in the room. Associated problems to this system are twofold: (1) While the robot is moving, the joints inevitably generate ego-motion noise due to its motors. (2) Recognizing target speech against other interfering speech signals is a difficult task. Since typical solutions to (1) and (2), motor noise suppression and sound source separation, both introduce distortion to the processed signals, the performance of automatic speech recognition (ASR) deteriorates. Instead of removing the ego-motion noise with conventional noise suppression methods, in this work, we investigate methods to eliminate the unreliable parts of the audio features that are contaminated by the ego-motion noise. For this purpose, we model masks that filter unreliable speech features based on the ratio of speech and motor noise energies. We analyze the performance of the proposed technique under various test conditions by comparing it to the performance of existing Missing Feature Theory-based ASR implementations. Finally, we propose an integration framework for two different masks that are designed to eliminate ego noise and to filter the leakage energy of interfering sound sources. We demonstrate that the proposed methods achieve a high ASR accuracy. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Hiroshi Tsujino, Jun-ichi Imura |
IROS | 1 |
| 2010 | Sound source separation and automatic speech recognition for moving sourcesabstractThis paper addresses sound source separation and speech recognition for moving sound sources. Real-world applications such as robots should cope with both moving and stationary sound sources. However, most studies assume only stationary sound sources. We introduce three key techniques to cope with moving sources, that is, Adaptive Step-size control (AS), Optima Controlled Recursive Average (OCRA), and Separation Parameter Switching (SPS). We implemented a real-time robot audition system with these techniques for our humanoid robot with an 8ch microphone array by using HARK which is our open-source software for robot audition. Preliminary results show that the performance of recognition of moving sound sources improved drastically, and also the performance of the system is shown through two speech dialog scenarios which requires sound source separation and automatic speech recognition for moving sources. Kazuhiro Nakadai, Hirofumi Nakajima, Gökhan Ince, Yuji Hasegawa |
IROS | 3 |
| 2010 | An easily-configurable robot audition system using Histogram-based Recursive Level EstimationabstractThis paper presents an easily-configurable robot audition system using the Histogram-based Recursive Level Estimation (HRLE) method. In order to achieve natural human-robot interaction, a robot should recognize human speeches even if there are some noises and reverberations. Since the precision of automatic speech recognizers (ASR) have been degraded by such interference, many systems applying speech enhancement processes have been reported. However, performance of most reported systems suffer from acoustical environmental changes. For example, an enhancement process optimized for steady-state noise, such as fan noise, yields low performance when the process is used for non-steady-state noises, such as background music. The primary reason is mismatches of parameters because the appropriate parameters change according to the acoustical environments. To solve this problem, we propose a robot audition system that optimizes parameters adaptively and automatically. Our system applies and non-linear enhancement sub-processes. For the linear sub-process, we used Geometric Source Separation with the Adaptive Step-size method (GSS-AS). This adjusts the parameters adaptively and does not have any manual parameters. For the non-linear sub-process, we applied a spectral subtraction-based enhancement method with the HRLE method that is newly introduced in this paper. Since HRLE controls the threshold level parameter implicitly based on the statistical characteristics of noise and speech levels, our system has high robustness against acoustical environmental changes. For robot audition systems, all processes should be performed in real-time. We also propose implementation techniques to make HRLE run in real-time and show the effectiveness. We evaluate performance of our system and compare it to conventional systems based on the Minima Controlled Recursive Average (MCRA) method and Minimum Mean Square Error (MMSE) method. The experimental results show that our system achieves better performance than the conventional systems. Hirofumi Nakajima, Gökhan Ince, Kazuhiro Nakadai, Yuji Hasegawa |
IROS | 2 |
| 2009 | Ego noise suppression of a robot using template subtractionabstractWhile a robot is moving, the joints inevitably generate noise due to its motors, i.e. ego-motion noise. This problem is very crucial, especially in humanoid robots, because it tends to have a lot of joints and the motors are located closer to the microphones than the sound sources. In this work, we investigate methods for the prediction and suppression of the ego-motion noise. In the first part, we analyze the performance of different noise subtraction strategies, assuming that the noise prediction problem has been solved. In the second part, we present some results for a noise prediction scheme based on the current robot joint status. Performance is evaluated for a number of criteria, including Automatic Speech Recognition (ASR). We demonstrate that our method improves recognition performance during ego-motion considerably. Gökhan Ince, Kazuhiro Nakadai, Tobias Rodemann, Yuji Hasegawa, Hiroshi Tsujino, Jun-ichi Imura |
IROS | 1 |
| 2008 | Using binaural and spectral cues for azimuth and elevation localizationabstractIt is a common assumption that with just two microphones only the azimuth angle of a sound source can be estimated and that a third, orthogonal microphone (or set of microphones) is necessary to estimate the elevation of the source. Recently, using specially designed ears and analyzing spectral cues several researchers managed to estimate sound source elevation with a binaural system. In this work, we show that with two bionic ears both azimuth and elevation angle can be determined using both binaural (e.g. IID and ITD) and spectral cues. This ability can also be used to disambiguate signals coming from the front or back. We present a detailed analysis of both azimuth and elevation localization performance for binaural and spectral cues in comparison. We demonstrate that with a small extension of a standard binaural system a basic elevation estimation capacity can be gained. Tobias Rodemann, Gökhan Ince, Frank Joublin, Christian Goerick |
IROS | 2 |