Makoto Kumon

dblp:94/993 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
5since 2021 · last 2025
0000-0003-4278-829XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 20 · 9 first-author · 5 since 2021Systems, architecture and hardware · 19 · 9 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2025 Swarm Active Audition with Robots and Drones: Real-World Performance Validation
abstract
Search and rescue (SAR) operations in large-scale disaster sites, such as areas affected by earthquakes, require rapid victim detection. While drones equipped with cameras are commonly used for SAR, their effectiveness is limited in visually obstructed environments, because of debris, smoke, or fog. Under such situations, auditory information can play a crucial role in locating victims who are not visible. Existing drone audition research has demonstrated the feasibility of detecting sound sources using onboard microphone arrays. However, most studies focus on single-drone systems, which face limitations in coverage and accessibility, particularly in complex environments such as collapsed buildings or urban canyons. Additionally, real-world validation of multi-drone audition systems remains limited, with prior studies relying primarily on simulations or controlled environments. To address these challenges, we propose and evaluate a Multi-Drone and Robot-Based Active Audition System (SAAS-RD: Swarm Active Audition System with Robots and Drones) that integrates multiple drones and ground robots to enhance acoustic search capabilities. Our work focuses on real-world performance validation, conducting field experiments in outdoor environments and analyzing system feasibility through case studies. The results demonstrate the potential of SAAS-RD as a practical solution for large-scale SAR operations.
Kazuhiro Nakadai, Kotaro Hoshiba, Benjamin Yen 0001, Makoto Kumon, Yoko Sasaki
IROS4
2022 A Data-Driven Multiple Model Framework for Intention Estimation
abstract
This paper presents a data-driven multiple model framework for estimating the intention of a target from observations. Multiple model (MM) state estimation methods have been extensively used for intention estimation by mapping one intention to one dynamic model assuming one-to-one relations. However, intentions are subjective to humans and it is difficult to establish the one-to-one relations explicitly. The proposed framework infers the multiple-to-multiple relations between intentions and models directly from observations that are labeled with intentions. For intention estimation, both the relations and model probabilities of an Interacting Multiple Model (IMM) state estimation approach are integrated into a recursive Bayesian framework. Taking advantage of the inferred multiple-to-multiple relations, the framework incorpo-rates more accurate relations and avoids following the strict one-to-one relations. Numerical and real experiments were performed to investigate the framework through the intention estimation of a maneuvered quadrotor. Results show higher estimation accuracy and superior flexibility in designing mod-els over the conventional approach that assumes one-to-one relations.
Yongming Qin, Makoto Kumon, Tomonari Furukawa
ICRA2
2022 Fast Scan Context Matching for Omnidirectional 3D Scan
abstract
Autonomous robots need to recognize the environment by identifying the scene. Scan context is one of global descriptors, and it encodes the three-dimensional scan data of the scene for the identification in a matrix form. Scan context is in a matrix form that is simple to store, but the matching of scan contexts can require computational effort because the descriptor is orientation-dependent. Because a scan context of an omnidirectional LiDAR scan becomes periodic in azimuth, this paper proposes to compute the scan context matching efficiently incorporating the cross-correlation with fast Fourier transform, and, hence, the method is named fast scan context matching. The effectiveness of the proposed method for computation time, accuracy, and robustness are reported in this paper. It is also shown that the method was also tested as a loop closure detector of a SLAM package as a practical application and that the proposed method outperformed the conventional scan context matching.
Hikaru Kihara, Makoto Kumon, Kei Nakatsuma, Tomonari Furukawa
IROS2
2022 Object Surface Recognition using Microphone Array by Acoustic Standing Wave
abstract
This paper proposes a microphone array with a speaker to recognize the shape of the surface of the target object by using the standing wave between the transmitted and the reflected acoustic signals. Because the profile of the distance spectrum encodes both the distance to the target and the distance to the edges of the target's surface, this paper proposes to fuse distance spectra using a microphone array to estimate the three-dimensional structure of the target surface. The proposed approach was verified through numerical simulations and outdoor field experiments. Results showed the effectiveness of the method as it could extract the shape of the board located 2m in front of the microphone array by using a chirp tone with 20kHz bandwidth.
Tomoya Manabe, Rikuto Fukunaga, Kei Nakatsuma, Makoto Kumon
IROS4
2021 Alternating Drive-and-Glide Flight Navigation of a Kiteplane for Sound Source Position Estimation
abstract
Drone audition, namely the hearing capability of a drone, is expected to compensate for the drawbacks of visual sensors in search-and-rescue missions. Current multi-rotor drones have limitations of flight duration and sound processing due to ego-noise generated by rotors and air-flow. Drone audition for a kiteplane, i.e., a fixed-wing drone that can fly slowly and stably, has not been investigated. This paper proposes "Alternating Drive-and-Glide Flight Navigation"(AltDGFNavi) of a kiteplane for sound source position estimation. AltDGFNavi consists of two functions: periodical switching rotor for driving and gliding to reduce ego-noise, and dynamic flight path generation to fly close to the target. AltDGFNavi was evaluated through numerical simulations, and the results of sound source position estimation demonstrated the effectiveness of AltDGFNavi.
Makoto Kumon, Hiroshi G. Okuno, Shuichi Tajima
IROS1
2019 Belief-Driven Control Policy of a Drone with Microphones for Multiple Sound Source Search
abstract
This paper proposes a belief-driven control policy of a drone with microphones for multiple sound source search. As the sound source localization by drones is uncertain because of the observation significantly distorted by noise such as rotor noise, the belief on the estimated targets may consist of multiple peaks that are spread over the bounded search area. The proposed control policy is formulated with a robust cost function so that the function encodes the search mission properly. A peak management mechanism is additionally introduced to keep tracking all targets by masking sufficiently observed and well estimated targets whose peaks normally become steep and high. The proposed control policy was evaluated by numerical simulations, and experiments, and those results have validated the efficacy of the proposed control policy.
Kenshiro Yamada, Makoto Kumon, Tomonari Furukawa
IROS2
2017 Development of microphone-array-embedded UAV for search and rescue task
abstract
This paper addresses online outdoor sound source localization using a microphone array embedded in an unmanned aerial vehicle (UAV). In addition to sound source localization, sound source enhancement and robust communication method are also described. This system is one instance of deployment of our continuously developing open source software for robot audition called HARK (Honda Research Institute Japan Audition for Robots with Kyoto University). To improve the robustness against outdoor acoustic noise, we propose to combine two sound source localization methods based on MUSIC (multiple signal classification) to cope with trade-off between latency and noise robustness. The standard Eigenvalue decomposition based MUSIC (SEVD-MUSIC) has smaller latency but less noise robustness, whereas the incremental generalized singular value decomposition based MUSIC (iGSVD-MUSIC) has higher noise robustness but larger latency. A UAV operator can use an appropriate method according to the situation. A sound enhancement method called online robust principal component analysis (ORPCA) enables the operator to detect a target sound source more easily. To improve the stability of wireless communication, and robustness of the UAV system against weather changes, we developed data compression based on free lossless audio codec (FLAC) extended to support a 16 ch audio data stream via UDP, and developed a water-resistant microphone array. The resulting system successfully worked in an outdoor search and rescue task in ImPACT Tough Robotics Challenge in November 2016.
Kazuhiro Nakadai, Makoto Kumon, Hiroshi G. Okuno, Kotaro Hoshiba, Mizuho Wakabayashi, Kai Washizaki, Takahiro Ishiki, Daniel Gabriel, Yoshiaki Bando, Takayuki Morito, Ryosuke Kojima, Osamu Sugiyama
IROS2
2016 Recursive Bayesian estimation of NFOV target using diffraction and reflection signals
Kuya Takami, Hangxin Liu, Makoto Kumon, Tomonari Furukawa, Gamini Dissanayake
FUSION3
2016 Non-field-of-view sound source localization using diffraction and reflection signals
abstract
This paper describes a non-field-of-view (NFOV) localization approach for a mobile robot in an unknown environment based on an acoustic signal combined with the geometrical information from an optical sensor. The approach estimates the location of a target through the mobile robot's sensor observation frame, which consists of a combination of diffraction and reflection acoustic signals and a 3-D environment geometrical description. This fusion of audio-visual sensor observation likelihoods allows the robot to estimate the NFOV target. The diffraction and reflection observations from the microphone array generate the acoustic joint observation likelihood. The observed geometry also determines far-field or near-field acoustic conditions to improve the estimation of the sound direction of arrival. A mobile robot equipped with a microphone array and an RGB-D sensor was tested in a controlled environment, an anechoic chamber, to demonstrate the NFOV localization capabilities. This resulted in +/- 18 degrees, and less than 0.75 m error in angle and distance estimation, respectively.
Kuya Takami, Hangxin Liu, Tomonari Furukawa, Makoto Kumon, Gamini Dissanayake
IROS4
2016 Position estimation of sound source on ground by multirotor helicopter with microphone array
abstract
Multirotor helicopters are expected to be utilized various tasks including rescue missions and surveillance. For those missions, sensors are equipped with helicopters in order to recognize the environment, and auditory information is one of such information that can be utilized to find the target sound source even if it is occluded by objects. One of the difficulty comes from the fact that the noise generated by rotating rotors distorts the target signal significantly, which leads inaccurate estimate of the target source direction. Besides, the estimate of the attitude and the position of the helicopter is inaccurate, which is another technical issue. This paper proposes a robust sound source position estimation taking such uncertainty into account, and experiment with real flight tests showed that the helicopter was able to estimate the target position within about 2m accuracy.
Kai Washizaki, Mizuho Wakabayashi, Makoto Kumon
IROS3
2015 Design model of microphone arrays for multirotor helicopters
abstract
Unmanned multirotor helicopters have been expected for various critical tasks such as rescue missions, and it is important to sense the environment to achieve those tasks. In order to realize such sensor systems, this paper considers acoustic information recognition from a hovering helicopter which a microphone array system. As the noise of rotors distorts the acoustic information, noise reduction is necessary. Since the rotor noise varies even at the hovering flight, an acoustic model of the rotor noise is derived taking the dynamics of the helicopter into account, and the model is utilized to evaluate the performance of the microphone array. The proposed approach was verified by evaluating the optimality of a real device that was empirically tuned in authors previous work, and the computed configuration of the microphone array and the developed one coincide well, which ensures the validity of the approach. As the validation through a practical application, the signal was processed by Delay-and-Sum Beam Former to localize the sound source. The system was able to find peaks of the power that corresponded to the sound source.
Takahiro Ishiki, Makoto Kumon
IROS2
2013 Bayesian non-field-of-view target estimation incorporating an acoustic sensor
abstract
This paper presents non-field-of-view (NFOV) target estimation incorporating an acoustic sensor, which consists of two microphones. The proposed approach derives the interaural level difference (ILD) of observations from the two microphones for different target positions and stores the ILDs as database a priori. Given a new acoustic observation on a target, an acoustic observation likelihood is created by calculating the correlation of the ILD of the new observation to the stored ILDs. A joint observation likelihood is then developed by fusing the optical and acoustic observation likelihoods, and the recursive Bayesian estimation updates and maintains belief on the target using the joint observation likelihood. The proposed approach detects a target positively using an acoustic sensor even if it is outside the field of view of the optical sensor and localizes the target accurately by estimating it within the RBE. The efficacy of the proposed approach was first validated by experimental studies. Further numerical demonstrations then show the applicability of the proposed approach to the NFOV target estimation.
Makoto Kumon, Daisuke Kimoto, Kuya Takami, Tomonari Furukawa
IROS1
2011 Active soft pinnae for robots
abstract
Sound localization is one of fundamental abilities of auditory robots that is necessary for human-machine communication. Especially for autonomous robots, it is also necessary to make the system as simple as possible, and binaural configuration is thought to be the minimum setup for auditory systems as animals achieve practical hearing with two ears. However, the limitation of the number of ears leads insufficient performance such as accuracy in sound localization, or sound source separation. In order to overcome those difficulties, this paper proposes a novel artificial soft pinna that can move, or deform its shape actively. The developed pinna has deformable skin of silicon rubber with wires controlled by motors, and it is covered by a fur cover. Considering applications of the proposed device for auditory robots, fundamental characteristics such as accuracy of motion and auditory characteristics were investigated through experiments, which is the main contribution of the paper.
Makoto Kumon, Yoshitaka Noda
IROS1
2010 Hovering control of vectored thrust aerial vehicles
abstract
In this paper, a vectored thrust aerial vehicle( VTAV) that has three ducted fans is considered. Since ducted fans are powerful and effective in providing lift, they are suitable for thrusters of UAVs, but modeling their aerodynamic effects such as ram drag is very difficult. The VTAV has one ducted fan fixed to its body and two ducted fans that can be tilted in order to make rotational moments, which makes the system dynamics even more complicated. This paper focuses on giving a precise dynamical model that includes aerodynamic effects of ducted fans. Then it presents a hovering controller based on the dynamical model developed. Since the horizontal dynamics are under actuated, a switching control approach is introduced to realize a stabilizing controller. Numerical simulations prove the validity of the approach.
Makoto Kumon, Hugh Cover, Jayantha Katupitiya
ICRA1
2010 Motion planning based on simultaneous perturbation stochastic approximation for mobile auditory robots
abstract
In this paper, a motion planning method for mobile auditory robots is proposed based on an optimization technique. Since it is one of the most important abilities for auditory robots to recognize vocal messages correctly, the proposed method is designed to maximize the confidence measure of a speech recognition since the measure is thought to be strongly related to the accuracy of the speech recognition. However, the cost function to optimize is hard to model explicitly, and it is difficult to obtain the gradient that is normally utilized to derive the motion. In order to overcome this difficulty, simultaneous perturbation stochastic approximation(SPSA) that does not require an explicit model of the cost function is applied to generate robot motion. The effectiveness of the approach was verified through real experiments: the robot could get better speech recognition rate after it approached the sound source by measuring the confidence measure.
Makoto Kumon, Keiichiro Fukushima, Sadaaki Kunimatsu, Mitsuaki Ishitobi
IROS1
2006 Robust Adaptive Output Feedback Control of MIMO Systems Using Multirate Sampling
abstract
In adaptive output feedback control based on almost strictly positive real (ASPR) conditions, a technical difficulty arises when the controlled multi-input multi-output (MIMO) system is non-square. To overcome this, the idea of multirate sampled-data control has been proposed. That is, through careful choice of faster input sampling rates create a lifted discrete-time system which has the same number of inputs and outputs and does not give rise to the causality constraint. The output feedback based adaptive control strategy can then be applied to this lifted system under certain conditions. In this report, we propose a robust adaptive controller design scheme for non-square MIMO systems using the multirate sampling strategy without the causality problem
Ikuro Mizumoto, Satoshi Ohdaira, Makoto Kumon, Zenta Iwai
ICARCV3
2006 Spectral Cues for Robust Sound Localization with Pinnae
abstract
An important ability in auditory robots is the localization of sound source. In this paper, a robust method to localize the vertical direction of the sound source with two microphones and pinnae is proposed. In order to achieve vertical sound source localization, the method of detecting spectral cues, the relationship between spectral cues and source direction are studied. In addition, this paper considers sound source separation in order to recognize and isolate spectral cues only from the sound source in order to make the method robust to extraneous noise. Furthermore the authors designed an audio servo with the proposed spectral cues and implemented the method in an actual robot. The experimental results confirmed the effectiveness of the method
Tomoko Shimoda, Toru Nakashima, Makoto Kumon, Ryuichi Kohzawa, Ikuro Mizumoto, Zenta Iwai
IROS3
2005 Wind Estimation by Unmanned Air Vehicle with Delta Wing
abstract
In this paper, an algorithm to estimate wind direction by using a small and light Unmanned Air Vehicle(UAV) called KITEPLANE was proposed. KITEPLANE had a big main wing which is a kite-like delta shape and, therefore, it was easy to be disturbed by wind. However, this disadvantage implies that the KITEPLANE has an ability to sense wind and that it is expected to use the KITEPLANE as a sensor for wind estimation. In order to achieve this feature, dynamics of the KITEPLANE under wind disturbance were derived and a numerical estimation method was proposed. Devices equipped on board were also developed and the proposed method was implemented. Results of an experiment showed the effectiveness of the proposed method.
Makoto Kumon, Ikuro Mizumoto, Zenta Iwai, Masanobu Nagata
ICRA1
2005 Audio servo for robotic systems with pinnae
abstract
Sound localization is one of important abilities for robots which utilize auditory information. Binaural information such as interaural time difference (ITD) or interaural intensity difference(IID) is effective for horizontal sound localization. In order to extend this ability to vertical sound localization, spectral cues were utilized in this paper instead of increasing number of microphones. Spectral cues were the information about the vertical position of the sound source on frequency domain. Since spectral cues strongly depended on the shape of pinnae, artificial pinnae for auditory robots were designed based on diffraction-reflection model. A method to detect spectral cues was also developed. By utilizing them, a controller which made the robot orient to the sound source was proposed. Results of an experiment showed the validity of the proposed method.
Makoto Kumon, Tomoko Shimoda, Ryuichi Kohzawa, Ikuro Mizumoto, Zenta Iwai
IROS1
2003 Adaptive audio servo for multirate robot systems
abstract
In this paper, audio servo robot system which was a servo system using audio information as feedback signal was considered. Since the sampling rate of audio signal and that of robot controller are usually different, the system is able to be treated as a multi-rate system. Since it is general that the sampling rate of the output is faster than that of the input, the controller obtains more information than a single-rate servo system. By using this property, an adaptive audio servo system was proposed. Results of a computer simulation and an experiment showed effectiveness of the method.
Makoto Kumon, Takahiro Sugawara, Katuhiro Miike, Ikuro Mizumoto, Zenta Iwai
IROS1
2001 Biped gait synthesis based on dynamic parametrization
abstract
Biped locomotion has two characteristics: 1) repeating impact at the moment when the foot contacts with the terrain, 2) limitation to the torque of the ankle. The second is required in order to make the foot of the supporting leg stay on the terrain. Since the impulsive effect can be a disturbance which makes the robot fall down, the controller of the biped robot is required to make the undesired effect of the impulsive phenomenon as small as possible. Although it is common to design the controller with a high-gain feedback in order to eliminate the disturbance, the torque of the ankle is not limited. In the paper, a control method is proposed so that the system with impact is stabilized with limited inputs. A simple biped model controlled by the proposed controller is simulated numerically and the result shows the efficiency of the method.
Makoto Kumon, Takayuki Nakata, Norihiko Adachi
IROS1