VLDB 2026 Research / reviewers in the wild / expert
Akira Watanabe
dblp:69/4102
· DBLP profile ↗
17ranked-venue papers
6as first author
3since 2021 · last 2023
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 1 since 2021Computer networks · 3 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Speech recognition and synthesis · 98% Reinforcement learning · 1% Legged, aerial and field robots · 0% | |
| Computer graphics and multimedia
4 papers |
Virtual and augmented reality · 90% Audio and music processing · 10% | |
| Computer networks
1 paper |
Internet architecture and protocols · 70% Cellular and mobile networks · 30% | |
| Human-computer interaction and pervasive computing
2 papers |
Interaction techniques and input · 88% Accessibility and assistive technology · 12% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Virtual and augmented reality
motion-to-photon latency |
0.7 | 1 | 2023 | Low-Latency Beaming Display: Implementation of Wearable, 133 μs Motion-to-Photon Latency Near-Eye Display · IEEE Trans. Vis. Comput. Graph. 2023 |
Virtual and augmented reality
near-eye display |
0.7 | 1 | 2023 | Low-Latency Beaming Display: Implementation of Wearable, 133 μs Motion-to-Photon Latency Near-Eye Display · IEEE Trans. Vis. Comput. Graph. 2023 |
Natural language and speech › Speech recognition and synthesis › speech analysis
acoustic feature analysis |
0.5 | 1 | 2021 | Vocal Tract Length Estimation Using Accumulated Means of Formants and Its Effects on Speaker-Normalization · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Natural language and speech › Speech recognition and synthesis › speech analysis
formant tracking |
0.5 | 1 | 2021 | Vocal Tract Length Estimation Using Accumulated Means of Formants and Its Effects on Speaker-Normalization · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Natural language and speech › Speech recognition and synthesis
speaker normalization |
0.5 | 1 | 2021 | Vocal Tract Length Estimation Using Accumulated Means of Formants and Its Effects on Speaker-Normalization · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Interaction techniques and input › input sensing
tracking |
0.2 | 1 | 2023 | Low-Latency Beaming Display: Implementation of Wearable, 133 μs Motion-to-Photon Latency Near-Eye Display · IEEE Trans. Vis. Comput. Graph. 2023 |
Internet architecture and protocols
end-to-end communication |
0.2 | 1 | 2013 | NTMobile: new end-to-end communication architecture in IPv4 and IPv6 networks · MobiCom 2013 |
Cellular and mobile networks
mobility management |
0.2 | 1 | 2013 | NTMobile: new end-to-end communication architecture in IPv4 and IPv6 networks · MobiCom 2013 |
Internet architecture and protocols › middlebox › middlebox traversal
NAT traversal |
0.2 | 1 | 2013 | NTMobile: new end-to-end communication architecture in IPv4 and IPv6 networks · MobiCom 2013 |
Audio and music processing
speech processing |
0.1 | 3 | 2006 | Reliable methods for estimating relative vocal tract lengths from formant trajectories of common words · IEEE Trans. Speech Audio Process. 2006 Formant estimation method using inverse-filter control · IEEE Trans. Speech Audio Process. 2001 Speech visualization by integrating features for the hearing impaired · IEEE Trans. Speech Audio Process. 2000 |
Audio and music processing › speech analysis
formant tracking |
0.0 | 1 | 2001 | Formant estimation method using inverse-filter control · IEEE Trans. Speech Audio Process. 2001 |
Accessibility and assistive technology
deaf and hard of hearing |
0.0 | 1 | 2000 | Speech visualization by integrating features for the hearing impaired · IEEE Trans. Speech Audio Process. 2000 |
Machine learning › Reinforcement learning
hybrid reinforcement learning |
0.0 | 1 | 1997 | Hybrid Reinforcement Learning and Its Application to Biped Robot Control · NIPS 1997 |
Audio and music processing › speech analysis
spectral envelope estimation |
0.0 | 1 | 2001 | Formant estimation method using inverse-filter control · IEEE Trans. Speech Audio Process. 2001 |
Robotics › Motion planning and robot control › locomotion control › legged robot control
bipedal robot control |
0.0 | 1 | 1997 | Hybrid Reinforcement Learning and Its Application to Biped Robot Control · NIPS 1997 |
Robotics › Legged, aerial and field robots › legged robots
biped robot |
0.0 | 1 | 1997 | Hybrid Reinforcement Learning and Its Application to Biped Robot Control · NIPS 1997 |
Methods — techniques the papers use, named apart from their topics
prototype implementation · 1.3formant trajectory analysis · 0.6inverse-filter control · 0.5time-delay neural network · 0.1dynamic time warping · 0.1zero-crossing frequency distribution · 0.0linear predictive coding · 0.0reinforcement learning · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Low-Latency Beaming Display: Implementation of Wearable, 133 μs Motion-to-Photon Latency Near-Eye DisplayabstractThis paper presents a low-latency Beaming Display system with a 133 μs motion-to-photon (M2P) latency, the delay from head motion to the corresponding image motion. The Beaming Display represents a recent near-eye display paradigm that involves a steerable remote projector and a passive wearable headset. This system aims to overcome typical trade-offs of Optical See-Through Head-Mounted Displays (OST-HMDs), such as weight and computational resources. However, since the Beaming Display projects a small image onto a moving, distant viewpoint, M2P latency significantly affects displacement. To reduce M2P latency, we propose a low-latency Beaming Display system that can be modularized without relying on expensive high-speed devices. In our system, a 2D position sensor, which is placed coaxially on the projector, detects the light from the IR-LED on the headset and generates a differential signal for tracking. An analog closed-loop control of the steering mirror based on this signal continuously projects images onto the headset. We have implemented a proof-of-concept prototype, evaluated the latency and the augmented reality experience through a user-perspective camera, and discussed the limitations and potential improvements of the prototype. Yuichi Hiroi, Akira Watanabe, Yuri Mikawa, Yuta Itoh 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Hovering and Contact Representation of Laser Contour-Based Hand with Swinging Tablet PC for Distant CommunicationabstractAs covid-19 spread worldwide, the growing trend to reduce people’s mobility and the establishment of new lifestyles such as remote work, contactless communication systems have been required. Using a projected hand makes it possible to realize remote communication with a sense of presence without loading a user’s physical body. However, conventional projected hand systems have a problem that the relative brightness of the hand image in a bright living room is low. Therefore, we propose a configuration that combines a video call system and a contour-based projected hand by a laser scanning projector. In addition, we incorporate a haptic feedback system with a tablet PC to improve the sense of contact with remote objects. To enhance the sense of contact in remote areas, we focus on providing the transition between hovering and contact states of the contour-based projected hand and users’ hands. According to the experiment, we confirmed that there is a possibility to express hovering and contact even with simple swinging flat-panel feedback and visual effects of the contour-based projected hand. Akira Watanabe, Takuya Uchida, Daisuke Iwai, Kosuke Sato |
TEI | 1 |
| 2021 | Vocal Tract Length Estimation Using Accumulated Means of Formants and Its Effects on Speaker-NormalizationabstractDifferences in vocal tract lengths (VTLs) in individual speakers cause variations in acoustic features of phonemes. In this paper, a simple method to estimate speaker-specific VTLs and to quantitatively evaluate some speaker-normalization effects of the VTLs is proposed. We employed accumulated means of formant trajectories to estimate the VTLs of speakers ranging from children to adults. For the formant estimation, the inverse-filter control (IFC) system was used. In the system, the decision of analysis order, which means number of formants to be estimated, is automated. Moreover, to evaluate the speaker-normalization effect of VTLs, we proposed the data reduction method, which can reasonably find dense areas of ellipses from distributions in the formant space. Using these ellipse areas, we evaluated the three normalization effects of VTLs: normalization by the mean of all VTLs as the standard, by speaker-categorical means of VTLs, and by individual VTLs. The area reduced from the standard area of the original data by 39.5% and 46.6% in the case of the categorical means and individual VTLs, respectively. As a result, our proposed method was used to provide a “normalized vowel map (NVM)” that visualizes universal vowel-distributions as a core image of linguistic information. Finally, we compared the estimated VTLs with those by another method based on magnetic resonance imaging (MRI) data, using the proposed methods. Tadashi Sakata, Naomitsu Ikeda, Yuichi Ueda, Akira Watanabe |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2018 | Development of certificate based secure communication for mobility and connectivity protocolabstractSome Internet of Things (IoT) devices for smart lock systems, home management systems have been released in recent years. Almost all services employ a client-server model because direct communication among these devices is difficult due to the Network Address Port Translation (NAPT) mechanism. Recently, some new ideas such as a fog computing and an edge computing have been proposed to create a distributed service. A direct communication is suitable for these ideas because each service devices should communicate with each other to realize a distributed processing. The authors have been developed Network Traversal with Mobility (NTMobile) which is a protocol for realizing a direct communication between devices and a secure communication. This paper extends the secure communication mechanism of NTMobile to realize a direct encryption key exchange mechanism between devices by employing the technology of Public Key Infrastructure (PKI). The developed protocol can provide a mutual authentication mechanism and secure encryption key exchange mechanism among devices. The evaluation results show that the overhead of the developed implementation is not large in practical services comparing to the conventional key exchange mechanism of NTMobile. Yuya Miyazaki, Katsuhiro Naito, Hidekazu Suzuki, Akira Watanabe |
CCNC | 4 |
| 2014 | End-to-end IP mobility platform in application layer for iOS and Android OSabstractSmartphones are a new type of mobile devices that users can install additional mobile software easily. In the almost all smartphone applications, client-server model is used because end-to-end communication is prevented by NAT routers. Recently, some smartphone applications provide real time services such as voice and video communication, online games etc. In these applications, end-to-end communication is suitable to reduce transmission delay and achieve efficient network usage. Also, IP mobility and security are important matters. However, the conventional IP mobility mechanisms are not suitable for these applications because most mechanisms are assumed to be installed in OS kernel. We have developed a novel IP mobility mechanism called NTMobile (Network Traversal with Mobility). NTMobile supports end-to-end IP mobility in IPv4 and IPv6 networks, however, it is assumed to be installed in Linux kernel as with other technologies. In this paper, we propose a new type of end-to-end mobility platform that provides end-to-end communication, mobility, and also secure data exchange functions in the application layer for smartphone applications. In the platform, we use NTMobile, which is ported as the application program. Then, we extend NTMobile to be suitable for smartphone devices and to provide secure data exchange. Client applications can achieve secure end-to-end communication and secure data exchange by sharing an encryption key between clients. Users also enjoy IP mobility which is the main function of NTMobile in each application. Finally, we confirmed that the developed module can work on Android system and iOS system. Katsuhiro Naito, Kazuo Mori, Hideo Kobayashi, Kazuma Kamienoo, Hidekazu Suzuki, Akira Watanabe |
CCNC | 6 |
| 2013 | NTMobile: new end-to-end communication architecture in IPv4 and IPv6 networksabstractWith the spread of mobile devices, there is a growing demand for direct communication between users. However, under recent complex IP networks, it is extremely hard to establish a direct connection between devices. In order to solve the problem, authors have been proposing a new end-to-end communication architecture called "Network Traversal with Mobility" (NTMobile). In NTMobile, applications in the mobile device establish an end-to-end connection by using virtual IPv4/IPv6 addresses in the NTMobile network independent from real IP networks. In this demo, we will show that Android smartphones can make free communication with each other without any constraint such as an NAT traversal problem in IPv4 networks and incompatibility between IPv4 and IPv6 architectures. Hidekazu Suzuki, Katsuhiro Naito, Kazuma Kamienoo, Tatsuya Hirose, Akira Watanabe |
MobiCom | 5 |
| 2012 | Proposal of seamless IP mobility schemes: Network traversal with mobility (NTMobile)abstractIP mobility is the technology that realizes a continuous communication even when communicating wireless nodes switch access networks. NTMobile (Network Traversal with Mobility) is a new type of IP mobility technology for the IPv4 and IPv6 networks. NTMobile nodes can also achieve IP mobility in global IP networks as well as in private IP networks by creating a tunnel route between a pair of NTMobile nodes. In NTMobile, network applications use virtual IP addresses to achieve a continuous communication at the time when real IP addresses change due to the switching of access networks or node movements. In NTMobile, we implement a packet manipulation method in the Linux kernel module so as to achieve high levels of throughput performance. Katsuhiro Naito, Takuya Nishio, Kazuo Mori, Hideo Kobayashi, Kazuma Kamienoo, Hidekazu Suzuki, Akira Watanabe |
GLOBECOM | 7 |
| 2006 | Reliable methods for estimating relative vocal tract lengths from formant trajectories of common wordsabstractThis paper describes reliable methods for estimating relative vocal tract lengths from speech signals. Two proposed methods are based on the simple principle that resonant frequencies in an acoustic tube are inversely proportional to the tube length in cases where the configuration is constant. We estimated the ratio between two speakers' vocal tract lengths using first and second formant trajectories of the same words uttered by them. In the first method, which is referred to as "strict estimation method", we sought instances at which the gross structures of two vocal tracts are analogous by applying dynamic time-warping to formant-trajectories of common words that were uttered at different speeds. In those instances, which were found from among more than 100 common words by two speakers, an average formant ratio proved to be an excellent estimate (about plusmn0.1% in errors) for a reciprocal of the vocal tract length ratio. Next, we examined a simplified method for estimating those ratios using all corresponding points of two formant-trajectories: it is the "direct estimation method". Estimation errors in the direct estimation were evaluated to be about plusmn0.3% at equal utterance-speeds and plusmn2% at most, within 2.0 of the ratios of "fast" to "slow". Finally, we estimated relative vocal tract lengths for four Japanese speaker groups, whose members differed in terms of age and gender. Experimental results showed that the average vocal tract length of adult females and that of 7-10-year-old boys and girls are 21%, 27%, and 30%, respectively, shorter than adult males' Akira Watanabe, Tadashi Sakata |
IEEE Trans. Speech Audio Process. | 1 |
| 2001 | An image sensor with fast extraction of objects' positions - rough vision processorabstractAn integration of the signal processing circuits with the image acquiring device, which is called the vision chip and can process information in parallel, is proposed for fast image processing. In applications for robot vision, not only the detailed information, such as shape or texture, but also the rough information, such as 'something is around here', are important and useful. We consider the detecting of centroids of objects in the focal plain as the rough vision processing, which is useful in practical application, and describe its implementation using two components; the centroid detector and the coordinate generator. First, we describe the fast flag generation algorithm indicating the centroid of objects, and its implementation using an analog parallel signal processing architecture. Next, we describe a novel encoding algorithm for flag positions indicating the centroids in order to obtain their coordinates. Akira Watanabe, Osamu Tooyama, Masayuki Miyama, Masahiko Yoshimoto, Junichi Akita |
ICIP (2) | 1 |
| 2001 | Formant estimation method using inverse-filter controlabstractThis paper proposes a new method for estimating formant frequencies of speech signals, based on inverse-filter control and zero-crossing frequency distributions. In this method, which is called the inverse-filter control (IFC) method, we use 32 basic inverse filters that are mutually controlled by weighted means of zero-crossing frequency distributions. After quick convergence of the inverse filters, we can gain four to six formant frequencies as final mean-values of the zero-crossing frequencies. The proposed method (IFC) has a specific feature that it directly estimates resonant frequencies of a vocal tract, unlike analysis-by-synthesis (A-b-S) or linear predictive coding (LPC) as a spectral matching method. Therefore, spectral shapes influence indirectly alone the formant estimation in the IFC. Although the superiority of IFC to LPC was not necessarily prominent in the systematic evaluation using synthetic speech, the estimates showed satisfactorily small errors for the practical analysis. On the other hand, when observing some analysis examples of real speech, we found many fewer gross errors in IFC than in LPC. Last, we describe in brief a method for estimating a spectral envelope (or formant bandwidths) based on the obtained formant frequencies and the spectrum to be analyzed. According to the results, it is understandable that the existence of the wide-band formants also contributes to stable formant trajectories. Akira Watanabe |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Proposal of Group Search Protocol Making Secure Communication Groups for IntranetabstractIn the intranet, there are demands for making secure communication groups that protect significant information in departmental and/or individual bases. To satisfy the demands, encryption elements in the system need to have exact and complicated process information. In general, this information is generated in a management server and downloaded to the encryption elements, every time the system configuration changes. We a propose group search protocol, by which encryption elements generate the process information automatically from the logical definition of the secure communication groups for themselves. By this method, secure communication groups for intranet with location freedom is realized, that is to say, users who have the function of encryption elements can move anywhere in the intranet freely. Akira Watanabe, Toru Inada, Tetsuo Ideguchi, Iwao Sasase |
ICC (2) | 1 |
| 2000 | Speech visualization by integrating features for the hearing impairedabstractDescribes development of a new speech visualization system that creates readable patterns by integrating different speech features into a single picture. The system extracts the phonemic and prosodic features from speech signals and converts them into a visual image using neither speech segmentation nor speech recognition. We used four time-delay neural networks (TDNNs) to generate phonemic features in the new system. Training of the TDNNs using three selected frames of eight kinds of acoustic parameters showed significant improvement in the performance. The TDNN outputs control the brightness of patterns used for consonants, that is, each of the consonant-patterns is represented by a different white texture whose brightness is weighted by the output of a corresponding TDNN. All the weighted consonant-patterns are simply added and then overlaid synchronously on colors due to the formant frequencies. When this is done, phonemic sequences and boundaries manifest themselves in the resulting visual patterns. In addition, the color of a single vowel sandwiched between consonants looks uniform. These visual phenomena are very useful for decoding the complex speech code, which is generated by the continuous movements of speech organs. We evaluated the visualized speech in a preliminary test. When three students read the patterns of 75 words uttered by four males (300 items), the learning curves showed a steep rise and the correct answer rate reached 96-99%. The learning effect was durable: after five months of absence from the system, a subject read 96.3% of the 300 tokens in a response time which averaged only 1.3 s/word. Akira Watanabe, Shingo Tomishige, Masahiro Nakatake |
IEEE Trans. Speech Audio Process. | 1 |
| 1997 | Hybrid Reinforcement Learning and Its Application to Biped Robot Control
Satoshi Yamada, Akira Watanabe, Michio Nakashima |
NIPS | 2 |
| 1994 | A Method to Interpret 3D Motion Using Neural NetworksabstractThis study proposes a 3D motion interpretation method which uses a neural network system consisting of three kinds of neural networks. This system estimates the solutions of 3D motion of an object by interpreting three optical flow (OF - motion vector field calculated from images) patterns of the same object obtained from three different view points. Though the interpretation system is trained using only basic 3D motions consisting of a single motion component, the system can interpret unknown multiple 3D motions consisting of several motion components. The generalization capacity of the proposed system is confirmed using diverse test patterns. Also the robustness of the system to noise is proved experimentally. The experimental results show that this method has suitable features for applying to real images.> Arata Miyauchi, Akira Watanabe, Minami Miyauchi |
ICIP (3) | 2 |
| 1994 | A hearing aid by single resonant analysis for telephonic speech
Takashi Ikeda, Kouji Tasaki, Akira Watanabe |
ICSLP | 3 |
| 1994 | A DSP-based amplitude compressor for digital hearing AIDS
Yuichi Ueda, Takayuki Agawa, Akira Watanabe |
ICSLP | 3 |
| 1986 | Assessment of electronic aids for the hearing impaired which transmit visible or tactile speechabstractMany electronic aids for the hearing impaired have been developed for the purpose of transmitting information as a visible or tactile speech. However, since the performance of such aids used to be individually evaluated with some restricted materials uttered by one or a few speakers, a variety of the results have been obtained even in almost the same kinds of aids. In order to investigate the most useful information to be transmitted, a common method in the assessment should be established. We propose, in this paper, a new method which uses synthetic speech (isolated and connected vowels) as speech materials. In this method, similarity between images obtained with the aids and with the normal auditory sense is adopted as an important factor for the assessment. The similarity is estimated by labelling test using the systematically changed stimuli sets. This method was applied to the four kinds of visual/tactual aids. As a result, the goodness of vowel information transmitted by each aid manifested itself in a general sense. Yuichi Ueda, Akira Watanabe |
ICASSP | 2 |