VLDB 2026 Research / reviewers in the wild / expert
Yuki Okamoto
dblp:27/7735
· DBLP profile ↗
10ranked-venue papers
4as first author
7since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
Yusuke Kanamori, Yuki Okamoto, Taisei Takano, Shinnosuke Takamichi, Yuki Saito 0001, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2024 | Environmental Sound Synthesis from Vocal Imitations and Sound Event LabelsabstractOne way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as rhythm and pitch, using vocal imitations, which cannot be expressed by conventional input information, such as sound event labels, images, or texts, in an environmental sound synthesis model. In this paper, we propose a framework for environmental sound synthesis from vocal imitations and sound event labels based on a framework of a vector quantized encoder and the Tacotron2 decoder. Using vocal imitations is expected to control the pitch and rhythm of the synthesized sound, which only sound event labels cannot control. Our objective and subjective experimental results show that vocal imitations effectively control the pitch and rhythm of synthesized sounds. Yuki Okamoto, Keisuke Imoto, Shinnosuke Takamichi, Ryotaro Nagase, Takahiro Fukumori, Yoichi Yamashita |
ICASSP | 1 |
| 2023 | Visual Onoma-to-Wave: Environmental Sound Synthesis from Visual Onomatopoeias and Sound-Source ImagesabstractWe propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective, onoma-to-wave has been proposed to synthesize environmental sounds from the desired onomatopoeia texts. Onomatopoeias have another representation: visual-text representations of sounds in comics, advertisements, and virtual reality. A visual onomatopoeia (visual text of onomatopoeia) contains rich information that is not present in the text, such as a long-short duration of the image, so the use of this representation is expected to synthesize diverse sounds. Therefore, we propose visual onoma-to-wave for environmental sound synthesis from visual onomatopoeia. The method can transfer visual concepts of the visual text and sound-source image to the synthesized sound. We also propose a data augmentation method focusing on the repetition of onomatopoeias to enhance the performance of our method. An experimental evaluation shows that the methods can synthesize diverse environmental sounds from visual text and sound-source images. Hien Ohnaka, Shinnosuke Takamichi, Keisuke Imoto, Yuki Okamoto, Kazuki Fujii, Hiroshi Saruwatari |
ICASSP | 4 |
| 2023 | CAPTDURE: Captioned Sound Dataset of Single Sources
Yuki Okamoto, Kanta Shimonishi, Keisuke Imoto, Kota Dohi, Shota Horiguchi, Yohei Kawaguchi |
INTERSPEECH | 1 |
| 2022 | Environmental Sound Extraction Using Onomatopoeic WordsabstractAn onomatopoeic word, which is a character sequence that phonetically imitates a sound, is effective in expressing characteristics of sound such as duration, pitch, and timbre. We propose an environmental-sound-extraction method using onomatopoeic words to specify the target sound to be extracted. By this method, we estimate a time-frequency mask from an input mixture spectrogram and an onomatopoeic word using a U-Net architecture, then extract the corresponding target sound by masking the spectrogram. Experimental results indicate that the proposed method can extract only the target sound corresponding to the onomatopoeic word and performs better than conventional methods that use sound-event classes to specify the target sound. Yuki Okamoto, Shota Horiguchi, Masaaki Yamamoto, Keisuke Imoto, Yohei Kawaguchi |
ICASSP | 1 |
| 2022 | Sound Event Detection Guided by Semantic Contexts of ScenesabstractSome studies have revealed that contexts of scenes (e.g., "home," "office," and "cooking") are advantageous for sound event detection (SED). Mobile devices and sensing technologies give useful information on scenes for SED without the use of acoustic signals. However, conventional methods can employ pre-defined contexts in inference stages but not undefined contexts. This is because one-hot representations of pre-defined scenes are exploited as prior contexts for such conventional methods. To alleviate this problem, we propose scene-informed SED where pre-defined scene-agnostic contexts are available for more accurate SED. In the proposed method, pre-trained large-scale language models are utilized, which enables SED models to employ unseen semantic contexts of scenes in inference stages. Moreover, we investigated the extent to which the semantic representation of scene contexts is useful for SED. Experimental results performed with TUT Sound Events 2016/2017 and TUT Acoustic Scenes 2016/2017 datasets show that the proposed method improves micro and macro F-scores by 4.34 and 3.13 percentage points compared with conventional Conformer- and CNN– BiGRU-based SED, respectively. Noriyuki Tonami, Keisuke Imoto, Ryotaro Nagase, Yuki Okamoto, Takahiro Fukumori, Yoichi Yamashita |
ICASSP | 4 |
| 2021 | Sound Event Detection Based on Curriculum Learning Considering Learning Difficulty of EventsabstractIn conventional sound event detection (SED) models, two types of events, namely, those that are present and those that do not occur in an acoustic scene, are regarded as the same type of the events. The conventional SED methods cannot effectively exploit the difference between the two types of events. The all time frames of sound events that do not occur in an acoustic scene are easily regarded as inactive in the scene, that is, the events are easy-to-train. The time frames of the events that are present in a scene must be classified as active in addition to inactive in the acoustic scene, that is, the events are difficult-to-train. To take advantage of the training difficulty, we apply curriculum learning into SED, where models are trained from easy- to difficult-to-train events. To utilize the curriculum learning, we propose a new objective function for SED, wherein the events are trained from easy-to difficult-to-train events. Experimental results show that the F-score of the proposed method is improved by 10.09 percentage points compared with that of the conventional binary cross entropy-based SED. Noriyuki Tonami, Keisuke Imoto, Yuki Okamoto, Takahiro Fukumori, Yoichi Yamashita |
ICASSP | 3 |
| 2017 | Subthreshold Operation of CAAC-IGZO FPGA by Overdriving of Programmable Routing Switch and Programmable Power SwitchabstractA field-programmable gate array (FPGA) using a crystalline oxide semiconductor of c-axis-aligned crystal indium-gallium-zinc oxide (CAAC-IGZO) has been developed, which is capable of subthreshold operation used for energy harvesting. To achieve subthreshold operation, the CAAC-IGZO FPGA has a structure designed as an extension of a boosting pass gate using a CAAC-IGZO FET and employs overdriving of a programmable routing switch and a programmable power switch for power gating (PG). A CAAC-IGZO FET is used to give an ideal floating gate with excellent charge retention. A chip fabricated using a 0.8-μm CAAC-IGZO/0.18-μm CMOS hybrid process achieves subthreshold operation while maintaining the features required for normally off computing proposed in our previous study. Specifically, these features are realized by fine-grained PG for individual programmable logic elements (PLEs), fast configuration switching between contexts, and load/store between a volatile register and a nonvolatile shadow register in the PLEs. The chip operation at a minimum operating voltage of 180 mV with a combinational circuit configuration is demonstrated. With a sequential circuit configuration, the chip operates at a minimum operating voltage of 190 mV with 12.5 kHz, and the minimum power-delay product is 3.40 pJ/operation at 330 mV. Munehiro Kozuma, Yuki Okamoto, Takashi Nakagawa, Takeshi Aoki, Yoshiyuki Kurokawa, Takayuki Ikeda, Yoshinori Ieda, Naoto Yamade, Hidekazu Miyairi, Masahiro Fujita 0004, Shunpei Yamazaki |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | A Boosting Pass Gate With Improved Switching Characteristics and No Overdriving for Programmable Routing Switch Based on Crystalline In-Ga-Zn-O TechnologyabstractA boosting pass gate (BPG) suitable for a programmable routing switch including a$c$-axis aligned crystal In-Ga-Zn-O (CAAC-IGZO) field effect transistor (FET) is proposed. The CAAC-IGZO is one of crystalline oxide semiconductors (OS). The proposed BPG (OS-based BPG, OS BPG) has a combination of a pass gate (PG) and a configuration memory (CM) cell utilizing a CAAC-IGZO FET with extremely low OFF-state current and a storage capacitor. This OS BPG achieves a routing switch with fewer transistors than a conventional routing switch having a combination of a PG and an static RAM (SRAM) cell. Owing to the boosting effect, the switching characteristics, at not only positive transition but also negative transition of input signals, of the OS BPG are improved without using overdriving. In circuits fabricated with a hybrid process of a CMOSFET and a CAAC-IGZO FET with gate lengths of 0.5 and 1.0$\mu $m, the net delays of the OS BPG, 75 and 58 ns, at driving voltages of 2.0 and 2.5 V have been found to be less than those of the conventional routing switch (SRAM-based PG, SRAM PG) by about 79% and 62%, respectively. It has also been confirmed that a field-programmable gate array (FPGA) chip utilizing the OS BPG as a routing switch reduces the layout areas of routing switches and the whole chip by 61% and 22%, respectively, and increases the maximum operating frequencies at driving voltage of 2.0 and 2.5 V by about 2.8 times and 1.6 times of those of the FPGA chip utilizing the SRAM PG as a routing switch. Yuki Okamoto, Takashi Nakagawa, Takeshi Aoki, Masataka Ikeda, Munehiro Kozuma, Takeshi Osada, Yoshiyuki Kurokawa, Takayuki Ikeda, Naoto Yamade, Yutaka Okazaki, Hidekazu Miyairi, Jun Koyama, Shunpei Yamazaki |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | Vision-based detection of finger touch for haptic device using transparent flexible sheetabstractA haptic device using transparent flexible sheet, which is an improvement of our previous one, presents visual and haptic sensations when a user pushes a virtual soft object. It consists of a transparent flexible sheet and a computer display; the display is placed behind the sheet. The user feels the softness of the object by pushing the sheet, and sees the stereoscopic CG image of the deformed object through the sheet. In the present study we propose a new method of detecting finger touch on the sheet and measuring the indentation of the contact point; the method is specialized for this haptic device. A camera is placed behind the sheet. A checkered pattern is reflected on the backside of the sheet like a two-way mirror. The camera monitors this reflected pattern. When the finger pushes the sheet, the sheet is deformed. Then the reflected pattern near the fingertip, seen from the camera, is distorted. Monitoring this distortion makes it possible to detect finger touch on the sheet. As the indentation of the contact point increases, the reflected pattern is more distorted. Thus measuring the amount of the distortion allows the measurement of the indentation. The proposed method is ascertained by some experiments. Kenji Inoue, Yuki Okamoto |
ICRA | 2 |