EDBT 2026 Demo / reviewers in the wild / expert
Hidehisa Nagano
dblp:04/6077
· DBLP profile ↗
15ranked-venue papers
6as first author
0since 2021 · last 2017
0000-0003-0617-6739ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-authorArtificial intelligence and machine learning · 3Systems, architecture and hardware · 3 · 2 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
4 papers |
Audio and music processing · 42% Multimedia analysis and retrieval · 32% Virtual and augmented reality · 24% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Human-computer interaction and pervasive computing
1 paper |
Accessibility and assistive technology · 100% | |
| Theoretical computer science
1 paper |
Coding theory · 100% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Virtual and augmented reality
user experience |
0.3 | 1 | 2017 | Visualizing Video Sounds With Sound Word Animation to Enrich User Experience · IEEE Trans. Multim. 2017 |
Accessibility and assistive technology
sound visualization |
0.3 | 1 | 2017 | Visualizing Video Sounds With Sound Word Animation to Enrich User Experience · IEEE Trans. Multim. 2017 |
Audio and music processing
music transcription |
0.2 | 1 | 2016 | Non-Negative Group Sparsity with Subspace Note Modelling for Polyphonic Transcription · IEEE ACM Trans. Audio Speech Lang. Process. 2016 |
Information retrieval › retrieval models › probabilistic retrieval model
BM25 |
0.2 | 1 | 2014 | BM25 With Exponential IDF for Instance Search · IEEE Trans. Multim. 2014 |
Information retrieval
retrieval models |
0.2 | 1 | 2014 | BM25 With Exponential IDF for Instance Search · IEEE Trans. Multim. 2014 |
Multimedia analysis and retrieval › visual search
instance search |
0.2 | 1 | 2014 | BM25 With Exponential IDF for Instance Search · IEEE Trans. Multim. 2014 |
Multimedia analysis and retrieval
video retrieval |
0.2 | 1 | 2014 | BM25 With Exponential IDF for Instance Search · IEEE Trans. Multim. 2014 |
Image and video coding › image compression
fractal image coding |
0.0 | 1 | 1999 | A Method for Implementing Fractal Image Compression on Reconfigurable Architecture · FPGA 1999 |
Hardware reliability and fault tolerance › error correction
error correction decoder |
0.0 | 1 | 1998 | Soft Decision Maximum Likelihood Decoders for Binary Linear Block Codes Implemented on FPGAs (Abstract) · FPGA 1998 |
Coding theory › error-correcting codes
decoding |
0.0 | 1 | 1998 | Soft Decision Maximum Likelihood Decoders for Binary Linear Block Codes Implemented on FPGAs (Abstract) · FPGA 1998 |
Coding theory › error-correcting codes › decoding › decoding algorithms › optimal decoding
maximum-likelihood decoding |
0.0 | 1 | 1998 | Soft Decision Maximum Likelihood Decoders for Binary Linear Block Codes Implemented on FPGAs (Abstract) · FPGA 1998 |
Coding theory › error-correcting codes › decoding
soft-decision decoding |
0.0 | 1 | 1998 | Soft Decision Maximum Likelihood Decoders for Binary Linear Block Codes Implemented on FPGAs (Abstract) · FPGA 1998 |
Reconfigurable computing and FPGAs
FPGA implementation |
0.0 | 2 | 1999 | A Method for Implementing Fractal Image Compression on Reconfigurable Architecture · FPGA 1999 Soft Decision Maximum Likelihood Decoders for Binary Linear Block Codes Implemented on FPGAs (Abstract) · FPGA 1998 |
Methods — techniques the papers use, named apart from their topics
text captioning · 0.6sound word animation · 0.6exponential IDF · 0.4bag of keypoints · 0.4BM25 · 0.4subspace modeling · 0.2nonnegative matrix factorization · 0.2group sparsity · 0.2gradient-based decomposition · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2017 | Visualizing Video Sounds With Sound Word Animation to Enrich User ExperienceabstractSound information in videos plays an important role in shaping the user experience. When sound is not accessible in videos, text captions can provide sound information. However, conventional text captions are not very expressive for nonverbal sounds because they are designed to visualize speech sounds. Here, we present a framework to automatically transform nonverbal video sounds into animated sound words and position them near the sound source objects in the video for visualization. This provides natural visual representation of nonverbal sounds with rich information about the sound category and dynamics. To evaluate how the animated sound words generated by our framework affect the user experience, we implemented an experimental system and conducted a user study involving over 300 participants from an online crowdsourcing service. The results of the user study show that the animated sound words can effectively and naturally visualize the dynamics of sound while clarifying the position of the sound source as well as contribute to making video-watching more enjoyable and increasing the visual impact of videos. Hidehisa Nagano, Kunio Kashino, Takeo Igarashi |
IEEE Trans. Multim. | 2 |
| 2016 | Non-Negative Group Sparsity with Subspace Note Modelling for Polyphonic TranscriptionabstractAutomatic music transcription (AMT) can be performed by deriving a pitch-time representation through decomposition of a spectrogram with a dictionary of pitch-labelled atoms. Typically, non-negative matrix factorisation (NMF) methods are used to decompose magnitude spectrograms. One atom is often used to represent each note. However, the spectrum of a note may change over time. Previous research considered this variability using different atoms to model specific parts of a note, or large dictionaries comprised of datapoints from the spectrograms of full notes. In this paper, the use of subspace modelling of note spectra is explored, with group sparsity employed as a means of coupling activations of related atoms into a pitched subspace. Stepwise and gradient-based methods for non-negative group sparse decompositions are proposed. Finally, a group sparse NMF approach is used to tune a generic harmonic subspace dictionary, leading to improved NMF-based AMT results. Ken O'Hanlon, Hidehisa Nagano, Nicolas Keriven, Mark D. Plumbley |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | A fast audio search method based on skipping irrelevant signals by similarity upper-bound calculationabstractIn this paper, we describe an approach to accelerate fingerprint techniques by skipping the search for irrelevant sections of the signal and demonstrate its application to the divide and locate (DAL) audio fingerprint method. The search result for the applied method, DAL3, is the same as that of DAL mathematically. Experimental results show that DAL3 can reduce the computational cost of DAL to approximately 25% for the task of music signal retrieval. Hidehisa Nagano, Ryo Mukai, Takayuki Kurozumi, Kunio Kashino |
ICASSP | 1 |
| 2015 | Visualizing video sounds with sound word animationabstractText captions are important means to provide sound information in videos when the sound is not accessible. However, conventional text captions are far less expressive for non-verbal sounds since they are designed to visualize speech sound. To address this problem, we propose a method for automatically transforming non-verbal video sounds to animated sound words, and positioning them near the sound source objects in the video for visualization. This provides natural visual representation of non-verbal sounds with rich information about the sound category and dynamics. We conducted a user study with over 300 participants using an online crowdsourcing service. The results showed that animated sound words could not only effectively and naturally visualize the dynamics of sound while clarify the position of the sound source, but also contribute to making video watching more enjoyable and increasing the visual impact of the video. Hidehisa Nagano, Kunio Kashino, Takeo Igarashi |
ICME | 2 |
| 2014 | Video Content Detection with Single Frame Level Accuracy Using Dynamic Thresholding TechniqueabstractThis paper proposes a video retrieval method that detects frame sections that correspond to shots in a query (video segment) with single frame level accuracy. The method adopts the coarse-to-fine strategy to decrease the processing time and the memory consumption, dynamic threshold with initial ranges for small segments is proposed to detect the exact beginning and end of each corresponding frame section to each shot in a query. Experiments on real videos show that our method can achieve accurate video detection with exact frame position while reducing processing time and memory consumption. Minoru Mori, Takayuki Kurozumi, Hidehisa Nagano, Kunio Kashino |
ICPR | 3 |
| 2014 | BM25 With Exponential IDF for Instance SearchabstractThis paper deals with a novel concept of an exponential IDF in the BM25 formulation and compares the search accuracy with that of the BM25 with the original IDF in a content-based video retrieval (CBVR) task. Our video retrieval method is based on a bag of keypoints (local visual features) and the exponential IDF estimates the keypoint importance weights more accurately than the original IDF. The exponential IDF is capable of suppressing the keypoints from frequently occurring background objects in videos, and we found that this effect is essential for achieving improved search accuracy in CBVR. Our proposed method is especially designed to tackle instance video search, one of the CBVR tasks, and we demonstrate its effectiveness in significantly enhancing the instance search accuracy using the TRECVID2012 video retrieval dataset. Masaya Murata, Hidehisa Nagano, Ryo Mukai, Kunio Kashino, Shin'ichi Satoh 0001 |
IEEE Trans. Multim. | 2 |
| 2012 | Structured sparsity for automatic music transcriptionabstractSparse representations have previously been applied to the automatic music transcription (AMT) problem. Structured sparsity, such as group and molecular sparsity allows the introduction of prior knowledge to sparse representations. Molecular sparsity has previously been proposed for AMT, however the use of greedy group sparsity has not previously been proposed for this problem. We propose a greedy sparse pursuit based on nearest subspace classification for groups with coherent blocks, based in a non-negative framework, and apply this to AMT. Further to this, we propose an enhanced molecular variant of this group sparse algorithm and demonstrate the effectiveness of this approach. Ken O'Hanlon, Hidehisa Nagano, Mark D. Plumbley |
ICASSP | 2 |
| 2010 | Statistical modeling of F0 dynamics in singing voices based on Gaussian processes with multiple oscillation basesabstractWe present a novel statistical model for dynamics of various singing behaviors, such as vibrato and overshoot, in a fundamental frequency (F0) contour. These dynamics are the important cues for perceiving individuality of a singer, and can be a useful measure for various applications, such as singing skill evaluation and singing voice synthesis. While most previous studies have modeled the dynamics using a second-order linear system, the automatic and accurate estimation of model parameters has yet to be accomplished. In this paper, we first develop a complete stochastic representation of the second-order system with Gaussian processes from parametric discretization, and propose a complete, efficient scheme for parameter estimation using the Expectation-Maximization (EM) algorithm. Experimental results show that the proposed method can decompose an F0 contour into a musical component and a dynamics component. Finally, we discuss estimating singing styles from the model parameters for each singer. Yasunori Ohishi, Hirokazu Kameoka, Daichi Mochihashi, Hidehisa Nagano, Kunio Kashino |
INTERSPEECH | 4 |
| 2007 | Robust Search Methods for Music Signals Based on Simple RepresentationabstractSignal similarity search is an important technique for music information retrieval. A basic task is finding identical signal segments on unlabeled music-signal archives, given a short music signal fragment as a query. In such a task, the search must be fast and sufficiently robust against possible signal fluctuations due to noise and distortions. In this special session paper, we describe a search method designed to cope with additive interfering sounds by spectral partitioning. Then, we introduce another method designed to be robust under multiplicative noise or distortion based on binary area representation. Kunio Kashino, Akisato Kimura, Hidehisa Nagano, Takayuki Kurozumi |
ICASSP (4) | 3 |
| 2003 | A fast search algorithm for background music signals based on the search for numerous small signal componentsabstractThe paper proposes a method for detecting and locating a known music signal in a long audio stream. Unlike existing methods, ours assumes that the music is used as background music (BGM) and overlapped by another sound such as speech and that the interfering sound is typically louder than the target music. The proposed method is based on time-series active search, which is a quick signal search method reported earlier (Kashino, K. et al., Proc. ICASSP-99, vol.VI, 1999). To realize the BGM search, however, a novel extension is introduced. That is, the music signal is first decomposed into a number of small time-frequency regions, and the search is carried out for each of those components. The results of the search are then integrated based on a voting scheme to find the target music locations. Experiments show that an accurate search is possible when SNR is -5 dB and that the search completes in about 8 s for a 30 min stored signal. Hidehisa Nagano, Kunio Kashino, Hiroshi Murase |
ICASSP (5) | 1 |
| 2003 | A fast search algorithm for background music signals based on the search for numerous small signal componentsabstractThis paper proposes a method for detecting and locating a known music signal in a long audio stream. Unlike existing methods, ours assumes that the music is used as background music (BGM) and overlapped by another sound such as speech and that the interfering sound is typically louder than the target music. The proposed method is based on time-series active search, which is a quick signal search method reported earlier. To realize the BGM search, however, a novel extension is introduced. That is, the music signal is firstly decomposed into a number of small time-frequency regions, and the search is carried out for each of those components. The results of the search are then integrated based on a voting scheme to find the target music locations. Experiments show that accurate search is possible when SNR is -5 dB and that the search completes in about 8 s for a 30-m stored signal. Hidehisa Nagano, Kunio Kashino, Hiroshi Murase |
ICME | 1 |
| 2002 | Fast music retrieval using polyphonic binary feature vectorsabstractWe propose a method for retrieving similar music from a polyphonic-music audio database using a polyphonic audio signal as a query. In this task, we must consider similarities among polyphonic signals of the music, and achieve quick retrieval. Therefore, we first introduce the polyphonic binary feature vector to represent the presence of multiple notes. This feature is suitable for the search based on the similarities among polyphonic audio signals. Then, we propose a new search method, which is quicker than the exhaustive use of DP matching. The search is accelerated using a "similarity matrix" to limit the search space. Experiments using a test database containing 216 music pieces show that the search accuracy of the proposed feature is 89%, which is approximately 26% higher than that of the conventional spectrum feature. It is also shown that the new search method retrieves similar music without significant accuracy degradation as well as the exhaustive search does and the computational complexity of the new search method is about 1/4 that of exhaustive search. Hidehisa Nagano, Kunio Kashino, Hiroshi Murase |
ICME (1) | 1 |
| 1999 | Acceleration of Linear Block Code Evaluations Using New Reconfigurable Computing ApproachabstractThis paper presents an approach to performing applications using reconfigurable computing (RC). Our RC approach is achieved by effective use of design automation systems. Logic circuits specialized for each individual application task are automatically implemented on FPGAs. Such circuits can quickly perform tasks that are time-consuming for general purpose computers. Decoding of binary linear block codes for the evaluation is taken up as an example application. Experimental results show that the time for decoding of the code specific decoding circuit implemented on FPGAs, in which computations are executed in parallel, is much shorter than that of the software decoder. Hidehisa Nagano, Takayuki Suyama, Akira Nagoya |
ASP-DAC | 1 |
| 1999 | A Method for Implementing Fractal Image Compression on Reconfigurable ArchitectureabstractNo abstract available. Akihiro Matsuura, Hidehisa Nagano, Akira Nagoya |
FPGA | 2 |
| 1998 | Soft Decision Maximum Likelihood Decoders for Binary Linear Block Codes Implemented on FPGAs (Abstract)
Hidehisa Nagano, Takayuki Suyama, Akira Nagoya |
FPGA | 1 |