Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Gou Koutaki

dblp:85/10109 · DBLP profile ↗
← Back
16ranked-venue papers
9as first author
4since 2021 · last 2025
0000-0002-3414-1085ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 8 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Geometric modeling and processing · 34% Audio and music processing · 26% Computational photography and imaging · 17%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Artificial intelligence
2 papers
3D vision · 50% Image recognition and object detection · 40% Representation and self-supervised learning · 10%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 50% Interaction techniques and input · 50%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing
music generation
0.912025
RoboSax Melody Slot Machine · ACM Multimedia 2025
Geometric modeling and processing
3d reconstruction
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Computational photography and imaging › image-based modeling
3d reconstruction from images
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Geometric modeling and processing
shape matching
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Mathematical optimization › combinatorial optimization
assignment problem
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Mathematical optimization
combinatorial optimization
0.612022
ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization · Int. J. Comput. Vis. 2022
Information retrieval › hashing
binary code learning
0.312018
Hadamard Coding for Supervised Discrete Hashing · IEEE Trans. Image Process. 2018
Information retrieval
hashing
0.312018
Hadamard Coding for Supervised Discrete Hashing · IEEE Trans. Image Process. 2018
Information retrieval
image retrieval
0.312018
Hadamard Coding for Supervised Discrete Hashing · IEEE Trans. Image Process. 2018
Information retrieval › hashing › supervised hashing
supervised discrete hashing
0.312018
Hadamard Coding for Supervised Discrete Hashing · IEEE Trans. Image Process. 2018
Rendering
display systems
0.212016
Binary continuous image decomposition for multi-view display · ACM Trans. Graph. 2016
Virtual and augmented reality › 3d display › stereoscopic display › autostereoscopic display
multi-view display
0.212016
Binary continuous image decomposition for multi-view display · ACM Trans. Graph. 2016
Computer vision › Image recognition and object detection
interest point detection
0.212015
Multiple-Hypothesis Affine Region Estimation with Anisotropic LoG Filters · ICCV 2015
Computer vision › 3D vision › low-level vision › feature detection
keypoint detection
0.212015
Multiple-Hypothesis Affine Region Estimation with Anisotropic LoG Filters · ICCV 2015
Image and video processing › multiscale analysis › multiresolution analysis
scale-space analysis
0.212014
Scale-Space Processing Using Polynomial Representations · CVPR 2014
Image and video processing
image decomposition
0.112016
Binary continuous image decomposition for multi-view display · ACM Trans. Graph. 2016
Computer vision › 3D vision › low-level vision
feature detection
0.112014
Scale-Space Processing Using Polynomial Representations · CVPR 2014
Machine learning › Representation and self-supervised learning › visual representation › image representation › handcrafted descriptor
SIFT
0.112014
Scale-Space Processing Using Polynomial Representations · CVPR 2014

Methods — techniques the papers use, named apart from their topics

MIDI · 1.7GTTM-based melody morphing · 1.7iterative closest point · 1.1assignment algorithm · 1.1spectral decomposition · 0.4principal component analysis · 0.4polynomial approximation · 0.4hadamard matrix · 0.3discrete cyclic coordinate descent · 0.3pulse-width modulation · 0.2optimization · 0.2binary continuous image decomposition · 0.2singular value decomposition · 0.2laplacian-of-gaussian filter · 0.2eigenfilter · 0.2
YearPublicationVenuePosition
2025 Automatic Fingering Saxophone Quartet System
Gou Koutaki, Masatoshi Hamanaka
ICEC1
2025 RoboSax Melody Slot Machine
abstract
RoboSax Melody Slot Machine is a mobile-to-acoustic performance system that connects an interactive tablet interface to live saxophone performance. Participants select melodic fragments on an iPad-based Melody Slot Machine, where related musical phrases are generated through GTTM-based melody morphing rather than arbitrary recombination. The selected melody is transmitted as MIDI to RoboSax, a robotic saxophone mechanism that actuates the instrument’s fingerings, while a human performer provides breath, tonguing, phrasing, dynamics, and tone color. This division of roles allows participants without instrumental training to influence the musical structure of a live performance while preserving the embodied expressiveness of acoustic wind performance. For SIGGRAPH Appy Hour, the work presents a participatory experience in which mobile interaction becomes immediately audible as live acoustic sound. By combining structurally coherent melody variation, audience-controlled interaction, and human–robot performance, RoboSax Melody Slot Machine explores how mobile music apps can extend beyond the screen and become part of a shared performative space.
Masatoshi Hamanaka, Gou Koutaki
ACM Multimedia2
2022 Hadamard-Coded Supervised Discrete Hashing on Complex and Quaternion Domain
abstract
This paper extends Hadamard-coded supervised discrete hashing on real domain (termed as ℝ-HCSDH) using a real-valued kernel trans-formation (ℝKT) to one on complex/quaternion domain (termed as ℂ-HCSDH/ℍ-HCSDH) using complex/quaternion-valued KTs (ℂKT/ℍKT). Supervised discrete hashing has recently been attracted for its efficiency in data retrieval. Efficient learning of a hashing function is at the core of SDH, and many methods have been proposed. Among them, HCSDH simplifies the learning process by introducing Hadamard codes and shows its efficiency. Although many studies on SDH, including HCSDH, focus on hashing function learning, KT, which is an initial step of SDH to generate a feature vector, also affects performance but has received less attention. This motivates us to establish more effective KTs in this work. Since conventional KTs are ℝKTs that only consider the distance be-tween the input data and each anchor chosen from a training dataset, it cannot distinguish two anchors being equidistant. To solve this problem, we introduce ℂKT/ℍKT to consider not only the distance but also the angle between the input data and each anchor. More-over, under the ℂKT and the ℍKT, we verify that Hadamard codes are still optimal for the HCSDH model. Experimental results show ℂ-HCSDH and ℍ-HCSDH outperform ℝ-HCSDH in (cross-modal) data retrieval.
Manabu Sueyasu, Seisuke Kyochi, Gou Koutaki
ICIP3
2022 ISHIGAKI Retrieval System Using 3D Shape Matching and Combinatorial Optimization
abstract
Abstract In April 2016, a massive earthquake with a magnitude of 7.3 struck Kumamoto region, Japan, causing major devastation. One of the structures that were damaged in Kumamoto was Kumamoto Castle, a cultural asset of great significance in Japan. The stone retaining wall “ishigaki” that formed the foundation of the castle collapsed, and the superstructure was destroyed. The number of stones is estimated to be more than 70,000, and restoration work is anticipated to take more than 20 years. Since each of the stones is an important cultural asset, the broken stone structure needed to be restored to its original state in order not to lose its cultural value forever. In addition, the fallen stones need to be returned to their original positions in the ishigakis. In similar cases, non-automatic visual verification was used. However, for Kumamoto Castle, this would have been impossible because a large number of stones were displaced as a result of the collapse. The purpose of this project is to provide support for the restoration work by matching the stones that fell down after the collapse with those before the collapse using information technology, such as computer vision and optimization technologies. Specifically, we captured photographic images of the stones before and after the collapse to match them. The technical contributions of this study are as follows: (a) To estimate the scale and surface orientation of the stones, we exploit 3D model construction from the images. (b) To solve the jigsaw-puzzle-like problem of reassembling the stone fragments, we exploit the combination of a customized iterative closest point (ICP) algorithm for shape position matching and an assignment algorithm to find the best pairs of stones before and after the collapse by using the matching degree obtained from ICP. Here, only the 2D shape of the stones before the collapse can be used due to the small number of photos available. In contrast, a detailed 3D shape can be obtained from the stones after their collapse. We matched these asymmetric data in 2D and 3D to enable a comprehensive reconstruction. (c) We developed a user-friendly graphical user interface system that was used by actual masons without special knowledge. The developed system was used to match the ishigaki of a turret, Iidamaru. As a result, we succeeded in identifying 337 stones, or approximately 90% of the 370 images. These results are expected to be useful for and were used as a blueprint during actual restoration work.
Gou Koutaki, Sakino Ando, Keiichiro Shirai, Tsuyoshi Kishigami
Int. J. Comput. Vis.1
2018 Hadamard Coded Discrete Cross Modal Hashing
abstract
Cross-modal retrieval is a hot topic in the fields of machine learning and media retrieval, making it possible to relate different types of media, such as image, text, and audio. A powerful method for the cross-modal retrieval, discrete cross-modal hashing (DCH), has recently been proposed. The DCH can encode different types of media feature vectors to binary codes. When stored in a database, the binary code makes searches efficient because the Hamming distance between the corresponding sections of two binary codes can be computed via a specialized CPU operation. Moreover, it has recently been shown that when optimizing hash functions for supervised discrete hashing (SDH), Hadamard matrices can be used, and this technique is named “HC-SDH”. In this study, we apply the HC-SDH to cross-modal hashing. Experimental results demonstrate that the proposed cross-modal hashing can achieve the same performance as the conventional DCH with 1/24 of the training time.
Koichi Eto, Gou Koutaki, Keiichiro Shirai
ICIP2
2018 Single image vehicle classification using pseudo long short-term memory classifier
Reza Fuad Rachmadi, Keiichi Uchimura, Gou Koutaki, Kohichi Ogata
J. Vis. Commun. Image Represent.3
2018 Hadamard Coding for Supervised Discrete Hashing
abstract
In this paper, we propose a learning-based supervised discrete hashing method. Binary hashing is widely used for large-scale image retrieval as well as video and document searches because the compact binary code representation is essential for data storage and reasonable for query searches using bit-operations. The recently proposed supervised discrete hashing (SDH) method efficiently solves mixed-integer programming problems by alternating optimization and the discrete cyclic coordinate descent (DCC) method. Based on some preliminary experiments, we show that the SDH method can be simplified without performance degradation. We analyze the simplified model and provide a mathematically exact solution thereof; we reveal that the exact binary code is provided by a "Hadamard matrix." Therefore, we named our method Hadamard codedsupervised discrete hashing (HC-SDH). In contrast to SDH, our model does not require an alternating optimization algorithm and does not depend on initial values. HC-SDH is also easier to implement than iterative quantization (ITQ). Experimental results involving a large-scale database show that Hadamard coding outperforms conventional SDH in terms of precision, recall, and computational time. On the large datasets SUN-397 and ImageNet, HC-SDH provides a superior mean average of precision (mAP) and top-accuracy compared to the conventional SDH methods with the same code length and FastHash. The training time of HC-SDH is 170 times faster than conventional SDH and the testing time including the encoding time is seven times faster than FastHash which encodes using a binary-tree.
Gou Koutaki, Keiichiro Shirai, Mitsuru Ambai
IEEE Trans. Image Process.1
2017 Texture Detection for Letter Carving Segmentation of Ancient Copper Inscriptions
abstract
As relics of history, ancient copper inscriptions are found in many countries. Information in the image or letter forms contained on copper ancient inscription has a very high value. The age and environmental factors caused damage to the surface of the inscription and also reduced the appearances of the image and letter. In this paper, we describe a novel segmentation methodology based on multi-texture features for ancient copper inscriptions which were severely damaged. The segmentation results of letters on ancient copper inscriptions by using the proposed method have an average accuracy of 90%. Based on these results, the proposed method is suitable for letter segmentation of the ancient copper inscriptions.
Susijanto T. Rasmana, Yoyon K. Suprapto, I Ketut Eddy Purnama, Keiichi Uchimura, Gou Koutaki
Int. J. Pattern Recognit. Artif. Intell.5
2016 Binary continuous image decomposition for multi-view display
abstract
This paper proposes multi-view display using a digital light processing (DLP) projector and new active shutter glasses. In conventional stereoscopic active shutter systems, active shutter glasses have a 0--1 (open and closed) state, and the right and left frames are temporally divided. However, this causes the display to flicker because the human eye perceives the appearance of black frames when the other shutter is closing. Furthermore, it is difficult to increase the number of views because the number of frames representing images is also divided. We solve these problems by extending the active shutter beyond the use of the 0--1 state to a continuous range of states [0, 1] instead. This relaxation leads to the formulation of a new DLP imaging model and an optimization problem. The special structure of DLP binary imaging and the continuous transmittance of the new active shutter glasses require the solution of a binary continuous image decomposition problem. Although it contains NP-hard problems, the proposed algorithm can efficiently solve the problem. The implementation of our imaging system requires the development of an active shutter device with continuous transmittance. We implemented the control of the transmittance of the liquid crystal display (LCD) shutter by using a pulse-width modulation (PWM). A simulation and the developed multi-view display system were used to show that our model can represent multi-view images more accurately than the conventional time-division 0-1 active shutter system.
Gou Koutaki
ACM Trans. Graph.1
2015 Multiple-Hypothesis Affine Region Estimation with Anisotropic LoG Filters
abstract
We propose a method for estimating multiple-hypothesis affine regions from a keypoint by using an anisotropic Laplacian-of-Gaussian (LoG) filter. Although conventional affine region detectors, such as Hessian/Harris-Affine, iterate to find an affine region that fits a given image patch, such iterative searching is adversely affected by an initial point. To avoid this problem, we allow multiple detections from a single keypoint. We demonstrate that the responses of all possible anisotropic LoG filters can be efficiently computed by factorizing them in a similar manner to spectral SIFT. A large number of LoG filters that are densely sampled in a parameter space are reconstructed by a weighted combination of a limited number of representative filters, called "eigenfilters", by using singular value decomposition. Also, the reconstructed filter responses of the sampled parameters can be interpolated to a continuous representation by using a series of proper functions. This results in efficient multiple extrema searching in a continuous space. Experiments revealed that our method has higher repeatability than the conventional methods.
Takahiro Hasegawa, Mitsuru Ambai, Kohta Ishikawa, Gou Koutaki, Yuji Yamauchi, Takayoshi Yamashita, Hironobu Fujiyoshi
ICCV4
2015 Optimization of color quantization with total luminance for DLP projector and its evaluation system
abstract
In order to design a color quantization method that addresses total luminance for digital light processing (DLP) projectors, we propose a framework for optimizing color quantization and light emitted diode (LED) luminance. We evaluate the proposed method and a system using a DLP projector and a CMOS camera. Experimental results indicate that our method improves the total luminance of a projected image approximately 120% when compared with results achieved using previous models and produces better image quality.
Gou Koutaki, Hiroshi Okajima, Nobutomo Matsunaga, Keiichi Uchimura
ICIP1
2015 Marker Identification Using ILEDs and RGB Color Descriptors
abstract
In optical motion capture systems, it is difficult to correctly recognize markers based on their unique identifiers (IDs) in a single frame. In this paper, we propose two types of light-emitting diodes (LEDs) and cameras, infrared (IR) and RGB, in order to correctly detect and identify all markers tracking objects in a given system. To detect and estimate the three-dimensional (3D) position of the marker, we measure IR LEDs using IR stereo cameras. Furthermore, in order to identify each marker, we calculate and compare the RGB color descriptor in the vicinity of its center. Our system consists of general IR and RGB cameras, and is easy to extend by increasing the number of cameras. We implemented an IR/RGB LED marker circuit and constructed a simple motion capture system to test the effectiveness of our system. The results show that our system can detect the 3D positions and unique IDs of markers in one frame.
Gou Koutaki, Shodai Hirata, Hiromu Sato, Keiichi Uchimura
ISMAR1
2014 Scale-Space Processing Using Polynomial Representations
abstract
In this study, we propose the application of principal components analysis (PCA) to scale-spaces. PCA is a standard method used in computer vision. The translation of an input image into scale-space is a continuous operation, which requires the extension of conventional finite matrix- based PCA to an infinite number of dimensions. In this study, we use spectral decomposition to resolve this infinite eigenproblem by integration and we propose an approximate solution based on polynomial equations. To clarify its eigensolutions, we apply spectral decomposition to the Gaussian scale-space and scale-normalized Laplacian of Gaussian (LoG) space. As an application of this proposed method, we introduce a method for generating Gaussian blur images and scale-normalized LoG images, where we demonstrate that the accuracy of these images can be very high when calculating an arbitrary scale using a simple linear combination. We also propose a new Scale Invariant Feature Transform (SIFT) detector as a more practical example.
Gou Koutaki, Keiichi Uchimura
CVPR1
2014 Scale-space filtering using a piecewise polynomial representation
abstract
Scale-space image processing is a basic technique used for object recognition and low-level feature extraction in computer vision. Many Gaussian filtering techniques have been proposed. Recently, the spectral decomposition method was proposed, which is an infinite version of principal components analysis. Using this method, Gaussian blurred images can be represented as polynomials with a scale parameter and a Gaussian blurred image with an arbitrary scale can be obtained from simple linear combinations of the convolved eigenimages. However, the scale is limited to a small range in this method. In this study, we propose an improvement to the spectral decomposition of a Gaussian kernel by widening the scale using a piecewise polynomial representation. We present an analysis of the continuous spectral decompositions of a Gaussian kernel and their eigensolutions. Experimental results show that the proposed method can generate accurate Gaussian blurred images with an arbitrary scale and a wide scale range.
Gou Koutaki, Keiichi Uchimura
ICIP1
2013 Scale-space compression and its application using spectral theory
abstract
In this paper, we propose the application of principal component analysis (PCA) to scale-spaces. PCA is a standard method used in computer vision tasks such as recognition of eigenfaces. Because the translation of an input image into scale-space is a continuous operation, it requires the extension of conventional finite matrix based PCA to an infinite number of dimensions. Here, we use spectral theory to resolve this infinite eigenproblem through the use of integration, and we propose an approximate solution based on polynomial equations. In order to clarify its eigensolutions, we apply spectral decomposition to gaussian scale-space. As an application of this proposed method we introduce a method for generating gaussian blur images, demonstrating that the accuracy of such an image can be made very high by using an arbitrary scale calculated through simple linear combination.
Gou Koutaki, Keiichi Uchimura
ICIP1
2012 Robust Face Recognition using Wavelet and DCT based Lighting Normalization, and Shifting-mean LDA
I Gede Pasek Suta Wijaya, Keiichi Uchimura, Gou Koutaki, Cuicui Zhang
ICPRAM (2)3