EDBT 2026 Demo / reviewers in the wild / expert
Akira Nakagawa
dblp:87/10305
· DBLP profile ↗
14ranked-venue papers
3as first author
4since 2021 · last 2024
0009-0008-1563-1573ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 36% Representation and self-supervised learning · 36% Probabilistic and Bayesian machine learning · 19% | |
| Computer graphics and multimedia
1 paper |
Image and video coding · 77% Multimedia systems and quality of experience · 23% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
variational autoencoder |
0.9 | 2 | 2021 | Quantitative Understanding of VAE as a Non-linearly Scaled Isometric Embedding · ICML 2021 Rate-distortion optimization guided autoencoder for isometric embedding in Euclidean latent space · ICML 2020 |
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
isometric embedding |
0.5 | 1 | 2021 | Quantitative Understanding of VAE as a Non-linearly Scaled Isometric Embedding · ICML 2021 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.5 | 1 | 2021 | Quantitative Understanding of VAE as a Non-linearly Scaled Isometric Embedding · ICML 2021 |
Machine learning › Representation and self-supervised learning › representation learning › representation geometry
latent space geometry |
0.4 | 1 | 2020 | Rate-distortion optimization guided autoencoder for isometric embedding in Euclidean latent space · ICML 2020 |
Image and video coding › video compression
video codec |
0.2 | 1 | 2013 | High-Performance Video Codec for Super Hi-Vision · Proc. IEEE 2013 |
Machine learning › Time series and sequential data
anomaly detection |
0.1 | 1 | 2020 | Rate-distortion optimization guided autoencoder for isometric embedding in Euclidean latent space · ICML 2020 |
Machine learning › Time series and sequential data › anomaly detection
unsupervised anomaly detection |
0.1 | 1 | 2020 | Rate-distortion optimization guided autoencoder for isometric embedding in Euclidean latent space · ICML 2020 |
Multimedia systems and quality of experience
video transmission |
0.0 | 1 | 2013 | High-Performance Video Codec for Super Hi-Vision · Proc. IEEE 2013 |
Methods — techniques the papers use, named apart from their topics
rate-distortion theory · 0.5differential geometry · 0.5rate-distortion optimization · 0.4orthonormal transform coding · 0.4temporal division · 0.2spatial division · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Towards Generated Image Provenance Analysis via Conceptual-Similar-Guided-SLIP RetrievalabstractWith the prevalence of state-of-the-art generative models, photorealistic synthetic images can now be easily generated. However, the generated images may replicate contents from the original training images, which can lead to potential legal issues. In this paper, we propose a novel method calledConceptual-Similar-guided Self-supervised Language-Image Pre-training(CS-SLIP) that leverages both image and text modalities for the generated image provenance. Besides the self-supervised learning branch and contrastive learning branch, a conceptual-similar branch is designed to guide the model to learn a better feature representation of image-text-pairs. We also adopt the re-ranking method to refine the initial matching candidates via the cross-modal bi-directional retrieval. Extensive qualitative and quantitative experiments are conducted, which demonstrate that the replication indeed exists in the generated images, and our proposed method can effectively retrieve the most similar images from the training corpus to achieve the goal of generated image provenance analysis. Xiaojie Xia, Liuan Wang, Jun Sun 0004, Akira Nakagawa |
IEEE Signal Process. Lett. | 4 |
| 2023 | An Auto-Encoder to Reconstruct Structure with Cryo-EM Images via Theoretically Guaranteed Isometric Latent Space, and Its Application for Automatically Computing the Conformational Pathway
Kimihiro Yamazaki, Yuichiro Wada, Atsushi Tokuhisa, Mutsuyo Wada, Takashi Katoh, Yuhei Umeda, Yasushi Okuno, Akira Nakagawa |
MICCAI (1) | 8 |
| 2023 | Improving Predicate Representation in Scene Graph Generation by Self-Supervised LearningabstractScene graph generation (SGG) aims to understand sophisticated visual information by detecting triplets of subject, object, and their relationship (predicate). Since the predicate labels are heavily imbalanced, existing supervised methods struggle to improve accuracy for the rare predicates due to insufficient labeled data. In this paper, we propose SePiR, a novel self-supervised learning method for SGG to improve the representation of rare predicates. We first train a relational encoder by contrastive learning without using predicate labels, and then fine-tune a predicate classifier with labeled data. To apply contrastive learning to SGG, we newly propose data augmentation in which subject-object pairs are augmented by replacing their visual features with those from other images having the same object labels. By such augmentation, we can increase the variation of the visual features while keeping the relationship between the objects. Comprehensive experimental results on the Visual Genome dataset show that the SGG performance of SePiR is comparable to the state-of-theart, and especially with the limited labeled dataset, our method significantly outperforms the existing supervised methods. Moreover, SePiR’s improved representation enables the model architecture simpler, resulting in 3.6x and 6.3x reduction of the parameters and inference time from the existing method, independently. So Hasegawa, Masayuki Hiromoto, Akira Nakagawa, Yuhei Umeda |
WACV | 3 |
| 2021 | Quantitative Understanding of VAE as a Non-linearly Scaled Isometric EmbeddingabstractVariational autoencoder (VAE) estimates the posterior parameters (mean and variance) of latent variables corresponding to each input data. While it is used for many tasks, the transparency of the model is still an underlying issue. This paper provides a quantitative understanding of VAE property through the differential geometric and information-theoretic interpretations of VAE. According to the Rate-distortion theory, the optimal transform coding is achieved by using an orthonormal transform with PCA basis where the transform space is isometric to the input. Considering the analogy of transform coding to VAE, we clarify theoretically and experimentally that VAE can be mapped to an implicit isometric embedding with a scale factor derived from the posterior parameter. As a result, we can estimate the data probabilities in the input space from the prior, loss metrics, and corresponding posterior parameters, and further, the quantitative importance of each latent variable can be evaluated like the eigenvalue of PCA. Akira Nakagawa, Keizo Kato, Taiji Suzuki |
ICML | 1 |
| 2020 | Rate-distortion optimization guided autoencoder for isometric embedding in Euclidean latent spaceabstractTo analyze high-dimensional and complex data in the real world, deep generative models, such as variational autoencoder (VAE) embed data in a low-dimensional space (latent space) and learn a probabilistic model in the latent space. However, they struggle to accurately reproduce the probability distribution function (PDF) in the input space from that in the latent space. If the embedding were isometric, this issue can be solved, because the relation of PDFs can become tractable. To achieve isometric property, we propose Rate-Distortion Optimization guided autoencoder inspired by orthonormal transform coding. We show our method has the following properties: (i) the Jacobian matrix between the input space and a Euclidean latent space forms a constantly-scaled orthonormal system and enables isometric data embedding; (ii) the relation of PDFs in both spaces can become tractable one such as proportional relation. Furthermore, our method outperforms state-of-the-art methods in unsupervised anomaly detection with four public datasets. Keizo Kato, Tomotake Sasaki, Akira Nakagawa |
ICML | 4 |
| 2018 | Speed-Up of Object Detection Neural Network with GPUabstractWe realized a speed-up of an object detection neural network with GPU. We improved the object detection speed of faster R-CNN [1], which is one of the most commonly used detection networks [2]. The speed of the original faster R-CNN (py - faster - rcnn [3]) was 72.4ms per image on our GPU server11OS: Ubuntu 14.04.5 LTS (GNU/Linux 4.2.0-42-generic x86_64), CPU: Intel(R) Xeon(R) CPU E5-2680 v4 @ 2.40GHz, GPU: GPU (Tesla P100-PCIE-16GB) x 1, Libraries: MKL, CUDA 8.0, cuDNN v5.1.5. We accelerated the detection speed by implementing our new algorithms that are suitable for GPUs. The speed-up is realized without sacrificing the object detection accuracy (mAP). Our GPU-accelerated faster R-CNN can detect objects with 55.8ms per image. This is nearly 30% speed-up. In detection networks, the processes of building scored candidate regions, sorting and non-maximum-suppression (nms) are commonly used. In faster R-CNN, these processes are executed in proposal layer. We reduced the processing time of the proposal layer from 5.6ms to 2.2ms. This is 2.5 times as fast as the original one. We also evaluated the detection speed with larger batch sizes. By applying batch size 16, it is accelerated to 44.9ms per image. This is 1.6 times as fast as the original faster R-CNN (py-faster-rcnn). Since we realized a speed-up of common basic methods for detection networks, our speed-up methods are also applicable to other detection networks such as R-FCN [4], YOLO [5] [6] and SSD [7]. Takuya Fukagai, Kyosuke Maeda, Satoshi Tanabe, Koichi Shirahata, Yasumoto Tomita, Atsushi Ike, Akira Nakagawa |
ICIP | 7 |
| 2013 | High-Performance Video Codec for Super Hi-VisionabstractTo help pave the way for Super Hi-Vision (SHV) broadcasting, we have developed a new codec system that can encode and decode SHV signals in real time. This is the third generation of SHV real-time hardware codec. This efficient compression system maintains high picture quality by using eight 1080/60p (60 frames/s) encoding units and a video format converter with signal compensation processing that takes the properties of the Dual Green format of SHV into account. The video format converter divides an SHV image spatially into eight 1920 × 1080 portions, each of which is fed to the encoding unit. In the previous SHV codec, the SHV image was divided into 16 portions (spatially eight and temporally two) and 16 1080/30p encoding units were used. Compared with the previous system, the new codec achieves a 50% bitrate saving and downsizes the codec by almost half. Furthermore, several new technologies were developed and installed in the codec. We conducted the world's first SHV international transmission over an advanced Internet connection using the codec at a TS rate of 260 Mb/s. The received picture quality was good enough to show any kind of SHV content on a large screen. Yoshiaki Shishikui, Kazuhisa Iguchi, Shinichi Sakaida, Kimihiko Kazui, Akira Nakagawa |
Proc. IEEE | 5 |
| 2012 | Coefficient sign bit compression in video codingabstractWe propose a novel compression technique on signs of DCT coefficients. The proposed technique utilizes a high correlation among pixels at boundaries of a current block and those of neighboring blocks and estimates the original signs. It encodes the difference between the original coefficient signs and the estimated signs. We also propose a speeding up technique for the estimation, utilizing concepts of Gray code and transform matrix decomposition. Compared to JM16.2, the proposed technique reduces 10% bitrate in sign bits of DCT coefficients and achieves 1% BD-bitrate reduction in total bitstream. With the speeding up technique, complexities increase only 8% and 5% on an encoder and a decoder side, respectively. Jumpei Koyama, Akihiro Yamori, Kimihiko Kazui, Satoshi Shimada, Akira Nakagawa |
PCS | 5 |
| 2012 | Enhancement on the motion vector derivation of HEVC for interlace formatabstractWe propose a derivation scheme of a motion vector for interlace format. The adjustment of the vertical component of a chroma MV for 4:2:0 format and the adjustment of the vertical component of a luma prediction MV are introduced in order to remove inconsistency between the current derivation process of motion vector predictor and the nature of interlace format in HM4.0. The proposed scheme achieves up to 1.0dB PSNR improvements in chroma components and up to 3.4% bit-rate reduction. Satoshi Shimada, Jumpei Koyama, Akihiro Yamori, Kimihiko Kazui, Hidenobu Miyoshi, Akira Nakagawa |
PCS | 6 |
| 2002 | A study of an adaptive algorithm for stereo signals with a power differenceabstractWith broadband networks spreading, high-presence acoustic communication like stereophonic full-duplex communication is desired. When loudspeakers and microphones are used for handsfree communication, a stereophonic acoustic echo canceller (SAEC) is required to avoid howling and echoes. However, when there is a large power difference between stereo signals received from the far-end, the performance of the acoustic echo canceller is degraded; that is the convergence speed of the coefficient error is low. In this paper, to overcome this problem, we propose a new adaptive algorithm for the SAEC that uses the adjustment vector. Each channel received a signal normalized by the root of their squared sum. Computer simulation shows that the proposed algorithm improves the convergence speed for stereo signals with a power difference. Akira Nakagawa, Youichi Haneda |
ICASSP | 1 |
| 2000 | Channel-number-compressed multi-channel acoustic echo canceller for high-presence teleconferencing system with large displayabstractSound localization is important to make conversation easy between local and remote sites in a teleconference. This requires a multi-channel sound system having a multi-channel acoustic echo canceller (MAEC). The appropriate number of channels is determined from a trade-off between high presence and MAEC performance, so it is not possible to increase the channel number by much. We propose a channel-number-compressed MAEC to provide teleconferencing systems that exhibit high presence. The channel number of the MAEC inputs is compressed and that of its outputs is expanded. Akira Nakagawa, Suehiro Shimauchi, Youichi Haneda, Shigeaki Aoki, Shoji Makino |
ICASSP | 1 |
| 1999 | A stereo echo canceller implemented using a stereo shaker and a duo-filter control systemabstractStereo echo cancellation has been achieved and used in daily teleconferencing. To overcome the non-uniqueness problem, a stereo shaker is introduced in eight frequency bands and adjusted so as to be inaudible and not affect stereo perception. A due-filter control system including a continually running adaptive filter and a fixed filter is used for double-talk control. A second-order stereo projection algorithm is used in the adaptive filter. A stereo voice switch is also included. This stereo echo canceller was tested in two-way conversation in a conference room, and the strength of the stereo shaker was subjectively adjusted. A misalignment of 20 dB was obtained in the teleconferencing environment, and changing the talker's position in the transmission room did not affect the cancellation. This echo canceller is now used daily in a high-presence teleconferencing system and has been demonstrated to more than 300 attendees. Suehiro Shimauchi, Shoji Makino, Youichi Haneda, Akira Nakagawa, Sumitaka Sakauchi |
ICASSP | 4 |
| 1997 | Subband stereo echo canceller using the projection algorithm with fast convergence to the true echo pathabstractThis paper proposes a new subband stereo echo canceller that converges to the true echo path impulse response much faster than conventional stereo echo cancellers. Since signals are bandlimited and downsampled in the subband structure, the time interval between the subband signals become longer, so the variation of the crosscorrelation between the stereo input signals becomes large. Consequently, convergence to the true solution is improved. Furthermore, the projection algorithm, or affine projection algorithm, is applied to further speed up the convergence. Computer simulations using stereo signals recorded in a conference room demonstrate that this method significantly improves convergence speed and almost solves the problem of stereo echo cancellation with low computational load. Shoji Makino, Klaus Strauss, Suehiro Shimauchi, Youichi Haneda, Akira Nakagawa |
ICASSP | 5 |
| 1996 | SSB subband echo canceller using low-order projection algorithmabstractThis paper proposes a new subband echo canceller that almost fully whitens the received input by using the low-order projection algorithm, or the affine projection algorithm. Since the projection algorithm can fully whiten the received input of a small-tap adaptive filter with a relatively small projection order, the proposed subband projection echo canceller achieves nearly maximum convergence with a small projection order. By reflecting the frequency characteristics of the speech and of the echo path, it allows a different projection order to be chosen in different subbands. This gives the proposed method optimum cost performance, which is promising for implementation. Shoji Makino, Josef Nöbauer, Youichi Haneda, Akira Nakagawa |
ICASSP | 4 |