Hiroshi Watanabe 0001

dblp:95/1079-1 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-9306-688XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 2Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 InterpIoU: Robust bounding box regression loss within an interpolation-based IoU framework
abstract
Bounding box regression (BBR) is central to object detection, where regression loss plays a key role in precise localization. Existing IoU-based losses often rely on handcrafted geometric penalties to provide gradients in non-overlapping cases and improve localization. However, these geometric penalties are inherently sensitive to box geometry, producing unstable gradients in extreme cases and a subtle misalignment with the IoU objective, which harms small objects detection and yields undesired converge behaviors such as bounding box enlargement. To address these limitations, we introduce InterpIoU, an interpolation-based IoU optimization framework that rethinks BBR beyond handcrafted penalties. By bridging predictions and ground truth with interpolated boxes, InterpIoU supplies meaningful gradients in non-overlapping cases while ensuring consistent alignment with the BBR objective. Crucially, our findings challenge the convention of using geometric penalties, demonstrating they are often unnecessary and suboptimal. Building on InterpIoU, we propose Dynamic InterpIoU, which adjusts interpolation coefficients based on IoU values, adapting to diverse object distributions. Experiments on COCO, VisDrone, and PASCAL VOC demonstrate that our methods consistently outperform state-of-the-art IoU-based losses across detection frameworks, including YOLOv8 and DINO, with notable improvements for small object detection.
Hiroshi Watanabe 0001
Neurocomputing2
2026 Enhancing RGB-IR object detection: a frozen backbone approach with multi-receptive field attention
abstract
Recent advancements in multimodal object detection have predominantly relied on end-to-end training paradigms, which, while effective, demand substantial computational resources and risk feature degradation. To address these challenges, we propose a frozen backbone paradigm, preserving pretrained representations as stable semantic anchors for efficient multimodal fusion. Our approach introduces a lightweight multi-receptive field attention (MRFA) mechanism, enhancing feature interaction and representation diversity without exhaustive retraining. Experiments on the FLIR Aligned and M $$^3$$ FD dataset demonstrate consistent improvements over state-of-the-art end-to-end models, highlighting the potential of pretrained backbones coupled with adaptive attention mechanisms for robust multimodal object detection. The project code is released at https://github.com/LuBingyu11/MRFA .
Bingyu Lu, Hiroshi Watanabe 0001
Vis. Comput.3
2025 Time Step Generating: A Universal Synthesized Deepfake Image Detector
abstract
The rise of high-fidelity text-to-image diffusion models has made synthetic images increasingly indistinguishable from real ones, posing serious threats in digital security and media integrity. Existing detection methods often rely on reconstruction-based pipelines, which are computationally expensive and brittle on out-of-distribution data. We propose Time Step Generating (TSG), a universal synthetic image detector that leverages a pre-trained diffusion model as a feature extractor. By inputting images at a fixed diffusion timestep, TSG captures semantic and structural differences in noise prediction behavior between real and generated images — all within a single forward pass, enabling lightweight and effective classification. To eliminate the reliance on the manually chosen timestep hyperparameter, we further introduce TSG++, an enhanced version that consolidates multi-timestep diffusion features through lightweight fine-tuning. TSG++ learns to align features across all timesteps, producing a unified representation that improves both robustness and generalization without additional inference cost. Experiments on GenImage and challenging multimedia datasets demonstrate that TSG and TSG++ outperform prior methods in both accuracy and efficiency, offering a strong and adaptable solution for diffusion-based synthetic image detection.
Ziyue Zeng, Yupei Guo, Dingjie Peng, Hiroshi Watanabe 0001
MMAsia5
2025 Structure-Preserving Patch Decoding for Efficient Neural Video Representation
abstract
Implicit neural representations (INRs) are the subject of extensive research, particularly in their application to modeling complex signals by mapping spatial and temporal coordinates to corresponding values. When handling videos, mapping compact inputs to entire frames or spatially partitioned patch images is an effective approach. This strategy better preserves spatial relationships, reduces computational overhead, and improves reconstruction quality compared to coordinate-based mapping. However, predicting entire frames often limits the reconstruction of high-frequency visual details. Additionally, conventional patch-based approaches based on uniform spatial partitioning tend to introduce boundary discontinuities that degrade spatial coherence. We propose a neural video representation method based on Structure-Preserving Patches (SPPs) to address such limitations. Our method separates each video frame into patch images of spatially aligned frames through a deterministic pixel-based splitting similar to PixelUnshuffle. This operation preserves the global spatial structure while allowing patch-level decoding. We train the decoder to reconstruct these structured patches, enabling a global-to-local decoding strategy that captures the global layout first and refines local details. This effectively reduces boundary artifacts and mitigates distortions from naive upsampling. Experiments on standard video datasets demonstrate that our method achieves higher reconstruction quality and better compression performance than existing INR-based baselines.
Taiga Hayami, Kakeru Koizumi, Hiroshi Watanabe 0001
MMSP3
2025 Adapting Image-to-Video Diffusion Models for Large-Motion Frame Interpolation
abstract
With the development of video generation models has advanced significantly in recent years, we adopt large- scale image-to-video diffusion models for video frame interpolation. We present a conditional encoder designed to adapt an image-to-video model for large-motion frame interpolation. To enhance performance, we integrate a dual-branch feature extractor and propose a cross-frame attention mechanism that effectively captures both spatial and temporal information, enabling accurate interpolations of intermediate frames. Our approach demonstrates superior performance on the Fréchet Video Distance (FVD) metric when evaluated against other state-of-the- art approaches, particularly in handling large motion scenarios, highlighting advancements in generative-based methodologies.
Luoxu Jin, Hiroshi Watanabe 0001
MMSP2
2025 Guided Diffusion for the Extension of Machine Vision to Human Visual Perception
abstract
Image compression technology eliminates redundant information to enable efficient transmission and storage of images, serving both machine vision and human visual perception. For years, image coding focused on human perception has been well-studied, leading to the development of various image compression standards. On the other hand, with the rapid advancements in image recognition models, image compression for AI tasks, known as Image Coding for Machines (ICM), has gained significant importance. Therefore, scalable image coding techniques that address the needs of both machines and humans have become a key area of interest. Additionally, there is increasing demand for research applying the diffusion model, which can generate human-viewable images from a small amount of data to image compression methods for human vision. Image compression methods that use diffusion models can partially reconstruct the target image by guiding the generation process with a small amount of conditioning information. Inspired by the diffusion model’s potential, we propose a method for extending machine vision to human visual perception using guided diffusion. Utilizing the diffusion model guided by the output of the ICM method, we generate images for human perception from random noise. Guided diffusion acts as a bridge between machine vision and human vision, enabling transitions between them without any additional bitrate overhead. The generated images then evaluated based on bitrate and image quality, and we compare their compression performance with other scalable image coding methods for humans and machines.
Takahiro Shindo, Yui Tatsumi, Taiju Watanabe, Hiroshi Watanabe 0001
MMSP4
2025 Explicit Residual-Based Scalable Image Coding for Humans and Machines
abstract
Scalable image compression is a technique that progressively reconstructs multiple versions of an image for different requirements. In recent years, images have increasingly been consumed not only by humans but also by image recognition models. This shift has drawn growing attention to scalable image compression methods that serve both machine and human vision (ICMH). Many existing models employ neural network-based codecs, known as learned image compression, and have made significant strides in this field by carefully designing the loss functions. In some cases, however, models are overly reliant on their learning capacity, and their architectural design is not sufficiently considered. In this paper, we enhance the coding efficiency and interpretability of ICMH framework by integrating an explicit residual compression mechanism, which is commonly employed in resolution scalable coding methods such as JPEG2000. Specifically, we propose two complementary methods: Feature Residual-based Scalable Coding (FR-ICMH) and Pixel Residual-based Scalable Coding (PR-ICMH). These proposed methods are applicable to various machine vision tasks. Moreover, they provide flexibility to choose between encoder complexity and compression performance, making it adaptable to diverse application requirements. Experimental results demonstrate the effectiveness of our proposed methods, with PR-ICMH achieving up to 29.57% BD-rate savings over the previous work.
Yui Tatsumi, Ziyue Zeng, Hiroshi Watanabe 0001
MMSP3
2024 Image Coding for Machines with Object Region Learning
abstract
Compression technology is essential for efficient image transmission and storage. With the rapid advances in deep learning, images are beginning to be used for image recognition as well as for human vision. For this reason, research has been conducted on image coding for image recognition, and this field is called Image Coding for Machines (ICM). We propose an image compression model that learns object regions. This compression model can cleanly decode only the object regions in the image. In the experiments, we verify the effectiveness of our model by comparing it with previous methods.
Takahiro Shindo, Taiju Watanabe, Kein Yamada, Hiroshi Watanabe 0001
CCNC4
2024 Improving Image Coding for Machines Through Optimizing Encoder Via Auxiliary Loss
abstract
Image coding for machines (ICM) aims to compress images for machine analysis using recognition models rather than human vision. Hence, in ICM, it is important for the encoder to recognize and compress the information necessary for the machine recognition task. There are two main approaches in learned ICM; optimization of the compression model based on task loss, and Region of Interest (ROI) based bit allocation. These approaches provide the encoder with the recognition capability. However, optimization with task loss becomes difficult when the recognition model is deep, and ROI-based methods often involve extra overhead during evaluation. In this study, we propose a novel training method for learned ICM models that applies auxiliary loss to the encoder to improve its recognition capability and rate-distortion performance. Our method achieves Bjøntegaard Delta rate improvements of $27.7 \%$ and $20.3 \%$ in object detection and semantic segmentation tasks, compared to the conventional training method.
Kei Iino, Shunsuke Akamatsu, Hiroshi Watanabe 0001, Shohei Enomoto, Akira Sakamoto, Takeharu Eda
ICIP3
2024 GEEG-YOLOv8: Gaussian Enhanced Euclidean Norm Ghost Attention for Real-Time Polyp Detection
abstract
Research on computer-aided polyp detection in gastrointestinal endoscopy has spanned the past few decades. Despite notable progress, the challenge of achieving automatic accurate and real-time polyp detection remains unresolved. This is because of the large differences in polyp characteristics such as shape, texture, size, and color, and the artifacts that are similar to polyp during endoscopy procedure. In this paper, we propose a novel Gaussian Enhanced Euclidean norm Ghost attention (GEEG) module for reliable real-time polyp detection on endoscopic images and videos. The new attention mechanism strengthens the features generated by Ghost convolution’s cheap operations by increasing the ability to extract inter-channel and spatial information inside the convolution layer. This module is integrated into the backbone of YOLOv8, creating a new model named GEEG-YOLOv8, to overcome above obstacles in polyp detection. Experiment results on three public datasets show that our proposed method outperforms existing state-of-the-art methods in both accuracy and speed.
Phuong Thao Nguyen, Hiroshi Watanabe 0001
ICIP2
2024 Image Coding For Machines With Edge Information Learning Using Segment Anything
abstract
Image Coding for Machines (ICM) is an image compression technique for image recognition. This technique is essential due to the growing demand for image recognition AI. In this paper, we propose a method for ICM that focuses on encoding and decoding only the edge information of object parts in an image, which we call SA-ICM. This is an Learned Image Compression (LIC) model trained using edge information created by Segment Anything. Our method can be used for image recognition models with various tasks. SA-ICM is also robust to changes in input data, making it effective for a variety of use cases. Additionally, our method provides benefits from a privacy point of view, as it removes human facial information on the encoder’s side, thus protecting one’s privacy. Furthermore, this LIC model training method can be used to train Neural Representations for Videos (NeRV), which is a video compression model. By training NeRV using edge information created by Segment Anything, it is possible to create a NeRV that is effective for image recognition (SANeRV). Experimental results confirm the advantages of SAICM, presenting the best performance in image compression for image recognition. We also show that SA-NeRV is superior to ordinary NeRV in video compression for machines. Code is available at https://github.com/final-0/SA-ICM.
Takahiro Shindo, Kein Yamada, Taiju Watanabe, Hiroshi Watanabe 0001
ICIP4
2024 Scalable Image Coding for Humans and Machines Using Feature Fusion Network
abstract
As image recognition models become more prevalent, scalable coding methods for machines and humans gain more importance. Applications of image recognition models include traffic monitoring and farm management. In these use cases, the scalable coding method proves effective because the tasks require occasional image checking by humans. Existing image compression methods for humans and machines meet these requirements to some extent. However, these compression methods are effective solely for specific image recognition models. We propose a learning-based scalable image coding method for humans and machines that is compatible with numerous image recognition models. We combine an image compression model for machines with a compression model, providing additional information to facilitate image decoding for humans. The features in these compression models are fused using a feature fusion network to achieve efficient image compression. Our method's additional information compression model is adjusted to reduce the number of parameters by enabling combinations of features of different sizes in the feature fusion network. Our approach confirms that the feature fusion network efficiently combines image compression models while reducing the number of parameters. Furthermore, we demonstrate the effectiveness of the proposed scalable coding method by evaluating the image compression performance in terms of decoded image quality and bitrate. Code is available at https://github.com/final-0/ICM-v1.
Takahiro Shindo, Taiju Watanabe, Yui Tatsumi, Hiroshi Watanabe 0001
MMSP4
2022 Application of Multi-modal Fusion Attention Mechanism in Semantic Segmentation
Yunlong Liu 0008, Osamu Yoshie, Hiroshi Watanabe 0001
ACCV (7)3
2019 Data Driven Cyber-Physical System for Landslide Detection
Zhi Liu 0002, Toshitaka Tsuda, Hiroshi Watanabe 0001, Satoko Ryuo, Nagateru Iwasawa
Mob. Networks Appl.3
2017 Adaptive Video Streaming in Hybrid Landslide Detection System with D-S Theory
abstract
Disaster detection is an important research topic and draws great attentions from both industry and academia. In this paper, we study the hybrid landslide detection system, which utilizes the video surveillance camera and multiple kinds of sensors and can detect the landslide automatically. Edge processing is adopted in this hybrid system to fuse the sensor data based on the Dempster-Shafer (D-S) theory, i.e. utilizing multiple sensors' information to calculate the possibility of the landslide. Edge processing can make faster decisions than the control center, the results of the edge processing are then used to schedule the sensor's transmission frequency and video transmission under the network constraints in this system. The simulation results show that the proposed scheme outperforms the competing schemes in typical network scenarios.
Zhi Liu 0002, Kenji Kanai, Masaru Takeuchi, Toshitaka Tsuda, Hiroshi Watanabe 0001
SMARTCOMP5
2012 A modified parabolic prediction based fractional pixel motion estimation using preset corrector
abstract
In general, a motion estimator in today's video coding includes fractional pixel motion estimation (FME) as well as integer pixel motion estimation (IME). FME can get better quality performance at the cost of higher computational complexity than IME alone. In this paper, a modified parabolic prediction based FME for H.264 video coding is proposed. The proposed method uses specific correction coefficients to improve the PSNR performance of the existing prediction based algorithm. In the simulation results, compared with the original algorithm, our technique shows more accurate predictive power, meanwhile, it does not require the interpolation process.
Chang-Uk Jeong, Hiroshi Watanabe 0001
APCC2
2011 Automatic preview generation of comic episodes for digitized comic search
abstract
This research proposes a novel method to present "thumbnails" of episodes of digitized comics, in order to improve the efficiency of comic search. Comic episode thumbnails are generated based on image analysis technologies developed especially for comic images. Namely, the following procedures are developed for our system: automatic comic frame segmentation, text balloon extraction, and a linear regression based model to calculate the importance score of each extracted frame. The system then selects frames from each episode with high importance score, and aligns the selected frames to create the episode thumbnail, which is presented to the system user as a compact preview of the episode. User experiments conducted with actual Japanese comic images prove that the proposed method significantly decreases the time necessary to search for specific episodes from a large scaled comic data collection.
Keiichiro Hoashi, Chihiro Ono, Daisuke Ishii, Hiroshi Watanabe 0001
ACM Multimedia4
2008 Divisible Load Scheduling with Result Collection on Heterogeneous Systems
abstract
Divisible Load Theory (DLT) is an established mathematical framework to study Divisible Load Scheduling (DLS). However, traditional DLT does not comprehensively deal with the scheduling of results back to source (i.e., result collection) on heterogeneous systems. In this paper, the DLSRCHETS (DLS with Result Collection on HETerogeneous Systems) problem is addressed. The few papers to date that have dealt with DLSRCHETS, proposed simplistic LIFO (Last In, First Out) and FIFO (First In, First Out) type of schedules as solutions to DLSRCHETS. In this paper, a new heuristic algorithm, ITERLP, is proposed as a solution to the DLSRCHETS problem. With the help of simulations, it is proved that the performance of ITERLP is significantly better than existing algorithms.
Abhay Ghatpande, Hidenori Nakazato, Hiroshi Watanabe 0001, Olivier Beaumont
IPDPS3
2006 Bit Rate Reduction of Vector Representation of Binary Images
abstract
Vector representation of binary images has an advantage of keeping high image quality for arbitrary scaling as well as editing capability of an object. However, the vector representation suffers from low compression efficiency compared with JBIG. In this paper, we show the main cause reducing coding efficiency and propose two methods to improve it. The proposed methods can reduce the file size of a binary image up to about 30-40 percent.
Yuki Yamamoto, Kei Kawamura, Hiroshi Watanabe 0001
ICIP3
2006 A Study on Spatial Scalable Coding using Vector Representation
abstract
The major advantage of vector representation of an image is that the image quality is maintained for arbitrary scaling. In recent years, a demand for scalable image coding has been increasing because of the wide variety of available digital contents and display terminals. Conventional scalable coding schemes are based on raster representation, and thus, line drawings deteriorate in quality when expanded and shrunk. In this paper, we propose an edge reconstruction method using vector representation for the purpose of keeping a consistent spatial scalability on transmission and display. We take an anti-aliasing into account in edge areas for approximation of luminance values around the edge. The proposed method can improve PSNR by up to 2 dB as compared to the conventional methods when image is expanded and shrunk
Yuki Yamamoto, Kei Kawamura, Hiroshi Watanabe 0001
ICME3
2005 Gradation approximation for vector based compression of comic images
abstract
In this paper, we propose a method to vectorize comic images including halftone dots. Our method can prevent the jaggy and moire phenomena when the images are enlarged and shrank. The method contains three modules: halftone dots separation, gradation approximation, and vectorization. At the first module, small isolated areas and altered areas by dilation and erosion operations are separated into halftone dots images. At the second module, areas of halftone are approximated by contours and gradation parameters. At the last module, both halftone areas and line drawings are vectorized and approximated by smooth contours. The size of compressed file produced by our method is equal or smaller than JBIG compression. Validity of the proposed method is confirmed by experimental results.
Kei Kawamura, Yuki Yamamoto, Hiroshi Watanabe 0001
ICIP (3)3
2004 Vector representation of binary images containing halftone dots
abstract
Vector representation of graphics has the advantage that the image can be displayed at any size. When the resolution of a bitmap image is changed, the lack of a line segment arises. In addition, moire occurs when the resolution of an image with halftone dots is changed. We propose a new technique to convert a binary image with halftone dots into its vector representation. Resolution conversion of a binary image can easily be performed without moire by using a continuous tone approximation of the halftone dots. First, we separate the area of halftone dots and line drawings in the image. Next, a continuous tone approximation is applied to the area of halftone dots. Then, conventional vectorization is applied to both continuous tone areas and line drawings. Finally, these components are mixed and reconstructed. Our approach provides an efficient way of displaying cartoon-like images at any size with a limited amount of data.
Kei Kawamura, Hiroshi Watanabe 0001, Hideyoshi Tominaga
ICME2
2002 A novel decoder-downloadable system for content-oriented coding
abstract
A new system architecture called decoder-downloadable system is described. The purpose of this system is to provide a uniform platform for the multimedia world based on image compressions using characteristics of the contents. This system enables dynamic decoder downloading by negotiation between servers and clients for seamless and minimum delay playback. This negotiation scheme is implemented by Java and CORBA (Common Object Request Broker). Thus, servers and clients in this system are independent from OSes and hardware specifications. This system is suitable for TV broadcasting as well as Internet streaming. We improve a decoder architecture in order to make the system more flexible. A decoder has some functions such as bitstream parsing, image compression algorithms (DCT, wavelet, motion estimation and so on). However, conventional decoders are monolithic software. Thus, it is impossible to share parts of decoders. The proposed method enables a module-based decoder and sharing some modules among some decoders. Consequently, the flexible decoder can make downloading time shorter by avoidance to download redundant parts of decoders. It also achieves scalability enabling this system to be used in some multimedia applications, such as a multimedia content search system, as well as a multimedia player.
Naoto Shimizu, Toshinori Miyazawa, Wataru Kameyama, Hiroshi Watanabe 0001, Hideyoshi Tominaga
GLOBECOM4
2002 A study on two-layer coding for animation images
abstract
A coding scheme specifically designed for animation images is proposed. Taking the characteristics of animation images into account, lines and homogeneous color regions are extracted from animation images. Lines and contours of the homogeneous color regions are approximated by straightline and spline functions. We found that smoothing operations are effective to extract homogeneous regions from background images. A connected filter is used for the smoothing operation while keeping edges of the original image correctly. Precise approximation of the contours is also required. For this purpose, we propose to use dynamic programming with a feedback algorithm. As a result, an animation image can be represented by two layers; significant points as a base layer, and discrete cosine transform (DCT) blocks for an additional layer. For the compensation of the loss caused by the smoothing operation, the DCT is applied to the difference between the original images and lines and homogeneous color regions images. High-quality images can be obtained by adding this differential data to the approximated images.
Ouji Nakagami, Toshinori Miyazawa, Hiroshi Watanabe 0001, Hideyoshi Tominaga
ICME (1)3
1992 Bit allocation and rate control based on human visual sensitivity for interframe coders
abstract
Video coding standards such as the CCITT H.261 and the ISO 11172 allow substantial freedom in the exact methodology used to allocate bits to different parameters and control the output bit rate. Thus, different standard conforming coders can give different picture quality depending on the bit allocation and rate control strategy used at the encoder. The authors propose a new bit allocation and rate control strategy that is based on some assumptions about the sensitivity of the human visual system to distortion. They use the ISO 11172 standard in their examples, although the strategy is equally applicable to other DCT based interframe coding schemes. The new scheme shows considerable quality improvement over conventional schemes in the informal subjective tests.>
Hiroshi Watanabe 0001, Sharad Singhal
ICASSP1
1989 64 kbit/s Video coding algorithm using adaptive gain/shape vector quantization
Hiroshi Watanabe 0001, Yutaka Suzuki
Signal Process. Image Commun.1