Masato Fujitake

dblp:214/5696 · DBLP profile ↗
← Back
10ranked-venue papers
10as first author
9since 2021 · last 2024
0000-0001-7702-499XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 LayoutLLM: Large Language Model Instruction Tuning for Visually Rich Document Understanding
abstract
This paper proposes LayoutLLM, a more flexible document analysis method for understanding imaged documents. Visually Rich Document Understanding tasks, such as document image classification and information extraction, have gained significant attention due to their importance. Existing methods have been developed to enhance document comprehension by incorporating pre-training awareness of images, text, and layout structure. However, these methods require fine-tuning for each task and dataset, and the models are expensive to train and operate. To overcome this limitation, we propose a new LayoutLLM that integrates these with large-scale language models (LLMs). By leveraging the strengths of existing research in document image understanding and LLMs’ superior language understanding capabilities, the proposed model, fine-tuned with multimodal instruction datasets, performs an understanding of document images in a single model. Our experiments demonstrate improvement over the baseline model in various document analysis tasks.
Masato Fujitake
LREC/COLING1
2024 RL-LOGO: Deep Reinforcement Learning Localization for Logo Recognition
abstract
This paper proposes a novel logo image recognition approach incorporating a localization technique based on reinforcement learning. Logo recognition is an image classification task identifying a brand in an image. As the size and position of a logo vary widely from image to image, it is necessary to determine its position for accurate recognition. However, because there is no annotation for the position coordinates, it is impossible to train and infer the location of the logo in the image. Therefore, we propose a deep reinforcement learning localization method for logo recognition (RL-LOGO). It utilizes deep reinforcement learning to identify a logo region in images without annotations of the positions, thereby improving classification accuracy. We demonstrated a significant improvement in accuracy compared with existing methods in several published benchmarks. Specifically, we achieved an 18-point accuracy improvement over competitive methods on the complex dataset Logo-2K+. This demonstrates that the proposed method is a promising approach to logo recognition in real-world applications.
Masato Fujitake
ICASSP1
2024 JaPOC: Japanese Post-OCR Correction Benchmark Using Vouchers
Masato Fujitake
PRICAI (4)1
2024 DTrOCR: Decoder-only Transformer for Optical Character Recognition
abstract
Typical text recognition methods rely on an encoder-decoder structure, in which the encoder extracts features from an image, and the decoder produces recognized text from these features. In this study, we propose a simpler and more effective method for text recognition, known as the Decoder-only Transformer for Optical Character Recognition (DTrOCR). This method uses a decoder-only Transformer to take advantage of a generative language model that is pre-trained on a large corpus. We examined whether a generative language model that has been successful in natural language processing can also be effective for text recognition in computer vision. Our experiments demonstrated that DTrOCR outperforms current state-of-the-art methods by a large margin in the recognition of printed, handwritten, and scene text in both English and Chinese.
Masato Fujitake
WACV1
2023 A3S: Adversarial Learning of Semantic Representations for Scene-Text Spotting
abstract
Scene-text spotting is a task that predicts a text area on natural scene images and recognizes its text characters simultaneously. It has attracted much attention in recent years due to its wide applications. Existing research has mainly focused on improving text region detection, not text recognition. Thus, while detection accuracy is improved, the end-to-end accuracy is insufficient. Texts in natural scene images tend to not be a random string of characters but a meaningful string of characters, a word. Therefore, we propose adversarial learning of semantic representations for scene text spotting (A3S) to improve end-to-end accuracy, including text recognition. A3S simultaneously predicts semantic features in the detected text area instead of only performing text recognition based on existing visual features. Experimental results on publicly available datasets show that the proposed method achieves better accuracy than other methods.
Masato Fujitake
ICASSP1
2023 DiffusionSTR: Diffusion Model for Scene Text Recognition
abstract
This paper presents Diffusion Model for Scene Text Recognition (DiffusionSTR), an end-to-end text recognition framework using diffusion models for recognizing text in the wild. While existing studies have viewed the scene text recognition task as an image-to-text transformation, we rethought it as a text-text one under images in a diffusion model. We show for the first time that the diffusion model can be applied to text recognition. Furthermore, experimental results on publicly available datasets show that the proposed method achieves competitive accuracy compared to state-of-the-art methods.
Masato Fujitake
ICIP1
2022 Temporal feature enhancement network with external memory for live-stream video object detection
Masato Fujitake, Akihiro Sugimoto
Pattern Recognit.1
2021 Real-Time Object Detection by Feature Map Forecast for Live Streaming Video
abstract
This paper proposes a method that jointly learns to detect objects at the current frame and forecast the next frame’s future feature map. Previous offline detectors have shown the effectiveness of utilizing future information in video object detection; however, we cannot take such an approach when dealing with live streaming videos. In contrast, we utilize the forecast feature map with the current and past frame feature maps for object detection, where forecast feature maps are learned using observation of the present and past frames. To maintain a reliable forecast, we introduce a scheduler network, which decides whether we use the forecast feature map as input or extract the feature map from the next frame. Evaluations of our proposed model on the ImageNet VID dataset demonstrate the superior performance of our model against the public benchmark at similar architectures, with achieving 65.7% mAP at 38.9 fps.
Masato Fujitake, Akihiro Sugimoto
ICME1
2021 Temporally-aware Convolutional Block Attention Module for Video Text Detection
abstract
Scene text in video carries rich semantic information that is of great value in various content-based video applications. Existing methods have been proposed to improve accuracy, such as combining tracking; however, many have complex structures and are difficult to execute in real-time. Therefore, to run in real-time, this paper proposes a simple and practical feature refinement module, named Temporally-aware Convolutional Block Attention Module (TCBAM), based on a novel self-attention recurrent neural network. The model exploits still-image-based feature maps to refine temporal constant feature maps for better capturing widely varied appearances of video text. For better generalization, we also provide the flow-based data augmentation method with artificial data. Explements on the scene text video datasets including ICDAR2013 Video, Minetto, and RoadText-1K demonstrate that the proposed methods perform the competitive accuracy to the state-of-the-art models within real-time running. Our method with ResNext-50 can run at 17 FPS with 73.11 F-score on ICDAR 2013 Video without complex tracking methods.
Masato Fujitake, Hongpeng Ge
SMC1
2020 Temporal Feature Enhancement Network with External Memory for Object Detection in Surveillance Video
abstract
Video object detection is challenging and essential in practical applications, such as surveillance cameras for traffic control and public security. Unlike the video in natural scenes, the surveillance video tends to contain dense and small objects (typically vehicles) in their appearances. Therefore, existing methods for surveillance object detection utilize still-image object detection approaches with rich feature extractors at the expense of their run-time speeds. The run-time speed, however, becomes essential when the video is being streamed. In this paper, we exploit temporal information in videos to enrich the feature maps, proposing the first temporal attention based external memory network for the live stream of video. Extensive experiments on real-world traffic surveillance benchmarks demonstrate the real-time performance of the proposed model while keeping comparable accuracy with state-of-the-art.
Masato Fujitake, Akihiro Sugimoto
ICPR1