Soma Shiraishi

dblp:20/10109 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Image recognition and object detection · 100%
Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
multi-object detection
0.412019
Robust Multi-Object Detection Based on Data Augmentation with Realistic Image Synthesis for Point-of-Sale Automation · AAAI 2019
Computer vision › Image recognition and object detection
object detection
0.412019
Robust Multi-Object Detection Based on Data Augmentation with Realistic Image Synthesis for Point-of-Sale Automation · AAAI 2019
Visual content generation and editing
image generation
0.412019
Robust Multi-Object Detection Based on Data Augmentation with Realistic Image Synthesis for Point-of-Sale Automation · AAAI 2019

Methods — techniques the papers use, named apart from their topics

image synthesis · 0.8data augmentation · 0.8
YearPublicationVenuePosition
2025 Mask augmented Object-Centric Contrastive Learning for Amodal Instance Segmentation
abstract
Human cognition is robust in estimating depth ordering and occluded regions of objects, including amodal instance segmentation (AIS). Object-centric representation learning (OCRL) is an unsupervised approach to obtaining a new representation that mimics human common sense, such as amodal perception. Nevertheless, a significant gap exists between OCRL and human perception, and there is room for improvement for AIS. We aim to empower OCRL with amodal perception by solving self-supervised learning via contrastive learning for OCRL and depth-order estimation. The proposed method calculates the training loss on two masks composed of the object representations extracted from the original image and the transformed image by artificial occluders. Moreover, our method efficiently acquires depth-aware estimation by simultaneously solving the depth-ordering problem and representation learning. We have applied the proposed method to several simulation datasets and confirmed that the accuracy of AIS achieves SOTA performance under weakly supervised learning conditions.
Tomokazu Kaneko, Ryosuke Sakai, Takashi Shibata 0001, Soma Shiraishi
ICASSP4
2019 Robust Multi-Object Detection Based on Data Augmentation with Realistic Image Synthesis for Point-of-Sale Automation
abstract
As an alternative to bar-code scanning, we are developing a real-time retail product detector for point-of-sale automation. The major challenge associated with image based object detection arise from occlusion and the presence of other objects in close proximity. For robust product detection under such conditions, it is crucial to train the detector on a rich set of images with varying degrees of occlusion and proximity between the products, which fairly represents a wide range of customer tendencies of placing products together. However, generating a fairly large database of such images traditionally requires a large amount of human effort. On the other hand, acquiring individual object images with their corresponding masks is a relatively easy task. We propose an realistic image synthesis approach which uses individual object images and their corresponding masks to create training images with desired properties (occlusion and congestion among the products). We train our product detector over images thus generated and achieve a consistent performance improvement across different types of test data. With the proposed approach, detector achieves an improvement of 46.2% (from 0.67 to 0.98) and 40% (from 0.60 to 0.84) over precision and recall respectively, compared to using a basic training dataset containing one product per image.
Saiprasad Koturwar, Soma Shiraishi, Kota Iwamoto
AAAI2
2016 Analysis of satellite images for disaster detection
abstract
Analysis of satellite images plays an increasingly vital role in environment and climate monitoring, especially in detecting and managing natural disaster. In this paper, we proposed an automatic disaster detection system by implementing one of the advance deep learning techniques, convolutional neural network (CNN), to analysis satellite images. The neural network consists of 3 convolutional layers, followed by max-pooling layers after each convolutional layer, and 2 fully connected layers. We created our own disaster detection training data patches, which is currently focusing on 2 main disasters in Japan and Thailand: landslide and flood. Each disaster's training data set consists of 30000~40000 patches and all patches are trained automatically in CNN to extract region where disaster occurred instantaneously. The results reveal accuracy of 80%~90% for both disaster detection. The results presented here may facilitate improvements in detecting natural disaster efficiently by establishing automatic disaster detection system.
Siti Nor Khuzaimah Binti Amit, Soma Shiraishi, Tetsuo Inoshita, Yoshimitsu Aoki
IGARSS2
2012 A Part-Based Skew Estimation Method
abstract
In this paper we propose a part-based skew estimation method which is more robust to larger varieties of text images, such as camera-captured scene images. Specifically, the skew angle at each local part of the input image is estimated independently by referring the local part of upright character images stored as a database. Then the global skew angle is estimated by aggregating the estimated local skews. The proposed method does not assume that characters are laid-out in straight lines and thus have more robustness to the varieties of text images than conventional methods. The experimental results show the advantage of the proposed method over the conventional methods under several conditions.
Soma Shiraishi, Yaokai Feng, Seiichi Uchida
Document Analysis Systems1
2012 On the Possibility of Instance-Based Stroke Recovery
abstract
This paper tackles the stroke recovery problem, which is a typical ill-posed reverse problem, by an instance-based method. The basic idea of the instance-based stroke recovery is to refer to the drawing order of a similar instance. The instance-based method has a strong merit that it can deal with multi-stroke characters and other complex characters without any special consideration. However, it requires a sufficient numbers of instances to cover those various characters. As an initial trial of the instance-based stroke recovery method, this paper describes the principle of the method and then provides several experimental results. The experimental results indicate the potential of the proposed method on recovering the drawing order of complex characters, as expected.
Yutaro Iwakiri, Soma Shiraishi, Yaokai Feng, Seiichi Uchida
ICFHR2
2011 A New Approach for Instance-Based Skew Estimation
Soma Shiraishi, Yaokai Feng, Seiichi Uchida
KES (4)1