Shin Katayama

dblp:229/1624 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0002-5614-4412ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Efficient Edge AI Based Annotation and Detection Framework for Logistics Warehouses
abstract
As global logistics demand increases, improving the efficiency of warehouse operations has become critical. To achieve this, it's crucial to identify inefficient tasks and layouts by recognizing various warehouse conditions. To address these challenges, we've constructed a large-scale camera infrastructure to convert object positions and movements into data. However, transmitting all video data to the cloud results in significant data transmission and power consumption. Edge AI cameras analyze and extract video data locally, transmitting only essential information and significantly reducing them. Edge AI cameras require a low-computation, high-accuracy object detection model due to limited computational power and complex warehouse environments. Furthermore, the same object appears differently based on the camera's position and angle. Therefore, customizing models for each camera improves accuracy, but the annotation cost would be very high. In this study, we propose a method to perform part of the training data generation on the camera, reducing data transmission and improving annotation efficiency.
Yuki Mori 0001, Yusuke Asai, Keisuke Higashiura, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
CCNC4
2025 CorVS: Person Identification via Video Trajectory-Sensor Correspondence in a Real-World Warehouse
abstract
Worker location data is key to higher productivity in industrial sites. Cameras are a promising tool for localization in logistics warehouses since they also offer valuable environmental contexts such as package status. However, identifying individuals with only visual data is often impractical. Accordingly, several prior studies identified people in videos by comparing their trajectories and wearable sensor measurements. While this approach has advantages such as independence from appearance, the existing methods may break down under real-world conditions. To overcome this challenge, we propose CorVS, a novel data-driven person identification method based on correspondence between visual tracking trajectories and sensor measurements. Firstly, our deep learning model predicts correspondence probabilities and reliabilities for every pair of a trajectory and sensor measurements. Secondly, our algorithm matches the trajectories and sensor measurements over time using the predicted probabilities and reliabilities. We developed a dataset with actual warehouse operations and demonstrated the method’s effectiveness for real-world applications.
Kazuma Kano, Yuki Mori 0001, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
IPIN3
2025 Digitization Methods for a Logistics Warehouse Towards Digital Twin-Driven Optimization
abstract
The rapid growth of e-commerce, rising consumer expectations, and labor shortages in Japan pose challenges for logistics warehouse optimization. While automation has advanced outbound operations, inbound processes remain inefficient. This study presents a comprehensive approach to digitizing and optimizing large-scale logistics warehouses, with a case study at TRUSCO NAKAYAMA Corp. in partnership with Nagoya University. Key technologies include a large-scale camera array, multi-camera object tracking, and smartphone-based task estimation. Our contributions include real-time tracking of personnel and packages, cooperative annotation for improved object recognition, and synthetic data augmentation. Additionally, truck berth analysis and indoor localization enhance operational efficiency. To optimize worker shifts and warehouse layout, we apply Factorization Machine Quantum Annealing (FMQA), achieving a 37.4% reduction in lead times and a 14.3% decrease in labor hours. A visualization tool enables warehouse operators to make data-driven decisions. This research demonstrates the potential of digital transformation in logistics and provides a scalable framework for broader industry adoption.
Nobuo Kawaguchi, Yusuke Asai, Kazuma Kano, Kairi Takaki, Yuki Mori 0001, Yuma Suzuki, Kisho Watanabe, Yuki Gushi, Shin Katayama, Kenta Urano, Takuro Yonezawa, Shintaro Hashiguchi
SMARTCOMP9
2025 Joint Black-Box Optimization of Warehouse Layout and Worker Assignment Using Quantum Annealing and Factorization Machines
abstract
As global demand for logistics continues to grow, improving the efficiency of warehouse operations has become increasingly important. While many processes in logistics warehouses have been automated, the receiving area still relies heavily on manual work. Therefore, optimizing this area is essential. When optimizing operations in a logistics warehouse, various problems must be considered, such as layout design and worker assignment. Previous research has typically focused on these problems individually. However, because they are closely related, jointly optimizing them can further improve overall efficiency. In this study, we focus on the receiving area and conduct joint optimization of layout design and worker assignment. Specifically, we extend and apply a black-box optimization method called FMQA, which combines Factorization Machines (FM) with Quantum Annealing (QA). By using Factorization Machine regression, we build an objective function from simulator data. We then use Quantum Annealing to minimize it and improve the receiving area. A comparison with an actual warehouse environment, using the multi-agent simulator, shows that our approach can reduce mean package processing time by up to 22.1%. This result demonstrates the effectiveness of joint optimization of layout design and worker assignment.
Kairi Takaki, Yusuke Asai, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
SMC3
2024 Unveiling Human Attributes through Life Pattern Clustering using GPS Data Only
abstract
Clustering people by their life patterns is valuable in government and business fields. Existing studies often rely on semantic data such as Point of Interest or stay purpose. However, they have the problem that obtaining large datasets is difficult due to the need for annotation work. Some studies try to use only location data. However, they do not reveal the semantics of the area where visitors stay because they only label visited areas by significance according to duration and frequency of stay. In this paper, we propose a framework, LPSeL, for clustering people's Life Patterns at a Semantic Level using only raw GPS location data. LPSeL is based on the idea that analyzing human mobility first requires understanding urban space. Therefore, it begins with area modeling, which models areas in a city based on people's activities. Then, treating human mobility as a sequence of area representations makes it possible to model individuals by semantic-level characteristics of their life patterns. We showed that LPSeL is capable of estimating people's attributes from their life patterns using a real-world dataset consisting of GPS data collected from tens of thousands of smartphone users.
Kazuyuki Shoji, Haru Terashima, Nobuo Kawaguchi, Shin Katayama, Kenta Urano, Takuro Yonezawa, Naoki Tamura
SIGSPATIAL/GIS4
2024 Additive Compositionality in Urban Area Embeddings Based on Human Mobility Patterns
abstract
Understanding the characteristics of various urban areas is crucial for applications such as urban planning, tourism policies, market analysis, and infection control. Techniques for embedding areas as vectors in a latent space based on human mobility patterns are actively researched. Many of these area embedding methods define areas as points, grids, or polygons on a geospatial plane and then embed them. However, existing methods do not allow for mutual transformation between these forms and sizes after the initial embedding. Additionally, if the characteristics of an area change due to events such as the opening of new buildings, re-embedding is necessary. Meanwhile, the Word2Vec technique, a representative word embedding method, has a property called additive compositionality. This property allows for the arithmetic operation of word meanings through the arithmetic operations of word embeddings. In this paper, we propose a method to apply this property to existing area embedding techniques, leveraging it for practical tasks such as area shape transformation and searching for areas with trends change.
Naoki Tamura, Haru Terashima, Kazuyuki Shoji, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
SIGSPATIAL/GIS4
2024 Demo: Assisting System for Creating Ceiling Plan Using a Video from a Smatrphone
abstract
We present an assisting system for creating a ceiling plan. Conventional methods of creating a ceiling plan are time-consuming and high-cost. Our system requires only two inputs from a user and outputs the panoramic ceiling image that shows the whole ceiling surface. The system detects the ceiling fixtures and depicts them seamlessly for a reliable resulting image. We confirmed the possibility of assisting in creating a ceiling plan with our system through the experiment.
Daiki Kohama, Yoshiteru Nagata, Kazushige Yasutake, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
MobiSys4
2024 Poster: Sustainable Data Management Flow for Spatio-Temporal Datasets
abstract
Spatio-temporal data is utilized in various fields, but its scale is continuously growing, leading to significant labor and costs in storage and processing. Therefore, the value that can be derived from spatio-temporal data is diluted due to management costs. We propose a new data management flow using various metadata and common programs for spatio-temporal data utilization. Traditionally, various spatio-temporal data processing have been implemented and processed according to each spatio-temporal data. We defined spatio-temporal data structure metadata and performed data processing based on metadata using a common data processing program. Furthermore, we automated the generation of data structure metadata by combining our data skeleton recognition method and generative AI model. Using this flow, we expect to improve the sustainability of utilizing spatio-temporal data.
Yoshiteru Nagata, Daiki Kohama, Yoshiki Watanabe, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
MobiSys4
2024 Semi-Automated Framework for Digitalizing Multi-Product Warehouses with Large Scale Camera Arrays
abstract
As global demand for logistics continues to grow, optimizing the automation and efficiency of distribution warehouse operations is of paramount importance. Digitalizing warehouse environments, which refers to the process of sensing the physical space and extracting meaningful information from the obtained data, offers a promising solution to this challenge. However, converting raw warehouse data, such as video footage captured inside the warehouse, into actionable metadata (e.g., tracking the movement paths of workers and products or analyzing the usage patterns of different warehouse locations) often necessitates significant human intervention. The rise of machine learning further complicates this, as it requires the manual preparation of extensive training datasets. In this paper, we introduce a framework that semi-automates the digitalization process in complex warehouse settings. This framework employs dense optical flow and representation learning to autonomously segment warehouse objects and cluster similar objects, thereby substantially cutting down on annotation costs. To evaluate our approach, we constructed a large-scale data collection platform with over 60 fixed cameras in a real-world logistics warehouse, and the video data from this platform was then applied to our framework. Our evaluations indicate that our method markedly reduces both the time and resources required for warehouse digitalization using the captured video data.
Keisuke Higashiura, Kodai Yokoyama, Yusuke Asai, Hironori Shimosato, Kazuma Kano, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
PerCom6
2022 ER-Chat: A Text-to-Text Open-Domain Dialogue Framework for Emotion Regulation
abstract
Emotions are essential for constructing social relationships between humans and interactive systems. Although emotional and empathetic dialogue generation methods have been proposed for dialogue systems, appropriate dialogue involves not only mirroring emotions and always being empathetic but also complex factors such as context. This paper proposes Emotion Regulation Chat (ER-Chat) as an end-to-end dialogue framework for emotion regulation. Emotion regulation is concerned with actions to approach appropriate emotional states. Learning appropriate emotion and intent when responding on the basis of the context of the dialogue enables the generation of more human-like dialogue. We conducted automatic and human evaluations to demonstrate the superiority of ER-Chat over the baseline system. The results show that inclusion of emotion and intent prediction mechanisms enable generation of dialogues with greater fluency, diversity, emotion awareness, and emotion appropriateness, which are greatly preferred by humans.
Shin Katayama, Shunsuke Aoki 0001, Takuro Yonezawa, Tadashi Okoshi, Jin Nakazawa, Nobuo Kawaguchi
IEEE Trans. Affect. Comput.1
2019 Situation-Aware Emotion Regulation of Conversational Agents with Kinetic Earables
abstract
Conversational agents are increasingly becoming digital partners of our everyday computing experiences offering a variety of purposeful information and utility services. Although rich on competency, these agents are entirely oblivious to their users' situational and emotional context today and incapable of adjusting their interaction style and tone contextually. To this end, we present a mixed-method study that informs the design of a situation- and emotion-aware conversational agent for kinetic earables. We surveyed 280 users, and qualitatively interviewed 12 users to understand their expectation from a conversational agent in adapting the interaction style. Grounded on our findings, we develop a first-of-its-kind emotion regulator for a conversational agent on kinetic earable that dynamically adjusts its conversation style, tone, volume in response to users emotional, environmental, social and activity context gathered through speech prosody, motion signals and ambient sound. We describe these context models, the end-to-end system including a purpose-built kinetic earable and their real-world assessment. The experimental results demonstrate that our regulation mechanism invariably elicits better and affective user experience in comparison to baseline conditions in different real-world settings.
Shin Katayama, Akhil Mathur, Marc Van den Broeck, Tadashi Okoshi, Jin Nakazawa, Fahim Kawsar
ACII1
2019 Situation-Aware Conversational Agent with Kinetic Earables
abstract
Conversational agents are increasingly becoming digital partners of our everyday computing experiences offering a variety of purposeful information and utility services. Although rich on competency, these agents are entirely oblivious to their users' situational and emotional context today and incapable of adjusting their interaction style and tone contextually. To this end, we present a first-of-its-kind situation-aware conversational agent on kinetic earable that dynamically adjusts its conversation style, tone, volume in response to users emotional, environmental, social and activity context gathered through speech prosody, ambient sound and motion signatures.
Shin Katayama, Akhil Mathur, Tadashi Okoshi, Jin Nakazawa, Fahim Kawsar
MobiSys1
2019 Motivating Long-term Dietary Habit Modification through Mobile MR Gamification
abstract
In correlation with the socio-economic development, changes in people's lifestyle brought about significant impact on dietary patterns. Though public concerns over healthy eating are increasing, many are still uncertain when choosing a well balanced meal amid welter of information. In this paper, we propose "KomaFLens'', a mobile system and application built for Microsoft HoloLens, which aims to motivate long term dietary habit modification through gamification. The primary purpose of this research is to enhance the users' nutritional knowledge and to guide them to make healthier choices in their diet. Our preliminary evaluation revealed interesting points for discussion regarding the procedure for capturing food labels. Streamlining the operational method to boost tractability will improve the accuracy when recording food intakes.
Kento Katsumata, Yusaku Eigen, Yuka Noda, Masayoshi Tsuruoka, Satsuki Hashiba, Shotaro Numoto, Shin Katayama, Tadashi Okoshi, Jin Nakazawa
MobiSys7