Nobuo Kawaguchi

dblp:85/5677 · DBLP profile ↗
← Back
52ranked-venue papers
11as first author
19since 2021 · last 2026
0000-0002-0444-2290ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-authorHuman-computer interaction and ubiquitous computing · 8 · 3 first-author · 4 since 2021Computer networks · 6 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Bridging Strange and Familiar: Design of Knowledge Expansion System via Daily Surroundings
Kohei Matsumoto, Hideki Deguchi, Yoshiki Watanabe, Kaiya Shimura, Nozomi Hayashida, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
COMPSAC8
2026 Towards Circular Accumulation of Latent Local Knowledge via Automated Interview System: A Preliminary Study
Aoi Sassa, Haru Terashima, Naoki Tamura, Kazuyuki Shoji, Kenta Urano, Takuro Yonezawa, Tadashi Yoshikawa, Nobuo Kawaguchi
COMPSAC8
2026 Time Series Forecasting of Sports Ticket Sales with Venue-Independent Fan Segments
Haru Terashima, Naoki Tamura, Kazuyuki Shoji, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
DATA (1)7
2025 Efficient Edge AI Based Annotation and Detection Framework for Logistics Warehouses
abstract
As global logistics demand increases, improving the efficiency of warehouse operations has become critical. To achieve this, it's crucial to identify inefficient tasks and layouts by recognizing various warehouse conditions. To address these challenges, we've constructed a large-scale camera infrastructure to convert object positions and movements into data. However, transmitting all video data to the cloud results in significant data transmission and power consumption. Edge AI cameras analyze and extract video data locally, transmitting only essential information and significantly reducing them. Edge AI cameras require a low-computation, high-accuracy object detection model due to limited computational power and complex warehouse environments. Furthermore, the same object appears differently based on the camera's position and angle. Therefore, customizing models for each camera improves accuracy, but the annotation cost would be very high. In this study, we propose a method to perform part of the training data generation on the camera, reducing data transmission and improving annotation efficiency.
Yuki Mori 0001, Yusuke Asai, Keisuke Higashiura, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
CCNC7
2025 Multi-City Next Location Prediction through Mobility-Derived Multi-Pattern Transfer Learning
abstract
In this study, we propose MoDeMIT (Mobility-Derived Multi-Insight Transfer), a human mobility prediction method that transfers various patterns shared across cities, extracted from mobility histories. Existing studies on human mobility prediction have primarily focused on learning the mobility patterns of users in the target city, without considering applications in cities with limited data. Moreover, relying solely on mobility patterns poses inherent limitations on prediction accuracy. The proposed MoDeMIT addresses these issues by defining and transferring multiple patterns shared across cities, such as lifestyle patterns and large-scale mobility patterns derived from mobility histories, as well as mobility patterns. This approach enables improvements in prediction accuracy compared to existing methods. We validate the effectiveness of MoDeMIT using real-world human mobility datasets. Furthermore, in the HuMob Challenge 2025 (GISCUP), MoDeMIT achieved a GEOBLEU score of [Average: 0.1632, CityA: 0.1504, CityB: 0.1471, CityC: 0.1801, CityD: 0.1753] and ranked within the top five teams.
Haru Terashima, Naoki Tamura, Kazuyuki Shoji, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
SIGSPATIAL/GIS6
2025 CorVS: Person Identification via Video Trajectory-Sensor Correspondence in a Real-World Warehouse
abstract
Worker location data is key to higher productivity in industrial sites. Cameras are a promising tool for localization in logistics warehouses since they also offer valuable environmental contexts such as package status. However, identifying individuals with only visual data is often impractical. Accordingly, several prior studies identified people in videos by comparing their trajectories and wearable sensor measurements. While this approach has advantages such as independence from appearance, the existing methods may break down under real-world conditions. To overcome this challenge, we propose CorVS, a novel data-driven person identification method based on correspondence between visual tracking trajectories and sensor measurements. Firstly, our deep learning model predicts correspondence probabilities and reliabilities for every pair of a trajectory and sensor measurements. Secondly, our algorithm matches the trajectories and sensor measurements over time using the predicted probabilities and reliabilities. We developed a dataset with actual warehouse operations and demonstrated the method’s effectiveness for real-world applications.
Kazuma Kano, Yuki Mori 0001, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
IPIN6
2025 Digitization Methods for a Logistics Warehouse Towards Digital Twin-Driven Optimization
abstract
The rapid growth of e-commerce, rising consumer expectations, and labor shortages in Japan pose challenges for logistics warehouse optimization. While automation has advanced outbound operations, inbound processes remain inefficient. This study presents a comprehensive approach to digitizing and optimizing large-scale logistics warehouses, with a case study at TRUSCO NAKAYAMA Corp. in partnership with Nagoya University. Key technologies include a large-scale camera array, multi-camera object tracking, and smartphone-based task estimation. Our contributions include real-time tracking of personnel and packages, cooperative annotation for improved object recognition, and synthetic data augmentation. Additionally, truck berth analysis and indoor localization enhance operational efficiency. To optimize worker shifts and warehouse layout, we apply Factorization Machine Quantum Annealing (FMQA), achieving a 37.4% reduction in lead times and a 14.3% decrease in labor hours. A visualization tool enables warehouse operators to make data-driven decisions. This research demonstrates the potential of digital transformation in logistics and provides a scalable framework for broader industry adoption.
Nobuo Kawaguchi, Yusuke Asai, Kazuma Kano, Kairi Takaki, Yuki Mori 0001, Yuma Suzuki, Kisho Watanabe, Yuki Gushi, Shin Katayama, Kenta Urano, Takuro Yonezawa, Shintaro Hashiguchi
SMARTCOMP1
2025 Joint Black-Box Optimization of Warehouse Layout and Worker Assignment Using Quantum Annealing and Factorization Machines
abstract
As global demand for logistics continues to grow, improving the efficiency of warehouse operations has become increasingly important. While many processes in logistics warehouses have been automated, the receiving area still relies heavily on manual work. Therefore, optimizing this area is essential. When optimizing operations in a logistics warehouse, various problems must be considered, such as layout design and worker assignment. Previous research has typically focused on these problems individually. However, because they are closely related, jointly optimizing them can further improve overall efficiency. In this study, we focus on the receiving area and conduct joint optimization of layout design and worker assignment. Specifically, we extend and apply a black-box optimization method called FMQA, which combines Factorization Machines (FM) with Quantum Annealing (QA). By using Factorization Machine regression, we build an objective function from simulator data. We then use Quantum Annealing to minimize it and improve the receiving area. A comparison with an actual warehouse environment, using the multi-agent simulator, shows that our approach can reduce mean package processing time by up to 22.1%. This result demonstrates the effectiveness of joint optimization of layout design and worker assignment.
Kairi Takaki, Yusuke Asai, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
SMC6
2024 Unveiling Human Attributes through Life Pattern Clustering using GPS Data Only
abstract
Clustering people by their life patterns is valuable in government and business fields. Existing studies often rely on semantic data such as Point of Interest or stay purpose. However, they have the problem that obtaining large datasets is difficult due to the need for annotation work. Some studies try to use only location data. However, they do not reveal the semantics of the area where visitors stay because they only label visited areas by significance according to duration and frequency of stay. In this paper, we propose a framework, LPSeL, for clustering people's Life Patterns at a Semantic Level using only raw GPS location data. LPSeL is based on the idea that analyzing human mobility first requires understanding urban space. Therefore, it begins with area modeling, which models areas in a city based on people's activities. Then, treating human mobility as a sequence of area representations makes it possible to model individuals by semantic-level characteristics of their life patterns. We showed that LPSeL is capable of estimating people's attributes from their life patterns using a real-world dataset consisting of GPS data collected from tens of thousands of smartphone users.
Kazuyuki Shoji, Haru Terashima, Nobuo Kawaguchi, Shin Katayama, Kenta Urano, Takuro Yonezawa, Naoki Tamura
SIGSPATIAL/GIS3
2024 Additive Compositionality in Urban Area Embeddings Based on Human Mobility Patterns
abstract
Understanding the characteristics of various urban areas is crucial for applications such as urban planning, tourism policies, market analysis, and infection control. Techniques for embedding areas as vectors in a latent space based on human mobility patterns are actively researched. Many of these area embedding methods define areas as points, grids, or polygons on a geospatial plane and then embed them. However, existing methods do not allow for mutual transformation between these forms and sizes after the initial embedding. Additionally, if the characteristics of an area change due to events such as the opening of new buildings, re-embedding is necessary. Meanwhile, the Word2Vec technique, a representative word embedding method, has a property called additive compositionality. This property allows for the arithmetic operation of word meanings through the arithmetic operations of word embeddings. In this paper, we propose a method to apply this property to existing area embedding techniques, leveraging it for practical tasks such as area shape transformation and searching for areas with trends change.
Naoki Tamura, Haru Terashima, Kazuyuki Shoji, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
SIGSPATIAL/GIS7
2024 Demo: Assisting System for Creating Ceiling Plan Using a Video from a Smatrphone
abstract
We present an assisting system for creating a ceiling plan. Conventional methods of creating a ceiling plan are time-consuming and high-cost. Our system requires only two inputs from a user and outputs the panoramic ceiling image that shows the whole ceiling surface. The system detects the ceiling fixtures and depicts them seamlessly for a reliable resulting image. We confirmed the possibility of assisting in creating a ceiling plan with our system through the experiment.
Daiki Kohama, Yoshiteru Nagata, Kazushige Yasutake, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
MobiSys7
2024 Poster: Sustainable Data Management Flow for Spatio-Temporal Datasets
abstract
Spatio-temporal data is utilized in various fields, but its scale is continuously growing, leading to significant labor and costs in storage and processing. Therefore, the value that can be derived from spatio-temporal data is diluted due to management costs. We propose a new data management flow using various metadata and common programs for spatio-temporal data utilization. Traditionally, various spatio-temporal data processing have been implemented and processed according to each spatio-temporal data. We defined spatio-temporal data structure metadata and performed data processing based on metadata using a common data processing program. Furthermore, we automated the generation of data structure metadata by combining our data skeleton recognition method and generative AI model. Using this flow, we expect to improve the sustainability of utilizing spatio-temporal data.
Yoshiteru Nagata, Daiki Kohama, Yoshiki Watanabe, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
MobiSys7
2024 Semi-Automated Framework for Digitalizing Multi-Product Warehouses with Large Scale Camera Arrays
abstract
As global demand for logistics continues to grow, optimizing the automation and efficiency of distribution warehouse operations is of paramount importance. Digitalizing warehouse environments, which refers to the process of sensing the physical space and extracting meaningful information from the obtained data, offers a promising solution to this challenge. However, converting raw warehouse data, such as video footage captured inside the warehouse, into actionable metadata (e.g., tracking the movement paths of workers and products or analyzing the usage patterns of different warehouse locations) often necessitates significant human intervention. The rise of machine learning further complicates this, as it requires the manual preparation of extensive training datasets. In this paper, we introduce a framework that semi-automates the digitalization process in complex warehouse settings. This framework employs dense optical flow and representation learning to autonomously segment warehouse objects and cluster similar objects, thereby substantially cutting down on annotation costs. To evaluate our approach, we constructed a large-scale data collection platform with over 60 fixed cameras in a real-world logistics warehouse, and the video data from this platform was then applied to our framework. Our evaluations indicate that our method markedly reduces both the time and resources required for warehouse digitalization using the captured video data.
Keisuke Higashiura, Kodai Yokoyama, Yusuke Asai, Hironori Shimosato, Kazuma Kano, Shin Katayama, Kenta Urano, Takuro Yonezawa, Nobuo Kawaguchi
PerCom9
2024 Performance Evaluation of KNIME Low Code Platform in Deep Learning Study and Optimal Hyperparameter Tuning
abstract
A low-code platform is a software development environment that allows for the creation of applications through graphical user interfaces and configuration instead of traditional hand-coded computer programming. In this study, the application to classify a dataset of traffic sign images using the KNIME low-code deep learning development platform will be discussed to represents this software performance especial in term of model optimization processes. By creates the workflow to perform image preprocessing, create the CNN layer under KERAS sequential API and finding the best set of key hyperparameters among traditional KNIME build-in optimization algorithm including Brute force, Hill climbing, random search, Bayesian Optimization and black-box optimizer Optuna optimization algorithm under 3 types CNN architecture as simple CNN, Resnet-50 and VGG16 to classify traffic sign images. The result demonstrates that both grid search and random search optimization can be effective, while both Optuna and Bayesian optimization stands out as a powerful method due to its ability to efficiently explore the hyperparameter space and achieve superior results to meet 99% accuracy under simple CNN environment, but Optuna is significantly improve optimization times than Bayesian about 7 - 8 times. The KNIME low-code platform provides a user-friendly environment for developing and fine-tuning models to contribute the ongoing progress in machine learning and deep learning research development.
Pornpawee Thongkhome, Takuro Yonezawa, Nobuo Kawaguchi
TENCON3
2023 IRIS: Interpretable Rubric-Informed Segmentation for Action Quality Assessment
abstract
AI-driven Action Quality Assessment (AQA) of sports videos can mimic Olympic judges to help score performances as a second opinion or for training. However, these AI methods are uninterpretable and do not justify their scores, which is important for algorithmic accountability. Indeed, to account for their decisions, instead of scoring subjectively, sports judges use a consistent set of criteria — rubric — on multiple actions in each performance sequence. Therefore, we propose IRIS to perform Interpretable Rubric-Informed Segmentation on action sequences for AQA. We investigated IRIS for scoring videos of figure skating performance. IRIS predicts (1) action segments, (2) technical element score differences of each segment relative to base scores, (3) multiple program component scores, and (4) the summed final score. In a modeling study, we found that IRIS performs better than non-interpretable, state-of-the-art models. In a formative user study, practicing figure skaters agreed with the rubric-informed explanations, found them useful, and trusted AI judgments more. This work highlights the importance of using judgment rubrics to account for AI decisions.
Hitoshi Matsuyama, Nobuo Kawaguchi, Brian Y. Lim
IUI2
2022 ER-Chat: A Text-to-Text Open-Domain Dialogue Framework for Emotion Regulation
abstract
Emotions are essential for constructing social relationships between humans and interactive systems. Although emotional and empathetic dialogue generation methods have been proposed for dialogue systems, appropriate dialogue involves not only mirroring emotions and always being empathetic but also complex factors such as context. This paper proposes Emotion Regulation Chat (ER-Chat) as an end-to-end dialogue framework for emotion regulation. Emotion regulation is concerned with actions to approach appropriate emotional states. Learning appropriate emotion and intent when responding on the basis of the context of the dialogue enables the generation of more human-like dialogue. We conducted automatic and human evaluations to demonstrate the superiority of ER-Chat over the baseline system. The results show that inclusion of emotion and intent prediction mechanisms enable generation of dialogues with greater fluency, diversity, emotion awareness, and emotion appropriateness, which are greatly preferred by humans.
Shin Katayama, Shunsuke Aoki 0001, Takuro Yonezawa, Tadashi Okoshi, Jin Nakazawa, Nobuo Kawaguchi
IEEE Trans. Affect. Comput.6
2021 How to monitor multiple autonomous vehicles remotely with few observers: An active management method
abstract
In this research, we proposed an active management method to tele-monitor and tele-operate more autonomous vehicles (AVs) with few observers by adjusting the movement of the AVs actively. A management system is created to get the status from the AVs and separate the monitoring requirement to the observers optimally. When the requirements might be intensive, the management system can also adjust the movements of the AVs actively to distribute the monitoring time, which can make the observers monitor more vehicles. We implement and verify the management system in an autonomous driving simulator - CARLA for the limited number of AVs and observers. Based on the data acquired from the driving simulator, we also create a numerical simulator and tested our method with the pairs of a large number of AVs and observers. The result of both simulators shows that our method can reduce the utilization degree of observers and make them monitor more AVs.
Ming Ding 0002, Eijiro Takeuchi, Yoshio Ishiguro, Yoshiki Ninomiya, Nobuo Kawaguchi, Kazuya Takeda
IV5
2021 DigiMobot: Digital Twin for Human-Robot Collaboration in Indoor Environments
abstract
Human-robot collaboration and cooperation are critical for Autonomous Mobile Robots (AMRs) in order to use them in indoor environments, such as offices, hospitals, libraries, schools, factories, and warehouses. Since a long transition period might be required to fully automate such facilities, we have to deploy AMRs while improving safety in the mixed environments of human and mobile robots. In addition, human behaviors in such environments might be difficult to predict. In this paper, we present a Digital Twin for Autonomous Mobile Robots system named DigiMobot to support, manage, monitor, and validate AMRs in indoor environments. First, DigiMobot can simulate human behaviors and robot movements to verify and validate AMRs to improve safety in a virtual world. Secondly, DigiMobot can monitor and manage AMRs in the physical world by collecting sensor data from each robot in real-time. Since DigiMobot enables us to test the robot systems in the virtual world, we can deploy and implement AMRs in each facility without any modifications. To show the feasibility of DigiMobot, we develop a software framework and two different types of autonomous mobile robots. Finally, we conduct real-world experiments in a warehouse located in Saitama, Japan, in which more than 400, 000 items are stored for commercial purposes.
Yuto Fukushima, Yusuke Asai, Shunsuke Aoki 0001, Takuro Yonezawa, Nobuo Kawaguchi
IV5
2021 Synthetic People Flow: Privacy-Preserving Mobility Modeling from Large-Scale Location Data in Urban Areas
Naoki Tamura, Kenta Urano, Shunsuke Aoki 0001, Takuro Yonezawa, Nobuo Kawaguchi
MobiQuitous5
2019 Basic Study of BLE Indoor Localization using LSTM-based Neural Network
abstract
In this paper, LSTM-based neural network is applied to indoor localization using mobile BLE tag's signal strength collected by multiple scanners. Stability of signal strength is a critical factor of wireless indoor localization for higher accuracy. While traditional methods like trilateration and fingerprinting suffer from noise and packet loss, deep learning based methods perform well. We focus on large-scale exhibition where wireless signal gets unstable due to many people. Proposed neural network consists of fully connected layers for noise removal and LSTM layers for time-series feature extraction. The network takes the time-series of signal strength as input and outputs the estimated location. In the evaluation, the number of layers is changed to find the optimal structure. As a result, the best configuration achieved the error of 2.44m at 75 percentile for the data of a large-scale exhibition in Tokyo.
Kenta Urano, Kei Hiroi, Takuro Yonezawa, Nobuo Kawaguchi
MobiSys4
2019 Capturing Subjective Time as Context and It's Applications
abstract
We propose an integrated framework for sensing, recognizing and utilizing of subjective time as context. Various studies on experimental psychology have showed several factors which affects subjective time. Those factors should be partially captured by ubiquitous sensors such as smartphones and wearable devices, therefore, we tackle to create common and individual model for subjective time based on the sensor data. We report our first prototype implementation for the framework based on AWARE framework with adding experience sampling method for subjective time recognition. In addition, we discuss potential applications which leveraging advantages of subjective time as context.
Takuro Yonezawa, Yuuki Nishiyama, Kei Hiroi, Nobuo Kawaguchi
MobiSys4
2019 Gait Dependency of Smartphone Walking Speed Estimation using Deep Learning
abstract
This paper proposes an accurate estimation method of walking speed using deep learning for smartphone-based Pedestrian Dead Reckoning (PDR).PDR requires to estimate speed and direction of pedestrians accurately using accelerometer and gyroscope.To improve the accuracy of PDR, existing works focused to improve the key factors of speed estimation (i.e., stride and/or step estimation) by adapting deep learning.On the contrary, our research proposes to adapt deep learning more directly to estimate walking speed from sensor data of smartphone. We evaluate the accuracy of proposed method by comparing with conventional PDR method. As a result, we confirmed that proposed method can estimate the speed more accurately.
Takuto Yoshida, Junto Nozaki, Kenta Urano, Kei Hiroi, Takuro Yonezawa, Nobuo Kawaguchi
MobiSys6
2018 Scale Estimation of Monocular SfM for a Multi-modal Stereo Camera
Shinya Sumikura, Ken Sakurada, Nobuo Kawaguchi, Ryosuke Nakamura
ACCV (3)3
2018 Image Translation Between Sar and Optical Imagery with Generative Adversarial Nets
abstract
In this paper, we propose a method for the translation from Synthetic Aperture Radar (SAR) to optical images using conditional Generative Adversarial Networks (cGANs). Satellite images have been widely utilized for various purposes, such as natural environment monitoring (pollution, forest or rivers), transportation improvement and prompt emergency response to disasters. However, the obscurity caused by clouds leads to unstable monitoring of the ground situation while using the optical camera. Images captured by a longer wavelength are introduced to reduce the effects of clouds. In particular, SAR images are known to be nearly unaffected by clouds and are often used for stably observing the ground situation. On the other hand, SAR images have lower spatial resolution and visibility than optical images. Therefore, we propose a deep neural network that generates optical images from SAR images. Finally, we confirm the feasibility of the proposed network on a dataset consisting of optical images and the corresponding SAR images.
Kenji Enomoto, Ken Sakurada, Nobuo Kawaguchi, Masashi Matsuoka, Ryosuke Nakamura
IGARSS4
2016 PIEM: Path Independent Evaluation Metric for Relative Localization
abstract
There are many methods for indoor positioning. These methods are divided into the relative localization and absolute localization. In the relative localization, one widely used method is Pedestrian Dead Reckoning (PDR). Relative localization estimates the moving distance, orientation, and height of the pedestrian. However, relative localization has a problem caused by an accumulated error: the longer the path, the worse the accuracy of relative localization. There is another problem in the existing evaluating metrics: they compare only the actual location and the estimated location of the destination. Relative localization also has this evaluation problem. We propose PIEM: Path Independent Evaluation Metric for Relative Localization. PIEM is a path independent evaluation metric, considering the complexity of the path; distance, orientation, and height. Then we evaluate these three factors of relative localization in addition to the position. Our proposed method showed more consistent results for the complexity of the path than the existing methods of relative localization evaluation.
Masaaki Abe, Katsuhiko Kaji, Kei Hiroi, Nobuo Kawaguchi
IPIN4
2016 Estimating 3D pedestrian trajectories using stability of sensing signal
abstract
A highly accurate estimation method of 3-D pedestrian trajectories from walking activity sensing data is proposed. This method uses data from an accelerometer, a gyrometer, and an air pressure sensor, and does not require detailed information on the building structure. In activity sensing using wearable sensors, higher accuracy can be expected from detection of zones in which there is continuously little change in the states of the sensor signals than from detection of zones in which there are large changes in the sensor signals within a short time. We focus on such stability of sensing signal and, as an application example, we used the concept to estimate walking direction on the basis of stable walking zone detection using a gyrometer and estimation of movement between the floors of a building by detection of stable floor zones with an air pressure sensor. We then integrated these estimations to obtain a 3-D pedestrian trajectory. The results of evaluation experiments using an indoor pedestrian sensing corpus (HASC-IPSC) showed that this method achieved higher accuracy for both walking direction and movement between floors than was achieved by a method based on large changes in the sensor signals. We also confirmed that the cumulative error rate for estimation of the 3-D pedestrian trajectory was 1 m per 10 seconds of movement.
Katsuhiko Kaji, Nobuo Kawaguchi
IPIN2
2016 Wi-Fi Human Behavior Analysis and BLE Tag Localization: A Case Study at an Underground Shopping Mall
abstract
Techniques for obtaining customers' behavior (dwell time, count, and flow) in a shopping mall or large exhibit are highly sought-after by organizers or shop owners. Additionally, ways of effectively directing customers from cyber space such as the Web or smartphone apps to physical retail stores are also in high demand. "Online to Offline (O2O) Marketing" has recently been attracting attention to make these a reality. However, the know-how accumulated through "O2O Marketing" is not shared widely as it can be sensitive for business and customer's privacy. In this paper, we provide the knowledge obtained through the demonstration experiments held by the "O2O Digital Marketing Study Group" organized by NPO Lisra at a large underground shopping mall in Nagoya. 250 BLE tags and 12 Wi-Fi scanners were installed in this mall. We also organized a "coupon campaign" to increase the number of participants. We were able to effectively collect visiting customer count measurements via Wi-Fi scanners irrespective of the anonymization of BSSID, as our results showed similar trends collected by optical human detectors inherently installed at the venue. This paper provides preliminary insight into understanding the behaviors of retail shoppers and we believe this is a firm starting point for this area.
Nobuo Kawaguchi, Kei Hiroi, Atsushi Shionozaki, Masamichi Asukai, Toshimune Nasu, Yu Hashimoto, Takeharu Nakamura, Tetsuya Gotou, Shinsuke Ando
MobiQuitous1
2012 Gaussian Mixture Model and Particle Filter
abstract
Up to now Scene Analysis has been based on the WiFi location estimation technique and it has been necessary to have a large scale database and a large amount of calculation. We propose a WiFi estimation method that uses little data or calculation. First of all we apply Gaussian Mixture Model to represent the large scale WiFi database to decrease the WiFi data by no less than 95%. Secondly, we apply Particle Filter to adjust the possible calculation quantity needed for the location estimation technique. As experimental result, we achieved real-time location estimation within 6~10m. Another important issue for Scene Analysis technique is the high cost of operation of the previous WiFi observation. Accordingly crowdsourcing approach was used, employing as system where some users could contribute and other uses could share. The ideal system is a composition of the Web and mobile terminal. WiFi data observed by mobile terminals is uploaded to a Web server where it is managed and integrated into GMM and large scale operations are carried out on data and calculations. When the lightweight modeled data is downloaded to the mobile-terminal, the mobile terminal then has the ability to carry out real-time location estimation independently.
Katsuhiko Kaji, Nobuo Kawaguchi
IPIN2
2011 Distributed human activity data processing using HASC tool
abstract
To accelerate and simplify human activity recognition research, we have been developing a data processing tool named "HASC Tool." As the activity corpus becomes huge, it is not simple to handle the large number of files because it takes a lot of time to process. In this paper, we propose a distributed data processing mechanism which is implemented in the HASC Tool. By using the system, we can simply scale the local system into distributed processing. We also show the preliminary experimental result.
Nobuo Kawaguchi, Nobuhiro Ogawa, Yohei Iwasaki, Katsuhiko Kaji
UbiComp1
2011 HASC2011corpus: towards the common ground of human activity recognition
abstract
Human activity recognition through the wearable sensor will enable a next-generation human-oriented ubiquitous computing. However, most of research on human activity recognition so far is based on small number of subjects, and non-public data. To overcome the situation, we have gathered 4897 accelerometer data with 116 subjects and compose them as HASC2011corpus. In the field of pattern recognition, it is very important to evaluate and to improve the recognition methods by using the same dataset as a common ground. We make the HASC2011corpus into public for the research community to use it as a common ground of the Human Activity Recognition. We also show several facts and results of obtained from the corpus.
Nobuo Kawaguchi, Tianhui Yang, Nobuhiro Ogawa, Yohei Iwasaki, Katsuhiko Kaji, Tsutomu Terada, Kazuya Murao, Sozo Inoue, Yoshihiro Kawahara, Yasuyuki Sumi, Nobuhiko Nishio
UbiComp1
2009 NAT Free Open Source 3D Video Conferencing using SAMTK and Application Layer Router
abstract
SAMTK: scalable adaptive multicast toolkit is a toolkit to bridge the gap between network researchers and application developers in the field of multi-point communication. SAMTK includes a Qt based GUI that ensures a single source multi-platform use (Linux, FreeBSD, Windows and MacOS). It provides interfaces for network plugins used by one-to-many network sockets and a simple application programming interface for application developers to develop multi-point communication applications quite easily. ALR: application layer router is a router which parses UDP packets and does a lookup in its internal forwarding table to duplicate and deliver the packets to multiple destinations. It provides NAT traversal function by using a singe UDP port both for session registration and packet delivery. The demo shows the feasibility of NAT free 3D video conferencing using application layer router and the ease of development of video conferencing applications using SAMTK.
Nobuo Kawaguchi, Shuntaro Nishiura, Odira Elisha Abade, Takahiro Kurosawa, Tatsuya Jinmei, Eiichi Muramoto
CCNC1
2009 Underground Positioning: Subway Information System Using WiFi Location Technology
abstract
We introduce a subway information system which utilize WiFi location technology for supporting a person in the underground. The system is composed of a mobile terminal with a WiFi device and a communication server. We have developed seven location aware applications for the mobile terminal. Each of the application helps the user with current location information. We have performed a demonstration experiment in the subway of Nagoya City with 35 subjects and got a positive acceptance of the system.
Nobuo Kawaguchi, Motoki Yano, Shogo Ishida, Takeshi Sasaki, Yohei Iwasaki, Kenji Sugiki, Shigeki Matsubara
Mobile Data Management1
2006 Layered Speech-Act Annotation for Spoken Dialogue Corpus
Yuki Irie, Shigeki Matsubara, Nobuo Kawaguchi, Yukiko Yamaguchi, Yasuyoshi Inagaki
LREC3
2006 Data Correction Method Using Ideal Wireless LAN Model in Positioning System
abstract
Over the last few years, many positioning systems and information support systems using wireless LAN have been developed. Some systems use the received signal strength of a wireless LAN for positioning. However, the received signal strength differs depending on each terminal's wireless LAN adapter. It is important to investigate the differences among the received signal strengths for different wireless LAN adapters, because the differences among each adapter may cause an increase in location estimation errors. In this paper, we examine the differences among the received signal strengths for a number of wireless LAN adapters, and propose wireless LAN data usage using an ideal wireless LAN adapter. Using our method, we can improve the estimation accuracy of the location estimation system
Seigo Ito, Nobuo Kawaguchi
PIMRC2
2005 Construction and utilization of bilingual speech corpus for simultaneous machine interpretation research
abstract
This paper describes the design, analysis and utilization of a simultaneous interpretation corpus. The corpus has been constructed at the Center for Integrated Acoustic Information Research (CIAIR) of Nagoya University in order to promote the realization of the multi-lingual communication supporting environment. The size of transcribed data is about 1 million words, and the corpus would deserve to be called the simultaneous interpretation corpus of the largest-in-the-world class. The discourse tag and the utterance time tag were given to the corpus, and some software tools for corpus analysis in order to support the practical use of the corpus have been developed. Therefore, the corpus is expected to be useful not only for the development of simultaneous interpreting systems but also for the construction of an interpreting theory.
Hitomi Tohyama, Shigeki Matsubara, Nobuo Kawaguchi, Yasuyoshi Inagaki
INTERSPEECH3
2005 Performance Evaluation of H.264 Video Streaming over Inter-Vehicular 802.11 Ad Hoc Networks
abstract
This paper evaluates the performance of video streaming in inter-vehicular environments using the 802.11 ad hoc network protocol. We performed transmission experiments while driving two cars equipped with 802.11b standard devices in urban and highway scenarios. Different sequences, bitrates and packctization policies have been tested. The experiments show that each scenario presents peculiar characteristics in terms of average link availability and SNR, which can be exploited to develop more efficient applications. In this paper we also determine the best packetization policies for the two scenarios, showing that large packets lead to better performance in the highway scenario and vice versa. Perceptual quality results indicate that the best packetization policy achieves consistent gains in terms of PSNR values (up to 5 dB), and reduced quality variations, with respect to a fixed-policy transmission technique
Paolo Bucciol, Enrico Masala, Nobuo Kawaguchi, Kazuya Takeda, Juan Carlos De Martin
PIMRC3
2004 Azim: Direction Based Service Using Azimuth Based Position Estimation
abstract
Here, we propose a system named Azim that provides service based on both location and direction, which uses position estimation method based on azimuth data. In this system, a user's position is estimated by having the user point to and measure azimuths of several markers or objects whose positions are already known. Because the measurements are naturally associated with some degree of error, the user's position is calculated as a probability distribution. Since both the user's position and azimuth data are obtained in this method, both sources of data are used to realize more advanced services such as identifying the object pointed to by a user. The proposed system utilizes a wireless LAN for supporting these advanced services. Finally, a prototype system was implemented using a direction sensor that combines a magnetic compass and accelerometer, and we exemplify the usefulness of our approach through an experiment.
Yohei Iwasaki, Nobuo Kawaguchi, Yasuyoshi Inagaki
ICDCS2
2004 Speech understanding, dialogue management and response generation in corpus-based spoken dialogue system
abstract
This paper presents construction of a spoken dialogue system using a large-scale spoken dialogue corpus with intention tags. In this system, all of main components, such as speech understanding, dialogue management, and response generation, are constructed with corpus-based methods. An evaluation experiment using a test set has shown that the performance of the corpus-based dialogue system is improved by adding examples.
Keita Hayashi, Yuki Irie, Yukiko Yamaguchi, Shigeki Matsubara, Nobuo Kawaguchi
INTERSPEECH5
2004 Speech intention understanding based on decision tree learning
abstract
Abstract This paper proposes a method of speech intention understand-ing based on a spoken dialogue corpus to which the intentiontags are given. The intention tag expresses the task-dependentintention of the speaker, and therefore, the proper understand-ing enables a spoken dialogue system to take appropriate ac-tions. We have tagged about 35000 utterances in the CIAIR in-car speech database. In our method, several decision trees forintention understanding are constructed. By constructing deci-sion trees and using them at the same time, the strong amount ofcharacteristic features related to intentions can be retrieved, andit can also be robustly coped with the diversity of the utterances.An experiment on inference of utterance intentions has shown73.1% accuracy. 1. Introduction In order to interact with a user naturally and smoothly, it is nec-essary for a spoken dialogue system to understand the intentionof the user exactly. As a method of speech intention under-standing, example-based approaches have been considered sofar [1, 3, 6]!%In general, these approaches involve comparinga spoken utterance with examples in a correctly-tagged corpus.The intention of the utterance is regarded as the intention tag ofthe most similar example in the corpus. However, it is difficultto infer the intention of the utterance to which any example inthe corpus is not similar.This paper proposes a method of speech intention under-standing based on a spoken dialogue corpus. The method con-structs several decision trees for intention understanding. Byconstructing several decision trees, the strong amount of char-acteristic features related to intentions can be retrieved, and itcan also be robustly coped with the diversity of the utterances.So far, we have designed an organization of the tags which iscalled Layered Intention Tag(LIT). These tags show more de-tailed utterance intention rather than the illocutionary act level,and have built the corpus[2, 3]. LIT is divided into several lay-ers considering the relevance between an intention and variousphenomena relevant to an utterance, such as a style, a keyword,a sentence structure. This method constructs several decisiontrees by using this corpus and infers the intention by combiningthem.In order to evaluate the effectiveness of our method, an ex-periment on inference of the utterance intentions was conductedusing the driver utterances about restaurant search recorded ona large-scale in-car spoken dialogue corpus of CIAIR[4, 5]. Asa result, the effectiveness of the method was confirmed.
Yuki Irie, Shigeki Matsubara, Nobuo Kawaguchi, Yukiko Yamaguchi, Yasuyoshi Inagaki
INTERSPEECH3
2004 CIAIR in-car speech database
Nobuo Kawaguchi, Shigeki Matsubara, Yukiko Yamaguchi, Kazuya Takeda, Fumitada Itakura
INTERSPEECH1
2004 Example-based spoken dialogue system with online example augmentation
abstract
In this paper, we propose a new method to expand an examplebased spoken dialogue system to handle context dependent utterances. The dialogue system refers to the dialogue examples to find an example that is suitable to promote dialogue. Here, the dialogue contexts are expressed in the form of dialogue slots. By constructing dialogue examples with the text of utterances and the dialogue slots, the system handle context dependent dialogue. And we also propose a new framework of spoken dialogue, named “GROW architecture” that consists of the dialogue system and a Wizard-of-OZ (WOZ) system. By using the WOZ system to add dialogue examples via network, it becomes efficient to augment dialogue examples.
Hiroya Murao, Nobuo Kawaguchi, Shigeki Matsubara, Yukiko Yamaguchi, Kazuya Takeda, Yasuyoshi Inagaki
INTERSPEECH2
2004 Robust dependency parsing of spontaneous Japanese speech and its evaluation
abstract
Grant-in-Aids for Young Scientists of the Ministry of Education, Science, Sports and Culture, Japan; The Tatematsu Foundation
Tomohiro Ohno, Shigeki Matsubara, Nobuo Kawaguchi, Yasuyoshi Inagaki
INTERSPEECH3
2003 Construction of an advanced in-car spoken dialogue corpus and its characteristic analysis
abstract
This paper describes an advanced spoken language corpus which has been constructed by enhancing an in-car speech database. The corpus has the following characteristic features: (1) series Advanced tag: Not only linguistic phenomena tags but also advanced discourse tags such as sentential structures, and utterance intentions, have been provided for the transcribed texts. (2) series Large-scale: The sentential structures and the intentions are currently provided for 45,053 phrases and 35,421 utterance units, respectively. (3) series Multi-layer: The corpus consists of different levels of spoken language data such as speech signals, transcribed texts, sentential structures, intentional markers and dialogue structures, moreover, they are related with each other. It allows a very wide variety of analysis of spontaneous spoken dialogue to utilize the multi-layered corpus. This paper also reports the result of investigation of the corpus, especially, focusing on the relations between the syntactic style and the intentional style of spoken utterances.
Itsuki Kishida, Yuki Irie, Yukiko Yamaguchi, Shigeki Matsubara, Nobuo Kawaguchi, Yasuyoshi Inagaki
INTERSPEECH5
2003 Touch-and-Connect: A Connection Request Framework for Ad-Hoc Networks and the Pervasive Computing Environment
abstract
The spread of short distance wireless communication technology is making it possible for various information appliances to communicate through a network. However, to connect and to use them, we must generally perform complicated setup operations such as setting addresses or names. In this paper, we propose a user-friendly connecting management framework called " Touch-and-Connect". With this method, we can connect any networked devices only by touching them. This system is designed so that it can work without a server and even if many individuals operate it independently, no incorrect connections occur. And we have exemplified the usefulness of this framework with subject experiments.
Yohei Iwasaki, Nobuo Kawaguchi, Yasuyoshi Inagaki
PerCom2
2002 Example-based Speech Intention Understanding and Its Application to In-Car Spoken Dialogue System
Shigeki Matsubara, Shinichi Kimura, Nobuo Kawaguchi, Yukiko Yamaguchi, Yasuyoshi Inagaki
COLING3
2002 Stochastic Dependency Parsing of Spontaneous Japanese Spoken Language
Shigeki Matsubara, Takahisa Murase, Nobuo Kawaguchi, Yasuyoshi Inagaki
COLING3
2002 Multi-Dimensional Data Acquisition for Integrated Acoustic Information Research
Nobuo Kawaguchi, Shigeki Matsubara, Kazuya Takeda, Fumitada Itakura
LREC1
2002 Bilingual Spoken Monologue Corpus for Simultaneous Machine Interpretation Research
Shigeki Matsubara, Akira Takagi, Nobuo Kawaguchi, Yasuyoshi Inagaki
LREC3
2001 Multimedia data collection of in-car speech communication
abstract
This paper reports the details of the collection of the multimedia data such as audio, video and auxiliary information of the vehicle during a spoken dialogue in a moving car. The system specially built in a Data CollectionVehicle (DCV) supports synchronous recording of multi-channel audio data from 16 microphones, 3-channel video data and the vehicle related data. Multimedia data has been collected for three sessions of spoken dialogue in about a 60-minute drive by each of 200 subjects. Data has been collected for two dialogue modes:(1) prompted dialogue between the driver and an accompanying operator and (2) natural dialogue between the driver and a telephone operator for information access over a cellular phone while driving a car. The corpus can be used for analysis of multimedia data in a moving car environment and also for modeling spoken dialogue in scenarios such as information access while driving a car.
Nobuo Kawaguchi, Shigeki Matsubara, Kazuya Takeda, Fumitada Itakura
INTERSPEECH1
2000 Spoken language corpus for machine interpretation research
abstract
This paper describes a database consisting of speech and language, which we are currently constructing for the purpose of the research on machine interpretation. The database contains bilingual data of lectures and dialogues. We have collected the speech of about 72 hours in total and transcribed it into the text manually. We have investigated the database in order to acquire empirical knowledge of human interpreting. In this paper, we report the characteristic features of spoken language by Japanese-to-English interpreters.
Yasuyuki Aizawa, Shigeki Matsubara, Nobuo Kawaguchi, Katsuhiko Toyama, Yasuyoshi Inagaki
INTERSPEECH3
2000 Construction of speech corpus in moving car environment
abstract
The Center for Integrated Acoustic Information Research (CIAIR) at Nagoya University has been collecting speech corpora in moving cars which are made available as resources to advance the research and development of robust ASRs and spoken dialogue systems under high-noise conditions. The speech corpus consists of (1) phonetically balanced sentences, (2) digit strings, (3) discrete words and (4) transcribed spoken dialogues between drivers and information systems for navigation and information retrieval. These data are collected in vehicles under both idling and driving situations. The language of the corpus is currently Japanese. The number of subjects is currently about 300, total recording time is over 200 hours and total corpus size is about 160GByte. We have also been recording video images from three different angles, vehicle-control signals, and vehicle location, all synchronized with the speech recording. We report the objective of the speech corpus, the recording methods and the recording vehicle developed.
Nobuo Kawaguchi, Shigeki Matsubara, Hiroyuki Iwa, Shoji Kajita, Kazuya Takeda, Fumitada Itakura, Yasuyoshi Inagaki
INTERSPEECH1
2000 MAGNET: ad hoc network system based on mobile agents
Nobuo Kawaguchi, Katsuhiko Toyama, Yasuyoshi Inagaki
Comput. Commun.1