Tatsuya Amano

dblp:190/1106 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 7 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Real-Time Predictive Crowd Twin Platform by Efficient Data Assimilation
Atsuhiro Yano, Tatsuya Amano, Hirozumi Yamaguchi
COMPSAC2
2026 Human-Flow Digital Twin for Predicting the Effects of Mobility Introduction on Visitor Circulation
Chiharu Shima, Haruki Yonekura, Fukuharu Tanaka, Tatsuya Amano, Hirozumi Yamaguchi
MDM4
2026 A Simulation-based Framework for Dynamic Light Pollution Prediction in Urban Air Mobility
Ying Chieh Wang, Tatsuya Amano, Hirozumi Yamaguchi
SmartComp2
2025 MobText-SISA: Efficient Machine Unlearning for Mobility Logs with Spatio-Temporal and Natural-Language Data
abstract
Modern mobility platforms have stored vast streams of GPS trajectories, temporal metadata, free-form textual notes, and other unstructured data. Privacy statutes such as the GDPR require that any individual's contribution be unlearned on demand, yet retraining deep models from scratch for every request is untenable. We introduce MobText-SISA, a scalable machine-unlearning framework that extends Sharded, Isolated, Sliced, and Aggregated (SISA) training to heterogeneous spatio-temporal data. MobText-SISA first embeds each trip's numerical and linguistic features into a shared latent space, then employs similarity-aware clustering to distribute samples across shards so that future deletions touch only a single constituent model while preserving inter-shard diversity. Each shard is trained incrementally; at inference time, constituent predictions are aggregated to yield the output. Deletion requests trigger retraining solely of the affected shard from its last valid checkpoint, guaranteeing exact unlearning. Experiments on a ten-month real-world mobility log demonstrate that MobText-SISA (i) sustains baseline predictive accuracy, and (ii) consistently outperforms random sharding in both error and convergence speed. These results establish MobText-SISA as a practical foundation for privacy-compliant analytics on multimodal mobility data at urban scale.
Haruki Yonekura, Ren Ozeki, Tatsuya Amano, Hamada Rizk, Hirozumi Yamaguchi
SIGSPATIAL/GIS3
2025 Collaborative Lightweight LLM Agents for Daily Activity Summarization on Edge Devices
abstract
This paper presents a privacy-preserving monitoring system for elderly individuals living alone, utilizing collaborative lightweight Large Language Model (LLM) agents deployed on edge devices. The system integrates non-invasive sensors with a novel three-agent architecture: an activity recognition agent processes sensor data, an hourly summarization agent generates intermediate reports in 3-hour segments, and a daily summarization agent produces comprehensive summaries. Our evaluation on real-world data from two elderly households demonstrates that the system achieves 95.8% accuracy in activity recognition and generates natural language summaries comparable to GPT-4o, while maintaining privacy through local processing on affordable Raspberry Pi hardware. The results indicate that our approach effectively balances monitoring accuracy, summary quality, and practical deployment constraints, making it suitable for widespread adoption in elderly care applications.
Kentaro Inohara, Tatsuya Amano, Hamada Rizk, Hirozumi Yamaguchi
IE2
2025 LLM - Driven Adaptive Autonomous Robot Navigation via Multimodal Fusion for Diverse Environments
abstract
This paper presents a novel autonomous navigation framework that integrates Large Language Models (LLMs) with multimodal sensor fusion to enable dynamic obstacle avoidance and human-aware path planning in diverse environments. The proposed system leverages an FPGA-accelerated fusion pipeline, combining LiDAR and vision data for real-time perception. A Hungarian algorithm-based object matching technique ensures robust tracking, while a bird's-eye view (BEV) representation enhances spatial reasoning and occlusion handling. The fused sensory inputs are processed by a fine-tuned LLM, which contextualizes pedestrian behavior and environmental constraints to generate adaptive, human-centric navigation strategies. Unlike traditional rule-based methods, LLMs provide generalization capabilities to novel scenarios, significantly improving interaction with vulnerable pedestrians such as children, elderly individuals, and wheelchair users. Extensive evaluations in both simulated and real-world scenarios confirm the system's ability to reduce collisions and enhance navigation efficiency in high-density environments. By bridging semantic reasoning and robotic control, this work lays the foundation for next-generation intelligent navigation systems that are both safety-aware and scalable across autonomous platforms.
Ahmed Farid, Riki Ukyo, Tatsuya Amano, Hamada Rizk, Hirozumi Yamaguchi
IV4
2025 A Digital Twin Approach for Crowd Flow Modeling on Railway Station Platforms
abstract
Effective crowd tracking at railway station platforms is essential for ensuring passenger safety and optimizing pedestrian flow, particularly in high-density urban transit hubs. However, traditional tracking methods, such as object detection and multi-object tracking, face limitations in congested environments due to severe occlusions and overlapping individuals. This paper proposes a novel approach for modeling pedestrian flow on train station platforms by coupling deep learning-based motion analysis with crowd simulation. In the proposed method, we utilize RAFT, a state-of-the-art optical flow model, to extract pixel-level motion vectors, which are clustered to identify human movement patterns. These motion data are mapped onto a calibrated 2D platform model, providing a top-down representation of pedestrian trajectories. To simulate realistic crowd dynamics, Unity's NavMesh is employed alongside an enhanced Simulated Annealing approach to generate high-accuracy origin-destination (OD) data. This is a new digital twin concept where the analysis from the vision in the real world is projected onto the virtual world model to simulate and reproduce the pedestrian flows. The proposed method was evaluated using synthetic crowd simulation data, demonstrating high accuracy in destination estimation. The experimental results indicate that the OD estimation outperforms conventional approaches, with error rates reduced to half of those observed in YOLOv8x-based tracking systems. These findings suggest that the integration of optical flow-based motion analysis with digital twin simulation can significantly enhance crowd monitoring and congestion management in railway stations.
Yu Yasuda, Tatsuya Amano, Hirozumi Yamaguchi
SMARTCOMP2
2025 LLM-Powered Embodied Intelligence for Socially-Aware Robot Navigation in Human-Robot Interaction
abstract
This doctoral research proposes a framework for developing sociallyaware robot navigation systems by integrating the cognitive capabilities of Large Language Models (LLMs) with the demands of real-world Human-Robot Interaction (HRI).Our work follows a four-stage plan that systematically addresses the challenges of applying LLMs to time-sensitive, safety-critical tasks.This paper details the completion of the first two stages, wherein we developed and evaluated a foundational navigation model.Our system features a meticulously designed multimodal fusion pipeline that integrates LiDAR and camera data, processed by a YOLO model and a Hungarian algorithm for semantic association, providing rich, contextual input to the LLM.Through knowledge distillation and fine-tuning on data from a custom simulator, our model demonstrates robust spatial reasoning and superior performance in low-frequency decision-making scenarios compared to traditional reinforcement learning methods.We successfully validated this foundational model and identified its inference latency as a key challenge.These results establish a solid basis for our future work.This includes developing a "brain-cerebellum" hybrid architecture for real-time performance and exploring multi-robot social compliance.This research contributes to HRI by creating more predictable and trustworthy robots, and to the LLM field by investigating the symbol grounding problem through embodied intelligence.
Ahmed Farid, Tatsuya Amano, Hamada Rizk, Hirozumi Yamaguchi
SSTD3
2025 Lightweight Safety Assistance System for E-Scooters: Current Results and Future Directions
abstract
Electric scooters (e-scooters) are reshaping urban mobility, providing eco-friendly and cost-effective transport.However, the surge in their usage in mixed-traffic environments poses significant safety risks, especially for vulnerable road users (VRUs).This paper presents a lightweight vision-based safety assistance system optimized for real-time inference on edge devices.The core modules include semantic segmentation enhanced by semi-supervised learning, accurate bird's-eye view (BEV) transformation, real-time motion prediction, and adaptive path planning.Extensive evaluations demonstrate high segmentation accuracy and low-latency execution suitable for resource-constrained hardware.Future research directions include knowledge distillation for model adaptability, cooperative multi-scooter perception, intersection-based AI infrastructure, and the application of large language models (LLMs) to optimize urban traffic flows.
Congzhi Ren, Hamada Rizk, Tatsuya Amano, Hirozumi Yamaguchi
SSTD3
2025 Real-time Path Prediction at the Edge for E-scooters
abstract
Electric scooters (e-scooters) are increasingly used in urban areas but face serious safety concerns, especially in mixed traffic environments. This paper presents a real-time vision-based safety assistance system designed for e-scooters and optimized for edge deployment. The system uses a lightweight segmentation model trained via semi-supervised learning to detect sidewalks and road areas from non-vehicle perspectives. Dynamic objects such as pedestrians and cyclists are tracked using Kalman filtering, and potential risks are assessed in real time. A hybrid path planning module uses A* and artificial potential fields to recommend safe forward paths. Experimental results show that the system achieves real-time performance on Jetson Orin NX while improving perception accuracy and navigation safety in sidewalk environments.
Congzhi Ren, Hamada Rizk, Tatsuya Amano, Hirozumi Yamaguchi
VTC2025-Fall3
2025 Robust Pedestrian Tracking With Severe Occlusions in Public Spaces Using 3D Point Clouds
abstract
The increasing need for pedestrian tracking within public and commercial domains necessitates innovative solutions that respect privacy concerns. Traditional surveillance methods employing RGB video footage for the monitoring of pedestrian movements have raised significant privacy issues, attributed to their capability to easily identify individuals. This paper shifts focus towards the utilization of 3D point cloud data as a less intrusive alternative for pedestrian monitoring, which inherently minimizes the risk of individual identification. Notwithstanding, traditional tracking methodologies frequently encounter difficulties in environments characterized by high pedestrian density, where occlusions caused by individuals obscuring one another from the sensors’ perspective are commonplace. This research introduces an advanced approach leveraging 3D point cloud data to address the challenges inherent in pedestrian tracking. A particular obstacle addressed is the frequent obstruction caused by pedestrians, who may be accompanied by objects such as baby strollers and/or luggage, thus complicating their detection and tracking when positioned behind others or in close proximity to the sensors. Such scenarios significantly undermine the efficacy of pedestrian segmentation and tracking, particularly in the context of Kalman-Filter-based methodologies. To mitigate these challenges, the proposed solution incorporates a method to spatially enhance the missing segments within the framework of Kalman-Filter-based multi-object tracking. The effectiveness of this approach has been rigorously assessed through the application of 3D point cloud data derived from diverse real-world settings, including laboratory environments, shopping mall entrances, and a densely populated urban area in Kita-ku, Osaka City, during a festival event.
Riki Ukyo, Tatsuya Amano, Hamada Rizk, Hirozumi Yamaguchi, Takao Moriya
IEEE Trans. Intell. Transp. Syst.2
2024 Simulating Urban Pedestrian Flows by Fusing Wide-Area Location Data and Spot Pedestrian Counts
Uegaki Masashi, Tatsuya Amano, Hirozumi Yamaguchi
MobiQuitous2
2024 Scalable and Distributed Optimization of Shared 3D Object Quality for Large-Scale Hybrid-Metaverses
abstract
Hybrid-metaverses, integrating physical and virtual spaces, face a critical challenge in managing shared 3D object quality across multiple users with diverse preferences and limited network resources. This paper addresses the problem of allocating limited bandwidth for transmitting point cloud representations while maximizing overall user satisfaction. We propose a distributed optimization method that dynamically adjusts 3D object quality based on contextual importance, available resources, and user preferences. Our approach uses Input Convex Neural Networks (ICNN) to model user utility functions and employs the Alternating Direction Method of Multipliers (ADMM) for distributed optimization. Key advantages include scalability, adaptability, and improved quality of experience. Evaluation using open dataset demonstrates significant improvements in user satisfaction and resource utilization compared to baseline approaches. Our method achieves 93-94.6% accuracy in modeling user utility and shows up to 60% faster convergence for scenarios with 30 users, contributing to the balance between high-fidelity representation and efficient data management in hybrid-metaverses.
Yui Maruyama, Tatsuya Amano, Hirozumi Yamaguchi
SECON2
2024 Advancing Smart Computing: A Comprehensive Tutorial to 3D Point Clouds from Installation to Efficient Processing and Context Recognition for Next-Gen Applications
abstract
3D point clouds from LiDAR sensors have emerged as a powerful representation for understanding and interacting with the surrounding environment in smart computing systems. This tutorial aims to provide an overview of fundamental techniques, considerations, and applications of 3D point cloud recognition, with a focus on both deep learning and nondeep learning approaches. We will begin by introducing the unique characteristics and challenges of point cloud processing, highlighting the differences from traditional image-based approaches. Through the example of PointNet [1], a pioneering deep learning model for point cloud recognition, we will discuss the specific considerations and best practices for handling point cloud data. Recognizing the limitations of deep learning in resource-constrained and mobile environments, we will also explore alternative statistical and probabilistic techniques, such as Fisher Vector-based approaches, which enable efficient and lightweight point cloud processing and context recognition. These techniques are particularly relevant for smart computing applications that require real-time processing on edge devices and in Internet of Things scenarios.
Tatsuya Amano, Hamada Rizk
SMARTCOMP1
2024 Privacy-preserving pedestrian tracking with path image inpainting and 3D point cloud features
abstract
Tracking pedestrian flow in large public areas is vital, yet ensuring privacy is paramount. Traditional visual-based tracking systems are raising concerns for potentially obtaining persistent and permanent identifiers that can compromise individual identities. Moreover, in areas such as the vicinity of restrooms, any form of data acquisition capturing human behavior should be refrained from, making it also crucial to appropriately address and complement these blind spots for a comprehensive analysis of pedestrian movement in the entire area. In this paper, we present our pedestrian tracking algorithm using distributed 3D LiDARs (Light Detection and Ranging), which capture pedestrians as 3D point clouds, omitting identifiable features. Our system bridges blind spots by leveraging historical movement data and 3D point cloud features, complemented by a generative diffusion model to predict trajectories in unseen areas. In a large-scale testbed with 70 LiDARs, the system achieved a 0.98 F-measure, highlighting its potential as a leading privacy-preserving tracking solution.
Masakazu Ohno, Riki Ukyo, Tatsuya Amano, Hamada Rizk, Hirozumi Yamaguchi
Pervasive Mob. Comput.3
2023 Multi-Person Tracking Method Robust to Dynamic Viewport Changes for AR apps
abstract
Augmented reality (AR) devices have gained a lot of attention in recent years due to their ability to enhance people’s abilities through 2D/3D spatial sensing and recognition functions. RGB cameras are most often used as sensors in this type of spatial recognition, and a particularly important task is the detection and tracking of objects and people in physical space. However, the camera positions and orientations on AR devices such as smartphones and smart glasses, frequently change due to the user wearing them on their head, leading to non-linear and complex motion in the video frames and reducing the accuracy of tracking people. To address this issue, the proposed method combines person re-identification based on deep metric learning with trajectory prediction to estimate the person’s sequential positions in 3D space around the camera. The experimental result shows 95.45% accuracy with our dataset.
Naoya Takahashi, Tatsuya Amano, Hirozumi Yamaguchi
IE2
2023 Fall Detection and Assessment Using Multitask Learning and Micro-sized LiDAR in Elderly Care
Hamada Rizk, Tatsuya Amano, Hirozumi Yamaguchi
MobiQuitous (2)3
2023 Privacy-preserving Pedestrian Tracking using Distributed 3D LiDARs
abstract
The growing demand for intelligent environments unleashes an extraordinary cycle of privacy-aware applications that makes individuals' life more comfortable and safe. Examples of these applications include pedestrian tracking systems in large areas. Although the ubiquity of camera-based systems, they are not a preferable solution due to the vulnerability of leaking the privacy of pedestrians. In this paper, we introduce a novel privacy-preserving system for pedestrian tracking in smart environments using multiple distributed LiDARs of non-overlapping views. The system is designed to leverage LiDAR devices to track pedestrians in partially covered areas due to practical constraints, e.g., occlusion or cost. Therefore, the system uses the point cloud captured by different LiDARs to extract discriminative features that are used to train a metric learning model for pedestrian matching purposes. To boost the system's robustness, we leverage a probabilistic approach to model and adapt the dynamic mobility patterns of individuals and thus connect their sub-trajectories. We deployed the system in a large-scale testbed with 70 colorless LiDARs and conducted three different experiments. The evaluation result at the entrance hall confirms the system's ability to accurately track the pedestrians with a 0.98 F-measure even with zero-covered areas. This result highlights the promise of the proposed system as the next generation of privacy-preserving tracking means in smart environments.
Masakazu Ohno, Riki Ukyo, Tatsuya Amano, Hamada Rizk, Hirozumi Yamaguchi
PERCOM3
2022 Demonstrating OmniCells: a resilient indoor localization system to devices' diversity
abstract
In this paper, we demonstrate OmniCells: a cellular-based indoor localization system designed to combat the device heterogeneity problem. OmniCells is a deep learning-based system that leverages cellular measurements from one or more training devices to provide consistent performance across unseen tracking phones. In this demo, we show the effect of device heterogeneity on the received cellular signals and how this leads to performance deterioration of traditional localization systems. In particular, we show how OmniCells and its novel feature extraction methods enable learning a rich and device-invariant representation without making any assumptions about the source or target devices. The system also includes other modules to increase the deep model's generalization and resilience to unseen scenarios.
Hamada Rizk, Tatsuya Amano, Hirozumi Yamaguchi, Moustafa Youssef 0001
MobiCom2
2022 Semantic Communication for Capacity-aware Remote Collaboration
abstract
The global spread of coronavirus has sparked a considerable interest in technologies that facilitate seamless communication between users which are physically or spatially distant. Using current remote collaboration systems that utilize 3D sensing with LiDAR and depth cameras, point cloud streaming, and MR/VR devices, distant users can communicate with each other as if they did in person. However, these systems may violate users' privacy since they can share information of their entire personal space with other users. In addition, although various point cloud compression methods have been proposed, remote transmission of 3D scenes still requires significant bandwidth. This paper proposes a 3D spatial data sharing system based on the paradigm of “semantic communication”, i.e., controlling communication in the units of semantic objects. Our system understands the semantics of the scene and leverages point cloud streaming, thereby enabling users to assert fine-grained control over their privacy. Further, the system adaptively controls the size of the data frame based on network capacity and scene context. The experimental results show that the network delay can be reduced by 96%. We have also tested our system in a commercial 4G network, showing that 3-D spatial sharing with point clouds over severe networks is possible.
Tatsuya Amano, Srikant Manas Kala, Teruhiro Mizumoto, Hirozumi Yamaguchi
WiMob1
2022 Object Recognition from 3D Point Cloud on Resource-Constrained Edge Device
abstract
This paper presents the design and development of a lightweight, portable 3D spatial sensing device equipped with a compact LiDAR-type sensor. The device provides a 3D point cloud representation of the surrounding environment as captured by the sensor. Based on the acquired 3D point cloud, we propose a real-time object recognition method. The technical challenge is how to process the 3D point data in real-time while pursuing the best trade-off between the processing overhead and recognition accuracy. To answer the question, we leverage the Fisher Vector to extract spatio-temporal features of different objects enabling efficient classification of these objects using the Support Vector Machine approach. The experimental results show that the proposed method achieves mean Average Precision of 0.961. The processing speed on the device was 59.3 frames/second, indicating that object detection can be done in real-time on the device.
Yuma Okochi, Hamada Rizk, Tatsuya Amano, Hirozumi Yamaguchi
WiMob3
2021 Smartwatch-Based Face-Touch Prediction Using Deep Representational Learning
Hamada Rizk, Tatsuya Amano, Hirozumi Yamaguchi, Moustafa Youssef 0001
MobiQuitous2
2021 Road Segment Re-Identification in Dashcam Videos
abstract
Due to the widespread of dashcams, we will have more videos that capture roads/streets in driving. Consequently, a vast amount of the videos will be available and can be utilized for analyzing road safety and similar purposes. For example, suppose different dashcams can take vehicle/pedestrian traffic at a risky intersection at different timings. In that case, the collection of such videos will effectively recognize the cause of dangerous situations without surveillance camera infrastructure. However, identifying a Road Segment of Interest (RSI) in the video, such as near the intersection region, is challenging as the video frames do not usually include location tags. In this paper, we present a unique approach to attack this challenge. Assume that a video segment, called reference video, captures an RSI. We re-identify the RSI taken in another video (called test video) that captures the roads containing that RSI. By this approach, we can automatically extract the video segment corresponding to RSI from a given test video, using the reference video. We introduce AKAZE features to assess frame-level similarity and develop an algorithm to find frame-by-frame matching between reference and test videos. We have evaluated our method using 10 reference videos that correspond to 10 RSIs, each with 5 test videos. The result has shown that the average frame error distance was only 3.03 in daytime and 4.93 in nighttime, which are sufficiently low to re-identify RSI in the newly obtained test videos.
Yukihiro Tsukamoto, Tatsuya Amano, Akihito Hiromori, Hirozumi Yamaguchi, Teruo Higashino
WiMob2
2018 Smartphone Applications Testbed Using Virtual Reality
abstract
Due to the nature of smartphones' portability and mobility, many mobile apps are usually utilized in the real field environment using GPS, Wi-Fi and embedded sensors. For example, any navigation app uses GPS and Wi-Fi to locate the user in the map, and streaming apps may be used in cafeteria or even outside to satisfy the users' demand to watch soccer games anywhere and anytime. To test the usability and performance of such mobile apps in in-situ environment, we need to bring the apps to such physical world and run (a number of) test scenarios, which is often cost-inefficient depending on the size, apps and situations assumed in those scenarios. In this paper, we design and develop a testbed to test mobile apps in VR space. The system allows developers to use a real smartphone in VR and to test and evaluate their apps at the interested locations, with various network environment. The system builds and reproduces the real world environment of 3D space and real networks in the VR environment, using the existing 3D city models and our original Wi-Fi database. Then it enables to real-timely integrate the screen of the VR user's smartphone in the VR space. The user can operate the app via the VR view, and test the usability and performance of the app in such an emulated environment. The experimental result shows our architecture could achieve such cyber-physical integration with 695.5 ms delay, which is negligible in many semi-realtime services such as navigations.
Tatsuya Amano, Shugo Kajita, Hirozumi Yamaguchi, Teruo Higashino, Mineo Takai
MobiQuitous1
2016 Wi-Fi Channel Selection Based on Urban Interference Measurement
abstract
Increasing availability and usability enhancement of Wi-Fi in public areas has become more active. However, due to the dense deployment of Wi-Fi access points (APs), there is a chaotic and disorderly environment in urban areas. In our previous work, we have designed a function that predicts the network performance at each Wi-Fi AP according to the measurement of IEEE802.11 MAC frames sensed in each Wi-Fi channel. However, it was not examined in such scenarios assuming urban environment. We should understand the situations of current Wi-Fi AP deployment and traffic conditions, and should confirm the effectiveness of channel migration in such realistic environment. In this study, we proposed urban Wi-Fi channel utilization model based on real urban Wi-Fi measurement. We show that our method can predict the best channels and APs can migrate to them in the urban scenario.
Shugo Kajita, Tatsuya Amano, Hirozumi Yamaguchi, Teruo Higashino, Mineo Takai
MobiQuitous2