Xiaodan Shi

dblp:217/4074 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 2 since 2021Computer networks · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Learn to Cluster Human Mobility Pattern for Post-Disaster Analysis
abstract
Disaster has a great impact on human mobility patterns. Clustering mobility patterns by generalizing group level characteristics from diverse individual trajectories plays a vital role in informing post-disaster recovery strategies. However, predefined clustering criteria or supervised labeling alone are insufficient to adequately capture the dynamics of mobility patterns during disaster events. Besides, the existing methods are not enough to capture the sensitive changes in mobility patterns in disaster events and handle the data of high-dimensional. This research investigates an embedded deep learning-based method which can automatically extract the groups' short-term feature of mobility patterns and achieves short-term mobility pattern clustering during disaster. The proposed method employs a Transformer-based temporal encoder to capture intra-day sequence patterns and integrates a VAE component with an embedded latent variable that directly encodes group-level mobility modes. We also design a compactness–separation loss that explicitly encourages within-mode feature compactness and between-mode feature separation. Based on massive mobile data, we conduct mobility pattern clustering on the case of 2011 Fukushima Earthquake. Compared to conventional clustering approaches, the proposed model structure is more discriminative and can capture the more sensitive changes between pre- and post- events. Compared to baselines with different loss functions, proposal methods can make more accurate fitting result and obtain more discrete clusters modes. Additionally, sensitivity analysis is conducted to examine the influence of key hyper parameters within the model. Based on the clustering outcomes, five representative mobility patterns are identified. We further analyze the spatial-temporal characteristics of mobility pattern changes during the disaster events and the recovery period of mobility pattern.
Wenjing Li 0006, Yuhao Yao, Hill Hiroki Kobayashi, Haoran Zhang 0002, Xuan Song 0001, Ryosuke Shibasaki, Xiaodan Shi
IEEE Trans. Big Data9
2024 Learning Social and Physical Compliant Multi-modal Futures
Xiaodan Shi, Haoran Zhang 0002, Ryosuke Shibasaki, Jinyue Yan
ICPR (24)1
2024 A Phone-Based Distributed Ambient Temperature Measurement System With an Efficient Label-Free Automated Training Strategy
abstract
Enhancing the energy efficiency of buildings significantly relies on monitoring indoor ambient temperature. The potential limitations of conventional temperature measurement techniques, together with the omnipresence of smartphones, have redirected researchers' attention towards the exploration of phone-based ambient temperature estimation methods. However, existing phone-based methods face challenges such as insufficient privacy protection, difficulty in adapting models to various phones, and hurdles in obtaining enough labeled training data. In this study, we propose a distributed phone-based ambient temperature estimation system which enables collaboration among multiple phones to accurately measure the ambient temperature in different areas of an indoor space. This system also provides an efficient, cost-effective approach with a few-shot meta-learning module and an automated label generation module. It shows that with just 5 new training data points, the temperature estimation model can adapt to a new phone and reach a good performance. Moreover, the system uses crowdsourcing to generate accurate labels for all newly collected training data, significantly reducing costs. Additionally, we highlight the potential of incorporating federated learning into our system to enhance privacy protection. We believe this study can advance the practical application of phone-based ambient temperature measurement, facilitating energy-saving efforts in buildings.
Dayin Chen, Xiaodan Shi, Haoran Zhang 0002, Xuan Song 0001, Dongxiao Zhang, Yuntian Chen, Jinyue Yan
IEEE Trans. Mob. Comput.2
2024 MobCovid: Confirmed Cases Dynamics Driven Time Series Prediction of Crowd in Urban Hotspot
abstract
Monitoring the crowd in urban hot spot has been an important research topic in the field of urban management and has high social impact. It can allow more flexible allocation of public resources such as public transportation schedule adjustment and arrangement of police force. After 2020, because of the epidemic of COVID-19 virus, the public mobility pattern is deeply affected by the situation of epidemic as the physical close contact is the dominant way of infection. In this study, we propose a confirmed case-driven time-series prediction of crowd in urban hot spot named MobCovid. The model is a deviation of Informer, a popular time-serial prediction model proposed in 2021. The model takes both the number of nighttime staying people in downtown and confirmed cases of COVID-19 as input and predicts both the targets. In the current period of COVID, many areas and countries have relaxed the lockdown measures on public mobility. The outdoor travel of public is based on individual decision. Report of large amount of confirmed cases would restrict the public visitation of crowded downtown. But, still, government would publish some policies to try to intervene in the public mobility and control the spread of virus. For example, in Japan, there are no compulsory measures to force people to stay at home, but measures to persuade people to stay away from downtown area. Therefore, we also merge the encoding of policies on measures of mobility restriction made by government in the model to improve the precision. We use historical data of nighttime staying people in crowded downtown and confirmed cases of Tokyo and Osaka area as study case. Multiple times of comparison with other baselines including the original Informer model prove the effectiveness of our proposed method. We believe our work can make contribution to the current knowledge on forecasting the number of crowd in urban downtown during the Covid epidemic.
Xiaodan Shi, Haoran Zhang 0002, Wenjing Li 0006, Yuhao Yao, Satoshi Miyazawa, Xuan Song 0001, Ryosuke Shibasaki
IEEE Trans. Neural Networks Learn. Syst.2
2023 Hybrid Feature Embedding for Automatic Building Outline Extraction
abstract
Building outline extracted from high-resolution aerial images can be used in various application fields such as change detection and disaster assessment. However, traditional CNN model cannot recognize contours very precisely from original images. In this paper, we proposed a CNN and Transformer based model together with active contour model to deal with this problem. We also designed a triple-branch decoder structure to handle different features generated by encoder. Experiment results show that our model outperforms other baseline model on two datasets, achieving 91.1% mIoU on Vaihingen and 83.8% on Bing huts.
Weihang Ran, Wei Yuan 0004, Xiaodan Shi, Zipei Fan, Ryosuke Shibasaki
IGARSS3
2023 Graph Encoding based Hybrid Vision Transformer for Automatic Road Network Extraction
abstract
This paper introduces a graph encoding-based hybrid vision transformer for automatic road network extraction via high-resolution remote sensing imagery. Given that high-resolution remote sensing images covered large urban areas, traditional segmentation-based road extraction methods usually can generate good binary classification maps in simple structured road surfaces but fail in complex highway and bridge-covered areas. We introduce a graph encoding-based mechanism to address the above issues, enabling the road extraction framework extracts the road segmentation feature and build the graph structure map jointly. Compared to only segmentation-based methods, our approach learns prior geometrical structure information from the extracted ViT feature maps and has a non-local awareness of the whole road network structure. Eexperimental results demonstrated that the proposed approach outperforms the traditional segmentation-based methods.
Wei Yuan 0004, Weihang Ran, Xiaodan Shi, Zipei Fan, Ryosuke Shibasaki
IGARSS3
2023 PredLife: Predicting Fine-Grained Future Activity Patterns
abstract
Activity pattern prediction is a critical part of urban computing, urban planning, intelligent transportation, and so on. Based on a dataset with more than 10 million GPS trajectory records collected by mobile sensors, this research proposed a CNN-BiLSTM-VAE-ATT-based encoder-decoder model for fine-grained individual activity sequence prediction. The model combines the long-term and short-term dependencies crosswise and also considers randomness, diversity, and uncertainty of individual activity patterns. The proposed results show higher accuracy compared to the ten baselines. The model can generate high diversity results while approximating the original activity patterns distribution. Moreover, the model also has interpretability in revealing the time dependency importance of the activity pattern prediction.
Wenjing Li 0006, Xiaodan Shi, Dou Huang, Hill Hiroki Kobayashi, Haoran Zhang 0002, Xuan Song 0001, Ryosuke Shibasaki
IEEE Trans. Big Data2
2023 Metagraph-Based Life Pattern Clustering With Big Human Mobility Data
abstract
Life pattern clustering is essential for abstracting the groups' characteristics of daily life patterns and activity regularity. Based on millions of GPS records, this research proposes a framework on the life pattern clustering which can efficiently identify the groups that have similar life patterns. The proposed method can retain original features of individual life pattern data without aggregation. Metagraph-based data structure is proposed for presenting the diverse life pattern. Spatial-temporal similarity includes significant places semantics, time-sequential properties and frequency are integrated into this data structure, which captures the uncertainty of an individual and the diversities between individuals. Non-negative-factorization-based method is utilized for reducing the dimension. The results show that our proposed method can effectively identify the groups that have similar life pattern in long term and takes advantage in computation efficiency and representational capacity compared with the traditional methods. We reveal the representative life pattern groups and analyze the group characteristics of human life patterns during different periods and different regions. We believe our work helps in future infrastructure planning, services improvement and policy making related to urban and transportation, thus promoting a humanized and sustainable city.
Wenjing Li 0006, Haoran Zhang 0002, Yuhao Yao, Xiaodan Shi, Mariko Shibasaki, Hill Hiroki Kobayashi, Xuan Song 0001, Ryosuke Shibasaki
IEEE Trans. Big Data6
2023 MetaTraj: Meta-Learning for Cross-Scene Cross-Object Trajectory Prediction
abstract
Long-term pedestrian trajectory prediction in crowds is highly valuable for safety driving and social robot navigation. The recent research of trajectory prediction usually focuses on solving the problems of modeling social interactions, physical constraints and multi-modality of futures without considering the generalization of prediction models to other scenes and objects, which is critical for real-world applications. In this paper, we propose a general framework that makes trajectory prediction models able to transfer well across unseen scenes and objects by quickly learning the prior information of trajectories. The trajectory sequences are closely related to the circumstance setting (e.g. exits, roads, buildings, entries etc.) and the objects (e.g. pedestrians, bicycles, vehicles etc.). We argue that those trajectory information varying across scenes and objects makes a trained prediction model not perform well over unseen target data. To address it, we introduce MetaTraj that contains carefully designed sub-tasks and meta-tasks to learn prior information of trajectories related to scenes and objects, which then contributes to accurate long-term future prediction. Both sub-tasks and meta-tasks are generated from trajectory sequences effortlessly and can be easily integrated into many prediction models. Extensive experiments over several trajectory prediction benchmarks demonstrate that MetaTraj can be applied to multiple prediction models and enables them generalize well to unseen scenes and objects.
Xiaodan Shi, Haoran Zhang 0002, Wei Yuan 0004, Ryosuke Shibasaki
IEEE Trans. Intell. Transp. Syst.1
2023 LTP-Net: Life-Travel Pattern Based Human Mobility Signature Identification
abstract
How to effectively extract identifiable information from human mobility data and distinguish different agents is a significant topic for location-based services and intelligent transportation systems, which is described as the Human Mobility Signature Identification problem. A deeper understanding of the identifiable information underlain in human mobility can help us lay the foundation for applications such as irregular user behavior detection and privacy protection. However, human mobility comprises a mixture of different mobility patterns, traditional methods usually pay more attention to spatial-temporal features, while pattern dimension feature is usually ignored, which makes the result very dependent on the population agglomeration degree. To bridge the research gap, in this paper, we propose a novel Life-Travel pattern-based learning module (LTP-Net), in which spatial-temporal-pattern dimension features are embedded together to provide more comprehensive information for individual identification. A real-world mobile phone location dataset is utilized to evaluate the performance of the proposed LTP-Net and traditional methods. Several case studies are also conducted to analyze the model performance, including the abnormal behavior detection for the east Japan earthquake.
Yuhao Yao, Haoran Zhang 0002, Xiaodan Shi, Wenjing Li 0006, Xuan Song 0001, Ryosuke Shibasaki
IEEE Trans. Intell. Transp. Syst.3
2022 Cross-Scale Attention-based Tree Crown Detection via UAV imagery
abstract
This paper introduces a cross-scale attention based end-to-end learning framework for tree crown detection via UAV imagery. Given that UAV images covered a large forests, the illumination variations, shadow obstacles and texture repetition always lead to inaccurate tree crown detection results. We introduce a cross-scale attention based mechanism to address the above issues, enabling the tree crown detection framework to reason about the RGB texture information and depth information introduced by the automatically generated depth map jointly. Compared to traditional image based tree crown detection methods, our approach learns prior over geometrical structure information from the real 3D world, which is robust to the texture repetition and small tree crowns. The experimental results demonstrated that the proposed approach outperforms the traditional CNN based method.
Wei Yuan 0004, Xiaodan Shi, Zhiling Guo, Zipei Fan, Jianya Gong, Ryosuke Shibasaki
IGARSS2
2021 Social-DPF: Socially Acceptable Distribution Prediction of Futures
abstract
We consider long-term path forecasting problems in crowds, where future sequence trajectories are generated given a short observation. Recent methods for this problem have focused on modeling social interactions and predicting multi-modal futures. However, it is not easy for machines to successfully consider social interactions, such as avoiding collisions while considering the uncertainty of futures under a highly interactive and dynamic scenario. In this paper, we propose a model that incorporates multiple interacting motion sequences jointly and predicts multi-modal socially acceptable distributions of futures. Specifically, we introduce a new aggregation mechanism for social interactions, which selectively models long-term inter-related dynamics between movements in a shared environment through a message passing mechanism. Moreover, we propose a loss function that not only accesses how accurate the estimated distributions of the futures are but also considers collision avoidance. We further utilize mixture density functions to describe the trajectories and learn the multi-modality of future paths. Extensive experiments over several trajectory prediction benchmarks demonstrate that our method is able to forecast socially acceptable distributions in complex scenarios.
Xiaodan Shi, Xiaowei Shao, Guangming Wu, Haoran Zhang 0002, Zhiling Guo, Renhe Jiang, Ryosuke Shibasaki
AAAI1
2020 Multimodal Interaction-Aware Trajectory Prediction in Crowded Space
abstract
Accurate human path forecasting in complex and crowded scenarios is critical for collision avoidance of autonomous driving and social robots navigation. It still remains as a challenging problem because of dynamic human interaction and intrinsic multimodality of human motion. Given the observation, there is a rich set of plausible ways for an agent to walk through the circumstance. To address those issues, we propose a spatio-temporal model that can aggregate the information from socially interacting agents and capture the multimodality of the motion patterns. We use mixture density functions to describe the human path and predict the distribution of future paths with explicit density. To integrate more factors to model interacting people, we further introduce a coordinate transformation to represent the relative motion between people. Extensive experiments over several trajectory prediction benchmarks demonstrate that our method is able to forecast various plausible futures in complex scenarios and achieves state-of-the-art performance.
Xiaodan Shi, Xiaowei Shao, Zipei Fan, Renhe Jiang, Haoran Zhang 0002, Zhiling Guo, Guangming Wu, Wei Yuan 0004, Ryosuke Shibasaki
AAAI1
2020 Learn to Recover Visible Color for Video Surveillance in a Day
Guangming Wu, Yinqiang Zheng, Zhiling Guo, Zekun Cai, Xiaodan Shi, Yifei Huang 0002, Ryosuke Shibasaki
ECCV (1)5
2020 An Adaptive Adjustment Algorithm of the Parameters in Alarm Association Rule Mining
abstract
With the rapid development of communication technology, communication networks are playing an increasingly important role in people's lives. Effective management of increasingly complex networks can improve the efficiency and stability of network operations. Fault management is one of the important functions of network management. The analysis of the alarms generated in the network can dig into the underlying rules to provide useful information for fault management. However, the existing alarm association rule mining algorithms often have the problem of parameter rigidity. In this paper, a two-level windows based alarm transaction extracting algorithm is proposed to solve the problem of low efficiency when using fixed size windows. Then this paper proposes an experience extraction method based on alarm priority in deep Q network (DQN), which can calculate the sampling probability according to the importance of alarm when the memory unit enters the queue. Aiming at the rare item problem caused by the fixed support threshold in association rule mining, the improved DQN is used to dynamically adjust the minimum support in rule mining algorithm. Experimental results show that the algorithm proposed in this paper can effectively improve the efficiency of alarm transaction extraction and the accuracy of alarm association rules mining.
Xiaodan Shi, Libin Jiao, Yang Yang 0006, Peng Yu 0001
IWCMC1
2019 A Fault Prediction Method Based on Load-capacity Model in the Communication Network
abstract
Due to the connectivity of the communication network, the occurrence of faults is often cascaded. The infrastructure of the communication network is extremely vulnerable to the failure of the hardware and software, resulting in the node to fail to work. Moreover, large-scale application service failure is often caused by the cascade effect after the failure occurs. Therefore, researching and predicting fault behavior becomes important for maintaining the reliability of the communication network. This paper analyzes the mechanism of fault cascading propagation in communication networks, and proposes a fault prediction model based on load capacity model, resetting the node load and capacity to exhibit a non-linear relationship and sets the node load to be dynamically changing, which is more suitable for the flow of the actual network. The original load capacity model has only two states of normal and fault, but the state can be between normal and fault in the real network. Therefore, this paper adds a third state-congestion state to the model to make the model more meet the actual characteristics of the communication network. In addition, this paper also proposes a load redistribution strategy based on Top K node intermediary, which is used to simulate network traffic redistribution. The experimental results show that the fault prediction model proposed in this paper has improved the precision and accuracy of the prediction, and has a good prediction effect.
Xilin Ji, Yonghua Huo, Xiaodan Shi, Yang Yang 0006
APNOMS5
2019 Geosr: A Computer Vision Package for Deep Learning Based Single-Frame Remote Sensing Imagery Super-Resolution
abstract
Recently, owing to the outstanding capability of deep learning in solving ill-posed problems, the single-frame super-resolution (SR) researches tend to focus on deep learning methods largely. However, related researches are implemented and evaluated through various datasets and different deep learning frameworks, which hinders the comparison of performance among different methods and heavily hampers the progress of SR techniques. In this study, we present GeoSR, an open source computer vision package for deep learning based single-frame remote sensing imagery super-resolution to facilitate the development of the SR community. As a unified, simple, and flexible package, GeoSR contains pipeline-like integrated tools from data retrieval to final result evaluation, which enables users to develop self-defined models conveniently; several state-of-the-art models trained through the same high-quality dataset are provided as the baseline in the package as well. Moreover, the proposed package could potentially serve as a viable backend for other related packages such as image segmentation with high efficiency.
Zhiling Guo, Guangming Wu, Xiaodan Shi, Mingzhou Sui, Xiaoya Song, Yongwei Xu, Xiaowei Shao, Ryosuke Shibasaki
IGARSS3
2019 Automatic Vectorization Extraction of Flat-Roofed Houses Using High-Resolution Remote Sensing Images
abstract
The vectorization of buildings provides quantitative disaster information and building damage extraction for the accurate assessment of disasters. This study presents an automatic extraction technology for flat-roofed houses by using high-resolution remote sensing images based on the single-point extraction of such houses and the concept of stepwise image segmentation. Vegetation removal, multi-scale segmentation, k-means clustering segmentation, rectangle recognition, and morphological processing are adopted to identify the center of an assumed house in the image by considering the characteristics of remote sensing images and the main features of house recognition (i.e., spectral characteristics and shape features). Furthermore, the single-point extraction of flat-roofed houses is used to acquire the surface vector diagram of the extracted houses, and finally, achieve the automatic extraction of flat-roofed houses. Experimental result shows that the accuracy of the vectorization extraction of flat-roofed houses through the combination is 83.33%.
Guorui Ma, Qinjie He, Xiaodan Shi, Xiaojie Fan
IGARSS3
2018 Semantic Segmentation for Urban Planning Maps Based on U-Net
abstract
The automatic digitizing of paper maps is a significant and challenging task for both academia and industry. As an important procedure of map digitizing, the semantic segmentation section is mainly relied on manual visual interpretation with low efficiency. In this study, we select urban planning maps as a representative sample and investigate the feasibility of utilizing U-shape fully convolutional based architecture to perform end-to-end map semantic segmentation. The experimental results obtained from the test area in Shibuya district, Tokyo, demonstrate that our proposed method could achieve a very high Jaccard similarity coefficient of 93.63% and an overall accuracy of 99.36%. For implementation on GPGPU and cuDNN, the required processing time for the whole Shibuya district can be less than three minutes. The results indicate the proposed method can serve as a viable tool for urban planning map semantic segmentation task with high accuracy and efficiency.
Zhiling Guo, Hiroaki Shengoku, Guangming Wu, Qi Chen 0012, Wei Yuan 0004, Xiaodan Shi, Xiaowei Shao, Yongwei Xu, Ryosuke Shibasaki
IGARSS6