VLDB 2026 Research / reviewers in the wild / expert
Wei Yuan 0004
dblp:67/2268-4
· DBLP profile ↗
14ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0002-1370-0079ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 9 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Taming Spatial Heterophily and Temporal Irregularity: A Curriculum Learning Approach for Traffic Forecasting
Hongjun Wang 0007, Zhiwen Zhang 0004, Jiyuan Chen, Zipei Fan, Renhe Jiang, Wei Yuan 0004, Ryosuke Shibasaki, Xuan Song 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2026 | Tiered Spatio-Temporal Difficulty: Curriculum Scheduler for Multi-Sensor Traffic Flow PredictionabstractThe development of the Internet of Things (IoT) has enhanced smart city services for traffic monitoring, leading to numerous schemes for accurate flow prediction based on traffic sensors. However, existing approaches primarily capture spatio-temporal (ST) dependencies from traffic graphs and train their models using randomly ordered data. This overlooks the fact that the modeling difficulty of each sensor/node in the ST traffic graph can vary significantly due to its spatial dependencies and temporal trends, resulting in unreliable and unstable predictions in IoT scenarios. In this context, we argue that a well-designed curriculum with an easy-to-difficult order can improve the training of ST models. Therefore, this paper introduces an ST difficulty measurer to score the node-level difficulty of traffic graph from both spatial and temporal aspects, and then implements a curriculum in the ST model training process. More specifically, based on the tiered ST difficulty score, the ST model training begins with a subgraph consisting of “easy” nodes characterized by relatively consistent spatial relationships and regular temporal patterns. Gradually, more difficult nodes are incorporated into the subgraph and participate in subsequent training stages. Comprehensive experiments and analysis on two real-world traffic flow datasets confirm the effectiveness of our proposed approach. Zhiwen Zhang 0004, Hongjun Wang 0007, Zipei Fan, Renhe Jiang, Wei Yuan 0004, Xuan Song 0001, Ryosuke Shibasaki |
IEEE Trans. Mob. Comput. | 5 |
| 2025 | PortVIS: An Interactive Platform for Port-to-Port Trajectory Imputation and Visual AnalyticsabstractMaritime traffic analysis plays a vital role in port operations and logistics coordination. Although Automatic Identification System (AIS) provides rich vessel movement data, analyzing port-to-port traffic remains challenging due to data heterogeneity and missing trajectory segments. We present PortVIS, a web-based interactive system that supports the comprehensive analysis and visualization of maritime traffic. PortVIS integrates multi-sourced datasets—including AIS records, port and anchorage metadata, and wind field data—to enable trajectory segmentation, regional analysis, and data imputation. Users can import raw AIS records and segment them into port-to-port trips using our system. These segmented trajectories can then be filtered and queried based on vessel information and trip attributes. Users can also define custom zones (e.g., anchorages or transit areas) and explore traffic patterns through maps and charts. Missing trajectory segments are reconstructed using our recent imputation approach to improve data quality. By integrating trajectory processing and various visual analytics, PortVIS provides a unified tool for maritime mobility analysis. Zhiwen Zhang 0004, Zipei Fan, Wei Yuan 0004, Shun Iwazaki, Ryosuke Shibasaki |
SIGSPATIAL/GIS | 3 |
| 2025 | You Always Recognize Me (YARM): Robust Texture Synthesis Against Multi-View CorruptionabstractDamage to imaging systems and complex external environments often introduce corruption, which can impair the performance of deep learning models pretrained on high-quality image data. Previous methods have focused on restoring degraded images or fine-tuning models to adapt to out-of-distribution data. However, these approaches struggle with complex, unknown corruptions and often reduce model accuracy on high-quality data. Inspired by the use of warning colors and camouflage in the real world, we propose designing a robust appearance that can enhance model recognition of low-quality image data. Furthermore, we demonstrate that certain universal features in radiance fields can be applied across objects of the same class with different geometries. We also examine the impact of different proxy models on the transferability of robust appearances. Extensive experiments demonstrate the effectiveness of our proposed method, which outperforms existing image restoration and model fine-tuning approaches across different experimental settings, and retains effectiveness when transferred to models with different architectures. Code will be available at https://github.com/SilverRAN/YARM. Weihang Ran, Wei Yuan 0004, Yinqiang Zheng |
ICML | 2 |
| 2025 | AISFuser: Encoding Maritime Graphical Representations With Temporal Attribute Modeling for Vessel Trajectory PredictionabstractMaritime transportation, vital for nearly 90% of global trade, necessitates precise vessel trajectory prediction for safety and efficiency. Although the Automatic Identification System (AIS) provides a comprehensive data source, how to model these multi-modal and heterogeneous time-varying sequences (such as vessels’ kinetic information and ocean weather factors) poses a formidable challenge. Moreover, most existing approaches are limited by the confined scope of vessel trajectory modeling, making it impossible to consider the unique characteristics of maritime transportation system. To tackle these challenges, we propose a novel framework called AISFuser to i) encode unique maritime traffic network into graphical representations, and ii) introduce the heterogeneity into multi-modal temporal embeddings through Self-Supervised Learning (SSL). Specifically, our AISFuser is constructed by combining an attention-based graph block with a transformer network to encode information across space and time, respectively. In terms of temporal dimension, one SSL auxiliary task is also designed to enhance the heterogeneity of temporal representations and supplement the main vessel prediction task. We validate the effectiveness of the proposed AISFuser on a real-world AIS dataset. Extensive experimental results demonstrate that our method can forecast multiple attributes of vessel trajectory for over 10 hours into the future, outperforming competitive baselines. Zhiwen Zhang 0004, Wei Yuan 0004, Zipei Fan, Xuan Song 0001, Ryosuke Shibasaki |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | Hybrid Network-Based Automatic Seamline Detection for Orthophoto MosaickingabstractSeamline detection is a crucial procedure for orthophoto mosaicking. To eliminate the seam effect caused by geometric discontinuities, seamlines must avoid crossing areas containing the obvious ground object, for which manual processing is usually required. Many existing seamline detection methods can generate seamlines bypassing most of obvious ground objects but always take pixel-level computation which may consume much time. To address this problem, this paper presents a seamline detection approach based on a hybrid network search. First, without auxiliary data, the semi-global block matching (SGBM) algorithm was adopted to generate a disparity map for pairwise orthophoto overlapping area. By using adaptive threshold segmentation, a binary cost map containing the ground objects was obtained. Subsequently, a hybrid network was constructed by edge points and uniform points extracted on the cost map. Finally, seamlines were detected by search on this network-based graph. The essential contribution of the proposed method is that the seamline is searched on a sparse hybrid network instead of a raster cost map. Thus, computational complexity can be significantly decreased and produce fine-tuned seamlines. A Series of comparison experiments were carried out between the proposed and well-established methods, using two benchmark datasets with different characteristics. The comparison results demonstrated that the proposed method could generate high-quality seamlines in terms of visual comparison and statistical evaluation. Moreover, the processing speed has a nearly tenfold improvement compared with the control group methods. Wei Yuan 0004, Jonathan Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Hybrid Feature Embedding for Automatic Building Outline ExtractionabstractBuilding outline extracted from high-resolution aerial images can be used in various application fields such as change detection and disaster assessment. However, traditional CNN model cannot recognize contours very precisely from original images. In this paper, we proposed a CNN and Transformer based model together with active contour model to deal with this problem. We also designed a triple-branch decoder structure to handle different features generated by encoder. Experiment results show that our model outperforms other baseline model on two datasets, achieving 91.1% mIoU on Vaihingen and 83.8% on Bing huts. Weihang Ran, Wei Yuan 0004, Xiaodan Shi, Zipei Fan, Ryosuke Shibasaki |
IGARSS | 2 |
| 2023 | Graph Encoding based Hybrid Vision Transformer for Automatic Road Network ExtractionabstractThis paper introduces a graph encoding-based hybrid vision transformer for automatic road network extraction via high-resolution remote sensing imagery. Given that high-resolution remote sensing images covered large urban areas, traditional segmentation-based road extraction methods usually can generate good binary classification maps in simple structured road surfaces but fail in complex highway and bridge-covered areas. We introduce a graph encoding-based mechanism to address the above issues, enabling the road extraction framework extracts the road segmentation feature and build the graph structure map jointly. Compared to only segmentation-based methods, our approach learns prior geometrical structure information from the extracted ViT feature maps and has a non-local awareness of the whole road network structure. Eexperimental results demonstrated that the proposed approach outperforms the traditional segmentation-based methods. Wei Yuan 0004, Weihang Ran, Xiaodan Shi, Zipei Fan, Ryosuke Shibasaki |
IGARSS | 1 |
| 2023 | MetaTraj: Meta-Learning for Cross-Scene Cross-Object Trajectory PredictionabstractLong-term pedestrian trajectory prediction in crowds is highly valuable for safety driving and social robot navigation. The recent research of trajectory prediction usually focuses on solving the problems of modeling social interactions, physical constraints and multi-modality of futures without considering the generalization of prediction models to other scenes and objects, which is critical for real-world applications. In this paper, we propose a general framework that makes trajectory prediction models able to transfer well across unseen scenes and objects by quickly learning the prior information of trajectories. The trajectory sequences are closely related to the circumstance setting (e.g. exits, roads, buildings, entries etc.) and the objects (e.g. pedestrians, bicycles, vehicles etc.). We argue that those trajectory information varying across scenes and objects makes a trained prediction model not perform well over unseen target data. To address it, we introduce MetaTraj that contains carefully designed sub-tasks and meta-tasks to learn prior information of trajectories related to scenes and objects, which then contributes to accurate long-term future prediction. Both sub-tasks and meta-tasks are generated from trajectory sequences effortlessly and can be easily integrated into many prediction models. Extensive experiments over several trajectory prediction benchmarks demonstrate that MetaTraj can be applied to multiple prediction models and enables them generalize well to unseen scenes and objects. Xiaodan Shi, Haoran Zhang 0002, Wei Yuan 0004, Ryosuke Shibasaki |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Exploring intercity regional similarity using worldwide location-based social network data (demo paper)abstractFinding out similar regions between cities is important to a variety of real-world applications, such as point-of-interest recommendations, site selection, and travel guidance. With the help of the increasing number of location-based social network users, we can measure the intercity similarity from a new perspective of spatiotemporal characteristics of human mobility. In this paper, we developed an interactive intercity regional similarity explorer (IRSE) that 1) visualizes regional spatiotemporal human mobility features, 2) searches similar region candidates in the target city, and 3) explores the regional similarity from different views in a quantitative and illustrative way. In this paper, we show how our system can be useful in exploring regional similarity across cities in the world by use cases, which will interest various users from different countries in the demonstration session. Demo available at: https://bit.ly/3OjwnGu Zipei Fan, Guixu Lin, Wei Yuan 0004, Ryosuke Shibasaki, Pengpeng E, Xuan Song 0001 |
SIGSPATIAL/GIS | 3 |
| 2022 | Online trajectory prediction for metropolitan scale mobility digital twinabstractKnowing "what is happening" and "what will happen" of the mobility in a city is the building block of a data-driven smart city system. In recent years, mobility digital twin that makes a virtual replication of human mobility and predicting or simulating the fine-grained movements of the subjects in a virtual space at a metropolitan scale in near real-time has shown its great potential in modern urban intelligent systems. However, few studies have provided practical solutions. The main difficulties are four-folds: 1) the daily variation of human mobility is hard to model and predict; 2) the transportation network enforces a complex constraints on human mobility; 3) generating a rational fine-grained human trajectory is challenging for existing machine learning models; and 4) making a fine-grained prediction incurs high computational costs, which is challenging for an online system. Bearing these difficulties in mind, in this paper we propose a two-stage human mobility predictor that stratifies the coarse and fine-grained level predictions. In the first stage, to encode the daily variation of human mobility at a metropolitan level, we automatically extract citywide mobility trends as crowd contexts and predict long-term and long-distance movements at a coarse level. In the second stage, the coarse predictions are resolved to a fine-grained level via a probabilistic trajectory retrieval method, which offloads most of the heavy computations to the offline phase. We tested our method using a real-world mobile phone GPS dataset in the Kanto area in Japan, and achieved good prediction accuracy and a time efficiency of about 2 min in predicting future 1h movements of about 220K mobile phone users on a single machine to support more higher-level analysis of mobility prediction. Zipei Fan, Wei Yuan 0004, Renhe Jiang, Quanjun Chen, Xuan Song 0001, Ryosuke Shibasaki |
SIGSPATIAL/GIS | 3 |
| 2022 | Cross-Scale Attention-based Tree Crown Detection via UAV imageryabstractThis paper introduces a cross-scale attention based end-to-end learning framework for tree crown detection via UAV imagery. Given that UAV images covered a large forests, the illumination variations, shadow obstacles and texture repetition always lead to inaccurate tree crown detection results. We introduce a cross-scale attention based mechanism to address the above issues, enabling the tree crown detection framework to reason about the RGB texture information and depth information introduced by the automatically generated depth map jointly. Compared to traditional image based tree crown detection methods, our approach learns prior over geometrical structure information from the real 3D world, which is robust to the texture repetition and small tree crowns. The experimental results demonstrated that the proposed approach outperforms the traditional CNN based method. Wei Yuan 0004, Xiaodan Shi, Zhiling Guo, Zipei Fan, Jianya Gong, Ryosuke Shibasaki |
IGARSS | 1 |
| 2020 | Multimodal Interaction-Aware Trajectory Prediction in Crowded SpaceabstractAccurate human path forecasting in complex and crowded scenarios is critical for collision avoidance of autonomous driving and social robots navigation. It still remains as a challenging problem because of dynamic human interaction and intrinsic multimodality of human motion. Given the observation, there is a rich set of plausible ways for an agent to walk through the circumstance. To address those issues, we propose a spatio-temporal model that can aggregate the information from socially interacting agents and capture the multimodality of the motion patterns. We use mixture density functions to describe the human path and predict the distribution of future paths with explicit density. To integrate more factors to model interacting people, we further introduce a coordinate transformation to represent the relative motion between people. Extensive experiments over several trajectory prediction benchmarks demonstrate that our method is able to forecast various plausible futures in complex scenarios and achieves state-of-the-art performance. Xiaodan Shi, Xiaowei Shao, Zipei Fan, Renhe Jiang, Haoran Zhang 0002, Zhiling Guo, Guangming Wu, Wei Yuan 0004, Ryosuke Shibasaki |
AAAI | 8 |
| 2018 | Semantic Segmentation for Urban Planning Maps Based on U-NetabstractThe automatic digitizing of paper maps is a significant and challenging task for both academia and industry. As an important procedure of map digitizing, the semantic segmentation section is mainly relied on manual visual interpretation with low efficiency. In this study, we select urban planning maps as a representative sample and investigate the feasibility of utilizing U-shape fully convolutional based architecture to perform end-to-end map semantic segmentation. The experimental results obtained from the test area in Shibuya district, Tokyo, demonstrate that our proposed method could achieve a very high Jaccard similarity coefficient of 93.63% and an overall accuracy of 99.36%. For implementation on GPGPU and cuDNN, the required processing time for the whole Shibuya district can be less than three minutes. The results indicate the proposed method can serve as a viable tool for urban planning map semantic segmentation task with high accuracy and efficiency. Zhiling Guo, Hiroaki Shengoku, Guangming Wu, Qi Chen 0012, Wei Yuan 0004, Xiaodan Shi, Xiaowei Shao, Yongwei Xu, Ryosuke Shibasaki |
IGARSS | 5 |