Yatao Zhang

dblp:166/4149 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MCCT-Net: A multi-perspective convolution and cross-connection transformer network for epilepsy detection
Guozhen Sun, Jinbiao Zhang, Yijun Ma, Jilin Wang, Yatao Zhang
Neurocomputing6
2025 A sequential MAE-clustering self-supervised learning method for arrhythmia detection
Yatao Zhang, Jilin Wang, Shipeng Jiang, Yijun Ma
Expert Syst. Appl.1
2025 Online dynamic influence maximization based on deep reinforcement learning
Nuan Song, Wei Sheng, Yanhao Sun, Zhanxue Xu, Yatao Zhang
Neurocomputing8
2025 A Novel Multi-Modal Population-Graph Based Framework for Patients of Esophageal Squamous Cell Cancer Prognostic Risk Prediction
abstract
Prognostic risk prediction is pivotal for clinicians to appraise the patient's esophageal squamous cell cancer (ESCC) progression status precisely and tailor individualized therapy treatment plans. Currently, CT-based multi-modal prognostic risk prediction methods have gradually attracted the attention of researchers for their universality, which is also able to be applied in scenarios of preoperative prognostic risk assessment in the early stages of cancer. However, much of the current work focuses only on CT images of the primary tumor, ignoring the important role that CT images of lymph nodes play in prognostic risk prediction. Additionally, it is important to consider and explore the inter-patient feature similarity in prognosis when developing models. To solve these problems, we proposed a novel multi-modal population-graph based framework leveraging CT images including primary tumor and lymph nodes combined with clinical, hematology, and radiomics data for ESCC prognostic risk prediction. A patient population graph was constructed to excavate the homogeneity and heterogeneity of inter-patient feature embedding. Moreover, a novel node-level multi-task joint loss was proposed for graph model optimization through a supervised-based task and an unsupervised-based task. Sufficient experimental results show that our model achieved state-of-the-art performance compared with other baseline models as well as the gold standard on discriminative ability, risk stratification, and clinical utility.
Shuai Wang 0003, Yaqi Wang 0002, Chengkai Wang, Huiyu Zhou 0001, Yatao Zhang
IEEE J. Biomed. Health Informatics6
2025 Context-Aware Knowledge Graph Framework for Traffic Speed Forecasting Using Graph Neural Network
abstract
Human mobility is intricately influenced by urban contexts spatially and temporally, constituting essential domain knowledge in understanding traffic systems. While existing traffic forecasting models primarily rely on raw traffic data and advanced deep learning techniques, incorporating contextual information remains underexplored due to insufficient integration frameworks and the complexity of urban contexts. This study proposes a novel context-aware knowledge graph (CKG) framework to enhance traffic speed forecasting by effectively modeling spatial and temporal contexts. Employing a relation-dependent integration strategy, the framework generates context-aware representations from the spatial and temporal units of CKG to capture spatio-temporal dependencies of urban contexts. A CKG-GNN model, combining the CKG, dual-view multi-head self-attention (MHSA), and graph neural network (GNN), is then designed to predict traffic speed utilizing these context-aware representations. Our experiments demonstrate that CKG’s configuration significantly influences embedding performance, with ComplEx and KG2E emerging as optimal for embedding spatial and temporal units, respectively. The CKG-GNN model establishes a benchmark for 10-120 min predictions, achieving average MAE, MAPE, and RMSE of 3.46 $\pm$ 0.01, 14.76 $\pm$ 0.09%, and 5.08 $\pm$ 0.01, respectively. Compared to the baseline DCRNN model, integrating the spatial unit improves the MAE by 0.04 and the temporal unit by 0.13, while integrating both units further reduces it by 0.18. The dual-view MHSA analysis reveals the crucial role of relation-dependent features from the context-based view and the model’s ability to prioritize recent time slots in prediction from the sequence-based view. Overall, this study underscores the importance of merging context-aware knowledge graphs with graph neural networks to improve traffic forecasting.
Yatao Zhang, Yi Wang 0123, Song Gao 0001, Martin Raubal
IEEE Trans. Intell. Transp. Syst.1
2024 MMFusion: Multi-modality Diffusion Model for Lymph Node Metastasis Diagnosis in Esophageal Cancer
Chengkai Wang, Huiyu Zhou 0001, Yatao Zhang, Yaqi Wang 0002, Shuai Wang 0003
MICCAI (5)4
2023 Unsupervised land-use change detection using multi-temporal POI embedding
abstract
Rapid land-use change detection (LUCD) is pivotal for refined urban planning and management. In this paper, we investigate LUCD through learning embeddings of points of interest (POIs) from multiple temporalities. There are several prominent challenges: (1) the co-occurrence problem of multi-temporal POIs, (2) the heterogeneity of POI categorization, and (3) The lack of human-crafted labels. Therefore, multi-temporal POIs need to be aligned in the embedding space for effective LUCD. This study proposes a multi-temporal POI embedding (MT-POI2Vec) technique for LUCD in a fully unsupervised manner. In MT-POI2Vec, we first utilize random walks in POI networks to capture their single-period co-occurrence patterns; then, we leverage manifold learning to capture (1) single-period categorical semantics of POIs to enforce semantically similar POI embedding to be close and (2) cross-period categorical semantics to align multi-temporal POI embedding in a unified embedding space. We conducted experiments in Shenzhen, China, which demonstrates that the proposed method is effective. Compared with several baseline models, MT-POI2Vec can better align multi-temporal POIs and thus achieve higher performance in LUCD. In addition, our model can effectively identify areas with unchanged land use and land use changes in residential and industrial areas at a fine scale.
Yao Yao 0004, Qia Zhu, Zijin Guo, Weiming Huang 0001, Yatao Zhang, Xiaoqin Yan, Anning Dong, Zhangwei Jiang, Qingfeng Guan 0001
Int. J. Geogr. Inf. Sci.5
2023 Incorporating multimodal context information into traffic speed forecasting through graph deep learning
abstract
Accurate traffic speed forecasting is a prerequisite for anticipating future traffic status and increasing the resilience of intelligent transportation systems. However, most studies ignore the involvement of context information ubiquitously distributed over the urban environment to boost speed prediction. The diversity and complexity of context information also hinder incorporating it into traffic forecasting. Therefore, this study proposes a multimodal context-based graph convolutional neural network (MCGCN) model to fuse context data into traffic speed prediction, including spatial and temporal contexts. The proposed model comprises three modules, ie (a) hierarchical spatial embedding to learn spatial representations by organizing spatial contexts from different dimensions, (b) multivariate temporal modeling to learn temporal representations by capturing dependencies of multivariate temporal contexts and (c) attention-based multimodal fusion to integrate traffic speed with the spatial and temporal context representations for multi-step speed prediction. We conduct extensive experiments in Singapore. Compared to the baseline model (spatial-temporal graph convolutional network, STGCN), our results demonstrate the importance of multimodal contexts with the mean-absolute-error improvement of 0.29 km/h, 0.45 km/h and 0.89 km/h in 30-min, 60-min and 120-min speed prediction, respectively. We also explore how different contexts affect traffic speed forecasting, providing references for stakeholders to understand the relationship between context information and transportation systems.
Yatao Zhang, Tianhong Zhao, Song Gao 0001, Martin Raubal
Int. J. Geogr. Inf. Sci.1
2021 Scale Effect on Fusing Remote Sensing and Human Sensing to Portray Urban Functions
abstract
The development of information and communication technologies has produced massive human sensing data sets, such as point of interest, mobile phone data, and social media data sets. These data sets provide alternative human perceptions of urban spaces; therefore, they have become effective supplements for remote sensing tasks. This letter presents an exploratory framework to examine the scale effect of fusing remote sensing and human sensing. The physical and social semantics are extracted from raw remote sensing images and human sensing data, respectively. A dynamic weighting strategy is developed to explore the fusion of remote sensing and human sensing. Taking urban function inference as an example, the scale effect is evaluated by weighting remote sensing and human sensing. The experiment demonstrates that fusing remote sensing and human sensing enables us to recognize multiple types of urban functions. Meanwhile, the results are significantly affected by the scale.
Wei Tu 0001, Yatao Zhang, Qingquan Li 0001, Ke Mai, Jinzhou Cao
IEEE Geosci. Remote. Sens. Lett.2
2021 Real-Time Route Recommendations for E-Taxies Leveraging GPS Trajectories
abstract
Electric vehicles (EVs) currently face formidable challenges in promotion, i.e., short driving ranges, long charging times, and few charging stations, thereby limiting their acceptability to taxi drivers. Leveraging massive-scale taxi GPS trajectory data, we present a novel real-time route recommendation system for electric taxi (ET) drivers. Taxi travel knowledge, including the probability of picking up passengers and the distribution of destinations, is learned from the raw GPS trajectories. Considering the cascading effect of route decision making, consecutive ET actions are modeled with an action tree. The corresponding expected net revenue is estimated based on the learned knowledge. A prototype online system is developed for providing route recommendations, e.g., when to go to a charging station or cruise on certain roads. An experiment in Shenzhen demonstrates that the average daily net revenue of ET drivers is better than those of 76.2% of gasoline taxi drivers. The presented approach not only increases the revenue of ET drivers in the short term but also improves the viability of EVs in the long run.
Wei Tu 0001, Ke Mai, Yatao Zhang, Yang Xu 0002, Jincai Huang 0002, Long Chen 0005, Qingquan Li 0001
IEEE Trans. Ind. Informatics3
2021 Heartbeats Classification Using Hybrid Time-Frequency Analysis and Transfer Learning Based on ResNet
abstract
The classification of heartbeats is an important method for cardiac arrhythmia analysis. This study proposes a novel heartbeat classification method using hybrid time-frequency analysis and transfer learning based on ResNet-101. The proposed method has the following major advantages over the afore-mentioned methods: it avoids the need for manual features extraction in the traditional machine learning method, and it utilizes 2-D time-frequency diagrams which provide not only frequency and energy information but also preserve the morphological characteristic within the ECG recordings, and it owns enough deep to make better use of performance of CNN. The method deploys a hybrid time-frequency analysis of the Hilbert transform (HT) and the Wigner-Ville distribution (WVD) to transform 1-D ECG recordings into 2-D time-frequency diagrams which were then fed into a transfer learning classifier based on ResNet-101 for two classification tasks (i.e., 5 heartbeat categories assigned by the ANSI/AAMI standard (i.e., N, V, S, Q and F) and 14 original beat kinds of the MIT/BIH arrhythmia database). For 5 heartbeat categories classification, the results show the F1-score of N, V, S, Q and F categories areF$_{N}$0.9899,F$_{V}$0.9845,F$_{S}$0.9376,F$_{Q}$0.9968,F$_{F}$0.8889, respectively, and the overall F1-score is 0.9595 using the combination data balancing. The results show the average values for accuracy, sensitivity, specificity, predictive value and F1-score on test set for 14 beat kinds the MIT-BIH arrhythmia database are 99.75%, 91.36%, 99.85%, 90.81% and 0.9016, respectively. Compared with other methods, the proposed method can yield more accurate results.
Yatao Zhang, Shoushui Wei, Fengyu Zhou 0002, Dong Li 0052
IEEE J. Biomed. Health Informatics1
2020 Parameters Analysis of Sample Entropy, Permutation Entropy and Permutation Ratio Entropy for RR Interval Time Series
Jian Yin 0003, Pengxiang Xiao, Yungang Liu, Chenggang Yan 0001, Yatao Zhang
Inf. Process. Manag.6
2017 Simulating urban land-use changes at a large scale by integrating dynamic land parcel subdivision and vector-based cellular automata
abstract
Cellular automata (CA) have been widely used to simulate complex urban development processes. Previous studies indicated that vector-based cellular automata (VCA) could be applied to simulate urban land-use changes at a realistic land parcel level. Because of the complexity of VCA, these studies were conducted at small scales or did not adequately consider the highly fragmented processes of urban development. This study aims to build an effective framework called dynamic land parcel subdivision (DLPS)-VCA to accurately simulate urban land-use change processes at the land parcel level. We introduce this model in urban land-use change simulations to reasonably divide land parcels and introduce a random forest algorithm (RFA) model to explore the transition rules of urban land-use changes. Finally, we simulate the land-use changes in Shenzhen between 2009 and 2014 via the proposed DLPS-VCA model. Compared to the advanced Patch-CA and RFA-VCA models, the DLPS-VCA model achieves the highest simulation accuracy (Figure-of-Merit = 0.232), which is 32.57% and 18.97% higher respectively, and is most similar to the actual land-use scenario (similarity = 94.73%) at the pattern level. These results indicate that the DLPS-VCA model can both accurately split the land during urban land-use changes and significantly simulate urban expansion and urban land-use changes at a fine scale. Furthermore, the land-use change rules that are based on DPLS-VCA mining and the simulation results of several future urban development scenarios can act as guides for future urban planning policy formulation.
Yao Yao 0004, Xiaoping Liu 0001, Xia Li 0001, Penghua Liu, Ye Hong, Yatao Zhang, Ke Mai
Int. J. Geogr. Inf. Sci.6
2017 Mapping fine-scale population distributions at the building level by integrating multisource geospatial big data
abstract
Fine-scale population distribution data at the building level play an essential role in numerous fields, for example urban planning and disaster prevention. The rapid technological development of remote sensing (RS) and geographical information system (GIS) in recent decades has benefited numerous population distribution mapping studies. However, most of these studies focused on global population and environmental changes; few considered fine-scale population mapping at the local scale, largely because of a lack of reliable data and models. As geospatial big data booms, Internet-collected volunteered geographic information (VGI) can now be used to solve this problem. This article establishes a novel framework to map urban population distributions at the building scale by integrating multisource geospatial big data, which is essential for the fine-scale mapping of population distributions. First, Baidu points-of-interest (POIs) and real-time Tencent user densities (RTUD) are analyzed by using a random forest algorithm to down-scale the street-level population distribution to the grid level. Then, we design an effective iterative building-population gravity model to map population distributions at the building level. Meanwhile, we introduce a densely inhabited index (DII), generated by the proposed gravity model, which can be used to estimate the degree of residential crowding. According to a comparison with official community-level census data and the results of previous population mapping methods, our method exhibits the best accuracy (Pearson R = .8615, RMSE = 663.3250, p < .0001). The produced fine-scale population map can offer a more thorough understanding of inner city population distributions, which can thus help policy makers optimize the allocation of resources.
Yao Yao 0004, Xiaoping Liu 0001, Xia Li 0001, Jinbao Zhang 0001, Zhaotang Liang, Ke Mai, Yatao Zhang
Int. J. Geogr. Inf. Sci.7
2016 FR-Index: A Multi-dimensional Indexing Framework for Switch-Centric Data Centers
Yatao Zhang, Jialiang Cao, Xiaofeng Gao 0001, Guihai Chen
DEXA (2)1
2015 Indexing Multi-dimensional Data in Modular Data Centers
Libo Gao, Yatao Zhang, Xiaofeng Gao 0001, Guihai Chen
DEXA (2)2
2014 ECG quality assessment based on a kernel support vector machine and genetic algorithm with a feature matrix
abstract
We propose a systematic ECG quality classification method based on a kernel support vector machine (KSVM) and genetic algorithm (GA) to determine whether ECGs collected via mobile phone are acceptable or not. This method includes mainly three modules, i.e., lead-fall detection, feature extraction, and intelligent classification. First, lead-fall detection is executed to make the initial classification. Then the power spectrum, baseline drifts, amplitude difference, and other time-domain features for ECGs are analyzed and quantified to form the feature matrix. Finally, the feature matrix is assessed using KSVM and GA to determine the ECG quality classification results. A Gaussian radial basis function (GRBF) is employed as the kernel function of KSVM and its performance is compared with that of the Mexican hat wavelet function (MHWF). GA is used to determine the optimal parameters of the KSVM classifier and its performance is compared with that of the grid search (GS) method. The performance of the proposed method was tested on a database from PhysioNet/Computing in Cardiology Challenge 2011, which includes 1500 12-lead ECG recordings. True positive (TP), false positive (FP), and classification accuracy were used as the assessment indices. For training database set A (1000 recordings), the optimal results were obtained using the combination of lead-fall, GA, and GRBF methods, and the corresponding results were: TP 92.89%, FP 5.68%, and classification accuracy 94.00%. For test database set B (500 recordings), the optimal results were also obtained using the combination of lead-fall, GA, and GRBF methods, and the classification accuracy was 91.80%.
Yatao Zhang, Chengyu Liu 0001, Shoushui Wei, Chang-zhi Wei
J. Zhejiang Univ. Sci. C1