Xinglin Piao

dblp:155/6693 · DBLP profile ↗
← Back
33ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0003-3774-5789ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 13 since 2021Artificial intelligence and machine learning · 9 · 3 first-author · 6 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Enhanced graph collaborative filtering for financial asset recommendation via Haar Kolmogorov-Arnold networks
Xinglin Piao, Wei Zhang 0320, Shiyu Zhao 0005, Yong Zhang 0029
Expert Syst. Appl.2
2026 Frequency-domain multi-scale graph learning with information-theoretic constraint for spatio-temporal prediction
Shun Wang 0004, Yong Zhang 0029, Xuanqi Lin, Guangyu Huo, Xinglin Piao, Yongli Hu
Pattern Recognit.5
2026 RAT: Residual Attention Transformer for Tabular Data
abstract
The effectiveness of Transformer-based methods in tabular data processing has been extensively validated. However, most existing models primarily adopt Transformer-based encoder architectures, which overly emphasize feature correlations while neglecting the unique explicit information representation characteristics of tabular data. The significantly higher information density of tabular data compared to textual data introduces noise that degrades performance. Consequently, directly applying language models to tabular data often yields suboptimal results. To address this limitation, we propose the Residual Attention Transformer (RAT), a Transformer-based mechanism specifically designed for tabular data. In subsequent sections, the Residual Attention Transformer will be referred to as RAT. RAT introduces residual connections into the self-attention mechanism, effectively preserving raw information while capturing dependencies with higher-order features, thus enhancing the richness of information representation. Additionally, we design a simplified Transformer output head, termed Light-Transformer, which consists of a lightweight Transformer block specifically dedicated to the output stage. This design not only significantly reduces the model's parameter count, but also greatly improves training and inference efficiency. Extensive experiments were conducted on six public datasets to evaluate our model. The results show that the RAT model consistently outperforms other models in various tasks, confirming its effectiveness and superiority.
Aiwen Wang, Xinglin Piao, Yujia Guo, Yong Zhang 0029
IEEE Trans. Big Data2
2025 MRAGNN: Refining urban spatio-temporal prediction of crime occurrence with multi-type crime correlation learning
Shun Wang 0004, Yong Zhang 0029, Xinglin Piao, Xuanqi Lin, Yongli Hu
Expert Syst. Appl.3
2025 SHKD: A framework for traffic prediction based on Sub-Hypergraph and Knowledge Distillation
Xiangyu Yao, Xinglin Piao, Qitan Shao, Yongli Hu, Yong Zhang 0029
Knowl. Based Syst.2
2025 Advancing Radar Echo Extrapolation With Hypergraph-Enhanced Latent Diffusion Model
abstract
Radar Echo Extrapolation (REE) can facilitate accurate and expeditious nowcasting of precipitation, reducing the reliance on complex Numerical Weather Prediction (NWP) models. Spatial-temporal forecasting methods dominate this task because they can fully exploit the spatio-temporal dependencies and complex dynamic patterns inherent in radar echo data. However, they struggle with handling uncertainty and incorporating domain-specific knowledge, often resulting in blurry or unrealistic predictions. We propose a Hypergraph-enhanced Latent Diffusion Model (HyDiff) to address these limitations. The EchoDiff has been utilized to aid in accurate extrapolation. To accurately describe precipitation microphysics while adhering to the hydro-microphysical and multiscale coupling principles, the method integrates additional semantic information into the model. Specifically, the Differential Reflectivity Factor (ZDR) and Differential Propagation Phase Shift (KDP) are incorporated into the model as additional semantic information. Furthermore, we introduce a Hypergraph Neural Network (HGNN) into the extrapolation method to capture correlation information across regions. Experiments show that HyDiff effectively handles uncertainty, incorporates domain-specific prior knowledge, and generates forecasts with high operational utility.
Xiaoni Sun, Yong Zhang 0029, Xin Di, Xinglin Piao, Guodong Jing, Dawei Lin
IEEE Trans. Geosci. Remote. Sens.4
2025 CMDNet: A Cross-Modality Spatiotemporal Graph Network for Enhanced Air Pollution Prediction With High-Resolution Satellite Data
abstract
Predicting air pollution plays a vital role in urban management and public health by providing early warnings on PM2.5, SO2, and NO2 concentrations, helping to mitigate the adverse effects of these pollutants. Traditional prediction methods, relying on physical and statistical models, often struggle to capture the complex spatio-temporal dependencies and dynamic characteristics of air pollution data. The application of deep learning methods, especially graph neural networks (GNNs), has shown promise in addressing these limitations. However, existing GNN-based methods ignore the integration of rich semantic information provided by high-resolution satellite data. To address this problem, we propose a Cross-Modality Dynamic Spatio-Temporal Graph Neural Network (CMDNet) for air pollution prediction. The model comprises two branches: a dynamic spatio-temporal graph neural network branch and a remote sensing image dynamic encoding network branch. The dynamic spatiotemporal graph neural network branch captures the spatiotemporal dependencies in air pollution data by constructing a dynamic graph structure. The remote sensing image dynamic encoding network branch extracts semantic information from high-resolution satellite data, which improves the model’s power to perceive air pollution conditions in different regions. Experiments on real-world datasets demonstrate that CMDNet achieves better air pollution prediction results than existing SOTA models, with maximum improvements of 2.4% (MAE), 1.8% (RMSE), 1.3% (CSI), 1.6% (FAR), and 1.4% (POD), providing more accurate prediction results.
Shun Wang 0004, Yong Zhang 0029, Xuanqi Lin, Xinglin Piao, Yongli Hu
IEEE Trans. Geosci. Remote. Sens.4
2025 HGSCO: Heterogeneous Graph Structure Contrast Optimization for Trajectory Prediction
abstract
Predicting and planning the future trajectories of various traffic participants is an important task with multiple applications, including autonomous vehicles, service robots, and intelligent transportation. However, the diversity of heterogeneous agents including pedestrians, bicycles, and vehicles in traffic scenarios presents substantial challenges to this task. Current models do not fully capture the implicit and explicit interaction relationships among these heterogeneous agents and often overlook the significance of extracting implicit correlations from agent features. To address these issues, we introduce a novel model for trajectory prediction: Heterogeneous Graph Structure Contrast Optimization (HGSCO). To accurately capture the interaction relationships among heterogeneous agents, HGSCO constructs semantic graph structures representing implicit relationships and meta-path graph structures representing explicit relationships. Then, HGSCO introduces a cross-view contrastive learning approach, which optimizes the heterogeneous graph structure by maximizing mutual information between the two types of graph structures. The model can provide precise interaction relationships among heterogeneous agents by effectively fusing these two graph representations with a gated fusion method. We utilized video data captured by camera sensors in complex environments with multiple agents to conduct experiments. Our proposed model achieved an 11.5% and 6.7% reduction in Average Displacement Error (ADE) across these datasets, respectively, and a reduction of 15.6% and 8.1% in Final Displacement Error (FDE). The results demonstrate that HGSCO significantly surpasses existing state-of-the-art methods regarding trajectory prediction accuracy.
Xuanqi Lin, Yong Zhang 0029, Shun Wang 0004, Xinglin Piao, Yongli Hu
IEEE Trans. Intell. Transp. Syst.4
2025 ChatTraffic: Text-to-Traffic Generation via Diffusion Model
abstract
The analysis of traffic situations under abnormal conditions is one of the bottleneck issues in Intelligent Transportation Systems (ITS). Influenced by the suddenness, randomness, and uncertainty, this issue is challenging to achieve through existing deep learning methods. It needs to be assisted by traffic simulation models for analysis. However, simulation models always require extensive scene modeling and calibration, making it difficult to meet the demands of natural human-machine interaction in the AIGC (Artificial Intelligence Generated Content) era, as well as the need for rapid and flexible implementation of situation analysis. With the accumulation of traffic data, the emergence of diffusion models offers a new entry point for the core method of data-driven analysis, namely Text-to-Traffic Generation (TTG). In this work, we explore how generative models combined with text describing the traffic system can be applied for traffic situation generation, and propose ChatTraffic, the first diffusion model for TTG. To guarantee the consistency between synthetic and real data, we augment a diffusion model with the Graph Convolutional Network (GCN) to extract spatial correlations of traffic data. In addition, we construct a large-scale dataset containing text-traffic pairs for TTG. We benchmarked ChatTraffic qualitatively and quantitatively on the released dataset. The experimental results indicate that ChatTraffic can rapidly and flexibly generate realistic traffic situations from text, which have practical significance in addressing bottlenecks in ITS. Our code and dataset are available athttps://github.com/ChyaZhang/ChatTraffic.
Yong Zhang 0029, Qitan Shao, Bo Li 0128, Xinglin Piao
IEEE Trans. Intell. Transp. Syst.6
2024 Cross-modal fusion encoder via graph neural network for referring image segmentation
abstract
Abstract Referring image segmentation identifies the object masks from images with the guidance of input natural language expressions. Nowadays, many remarkable cross‐modal decoder are devoted to this task. But there are mainly two key challenges in these models. One is that these models usually lack to extract fine‐grained boundary information and gradient information of images. The other is that these models usually lack to explore language associations among image pixels. In this work, a Multi‐scale Gradient balanced Central Difference Convolution (MG‐CDC) and a Graph convolutional network‐based Language and Image Fusion (GLIF) for cross‐modal encoder, called Graph‐RefSeg, are designed. Specifically, in the shallow layer of the encoder, the MG‐CDC captures comprehensive fine‐grained image features. It could enhance the perception of target boundaries and provide effective guidance for deeper encoding layers. In each encoder layer, the GLIF is used for cross‐modal fusion. It could explore the correlation of every pixel and its corresponding language vectors by a graph neural network. Since the encoder achieves robust cross‐modal alignment and context mining, a light‐weight decoder could be used for segmentation prediction. Extensive experiments show that the proposed Graph‐RefSeg outperforms the state‐of‐the‐art methods on three public datasets. Code and models will be made publicly available at https://github.com/ZYQ111/Graph_refseg .
Yong Zhang 0029, Xinglin Piao, Yongli Hu
IET Image Process.3
2024 Multiagent trajectory prediction with global-local scene-enhanced social interaction graph network
abstract
Abstract Trajectory prediction is essential for intelligent autonomous systems like autonomous driving, behavior analysis, and service robotics. Deep learning has emerged as the predominant technique due to its superior modeling capability for trajectory data. However, deep learning‐based models face challenges in effectively utilizing scene information and accurately modeling agent interactions, largely due to the complexity and uncertainty of real‐world scenarios. To mitigate these challenges, this study presents a novel multiagent trajectory prediction model, termed the global‐local scene‐enhanced social interaction graph network (GLSESIGN), which incorporates two pivotal strategies: global‐local scene information utilization and a social adaptive attention graph network. The model hierarchically learns scene information relevant to multiple intelligent agents, thereby enhancing the understanding of complex scenes. Additionally, it adaptively captures social interactions, improving adaptability to diverse interaction patterns through sparse graph structures. This model not only improves the understanding of complex scenes but also accurately predicts future trajectories of multiple intelligent agents by flexibly modeling intricate interactions. Experimental validation on public datasets substantiates the efficacy of the proposed model. This research offers a novel model to address the complexity and uncertainty in multiagent trajectory prediction, providing more accurate predictive support in practical application scenarios.
Xuanqi Lin, Yong Zhang 0029, Shun Wang 0004, Xinglin Piao
Comput. Animat. Virtual Worlds4
2024 Multi-scale hypergraph-based feature alignment network for cell localization
Bo Li 0128, Yong Zhang 0029, Xinglin Piao, Yongli Hu
Pattern Recognit.4
2024 CSAT: Contrastive Sampling-Aggregating Transformer for Community Detection in Attribute-Missing Networks
abstract
Community detection aims to identify dense subgroups of nodes within a network. However, in real-world networks, node attributes are often missing, making traditional methods less effective. In networks with missing attributes, the main challenge of community detection is to deal with the missing attribute information efficiently and use network structure information to make accurate predictions. This article proposes an innovative method called contrastive sampling-aggregating transformer (CSAT) for community detection in attribute-missing networks. CSAT incorporates the contrastive learning principle to capture hidden patterns among nodes and to aggregate information from different samples to create a more robust and accurate methodology for community detection. Specifically, CSAT utilizes a sampling and propagation strategy to obtain different samples and smooth attribute features of the network structure and leverages the Transformer architecture to model the pairwise relationships between nodes. Therefore, our method can address the attribute-missing issue by integrating the auxiliary information from both the network structure and other sources. Extensive experiments on several benchmark datasets demonstrate CSAT’s superior performance compared to the state-of-the-art methods for community detection.
Mengran Li 0001, Yong Zhang 0029, Wei Zhang 0320, Shiyu Zhao 0005, Xinglin Piao
IEEE Trans. Comput. Soc. Syst.5
2024 Contextual Semantics Interaction Graph Embedding Learning for Recommender Systems
abstract
Recommender systems have become an indispensable tool in today's digital age, significantly enhancing user engagement on various online platforms by curating personalized item recommendations tailored to individual preferences. While the field has long been dominated by the collaborative filtering technique, which primarily leverages user–item interaction data, it often falls short in encapsulating the rich contextual intricacies and evolving dynamics inherent to these interactions. Recognizing this limitation, our research introduces the contextual semantic interaction graph embedding (CSI-GE) method. This advanced model incorporates a dynamic hop window within a multilayer graph convolutional network, ensuring a comprehensive extraction of both immediate and evolving contextual features. By amalgamating self-supervised contrastive learning, we achieve a refinement of user and item embeddings. Furthermore, our innovative variance–invariance–covariance (VIC) regularization-based loss function fortifies the robustness of these embeddings. Through rigorous testing, CSI-GE consistently outperformed contemporary methods, underscoring its superior accuracy and stability.
Shiyu Zhao 0005, Yong Zhang 0029, Mengran Li 0001, Xinglin Piao
IEEE Trans. Comput. Soc. Syst.4
2024 Multi-Information Aggregation and Estrangement HyperGraph Convolutional Networks for Spatiotemporal Weather Forecasting
abstract
Weather forecasting is inextricably linked to human lives and represents a quintessential task of spatiotemporal modeling, necessitated by the spatial and temporal dependencies inherent in meteorological data. Recent studies have consistently shown the excellent performance of graph-based neural networks in accurately modeling spatiotemporal data across various applications. Yet, traditional graph neural networks (GNNs) are unable to handle the high-order diffusion and aggregation phenomena between meteorological data caused by advection. Moreover, the impacts of spatial correlation among multisource information and the presence of noise in meteorological data are often overlooked. This study proposes a novel approach for modeling the spatiotemporal dependencies in meteorological data using the multi-information spatiotemporal aggregation and estrangement hypergraph convolution network. This method employs a novel representation of meteorological data using hypergraphs to address the aforementioned challenges. Specifically, we construct adjacency and semantic hypergraphs to represent spatial correlations and then introduce aggregation and estrangement hypergraph convolution networks to effectively capture multi-information spatial correlations. A new reconstruction feature attention module has been developed to fuse aggregation and estrangement semantic spatial information across various subspaces. In addition, the hypergraph convolution is embedded within a recurrent neural network architecture to model the temporal correlations. Extensive experiments have been conducted on four weather datasets, and state-of-the-art performance has been achieved in comparison to mainstream baseline methods.
Zhuangzhuang Miao, Yong Zhang 0029, Jiayi Wu 0004, Guodong Jing, Xinglin Piao
IEEE Trans. Geosci. Remote. Sens.5
2024 PN-HGNN: Precipitation Nowcasting Network Via Hypergraph Neural Networks
abstract
Precipitation nowcasting within 2 hours is an important and hard issue in weather research area. Benefiting from the outstanding nonlinear relationship modeling capability, methods based on deep learning have achieved significant success in the task of precipitation nowcasting compared to the others. However, existing deep learning based methods always disregard the intricate high-order correlations and lack substantial connections with the evolution of the precipitation system, which would lead blurred forecasts and implausible predictions. To address these issues, we proposed a new Precipitation Nowcasting Network within 2 hours model based on Hypergraph Neural Network (PN-HGNN). In this work, Hypergraph Neural Network is firstly adopted for extracting spatio-temporal dynamic echo features. Secondly, regulation evolution is in charge of capturing the memory features to guide the extrapolation. Finally, we design a dual branch module to extrapolate the radar echoes. The proposed model has been assessed on the dataset HKO-7. The experimental results demonstrate that PN-HGNN achieved better prediction performance than the six representative echo extrapolation models.
Xiaoni Sun, Yong Zhang 0029, Xinglin Piao, Jiayi Wu 0004, Guodong Jing
IEEE Trans. Geosci. Remote. Sens.3
2024 BjTT: A Large-Scale Multimodal Dataset for Traffic Prediction
abstract
Traffic prediction plays a significant role in Intelligent Transportation Systems (ITS). Although many datasets have been introduced to support the study of traffic prediction, most of them only provide time-series traffic data. However, urban transportation systems are always susceptible to various factors, including unusual weather and traffic accidents. Therefore, relying solely on historical data for traffic prediction greatly limits the accuracy of the prediction. In this paper, we introduce Beijing Text-Traffic (BjTT), a large-scale multimodal dataset for traffic prediction. BjTT comprises over 32,000 time-series traffic records, capturing velocity and congestion levels on more than 1,200 roads within the 5th ring area of Beijing. Meanwhile, each piece of traffic data is coupled with a text describing the traffic system (including time, location, and events). We detail the data collection and processing procedures and present a statistical analysis of the BjTT dataset. Furthermore, we conduct comprehensive experiments on the dataset with state-of-the-art traffic prediction methods and text-guided generative models, which reveal the unique characteristics of the BjTT. The dataset is available athttps://github.com/ChyaZhang/BjTT.
Yong Zhang 0029, Qitan Shao, Jiangtao Feng, Bo Li 0128, Xinglin Piao
IEEE Trans. Intell. Transp. Syst.7
2024 CrowdGraph: Weakly supervised Crowd Counting via Pure Graph Neural Network
abstract
Most existing weakly supervised crowd counting methods utilize Convolutional Neural Networks (CNN) or Transformer to estimate the total number of individuals in an image. However, both CNN-based (grid-to-count paradigm) and Transformer-based (sequence-to-count paradigm) methods take images as inputs in a regular form. This approach treats all pixels equally but cannot address the uneven distribution problem within human crowds. This challenge would lead to a decline in the counting performance of the model. Compared with grid and sequence, the graph structure could better explore the relationship among features. In this article, we propose a new graph-based crowd counting method named CrowdGraph, which reinterprets the weakly supervised crowd counting problem from a graph-to-count perspective. In the proposed CrowdGraph, each image is constructed as a graph, and a graph-based network is designed to extract features at the graph level. CrowdGraph comprises three main components: a dynamic graph convolutional backbone, a multi-scale dilated graph convolution module, and a regression head. To the best of our knowledge, CrowdGraph is the first method that is completely formulated based on the Graph Neural Network (GNN) for the crowd counting task. Extensive experiments demonstrate that the proposed CrowdGraph outperforms pure CNN-based and pure Transformer-based weakly supervised methods comprehensively and achieves highly competitive counting performance.
Yong Zhang 0029, Bo Li 0128, Xinglin Piao
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Region feature smoothness assumption for weakly semi-supervised crowd counting
abstract
Abstract Crowd counting is a hot issue in visual data processing. It also plays an important role in the field of video surveillance, social security, and traffic control. However, most of the existing crowd counting methods always adopt a mount of training data or point‐level annotation to learn the mapping relationships between images and density maps, which would cost much human labor. In this paper, we propose a new weakly semi‐supervised crowd counting method which uses less count‐level data for data training. In particular, we extend the classical smoothness assumption and design a many‐to‐many Region Feature Smoothness Assumption to deal with the uneven density distribution problem within crowd region. Further, we adopt hypergraph representation to explore the complex high‐order relationship for different crowd regions. Besides, we design a multi‐scale dynamic hypergraph convolutional module and hyperedge contrastive loss. Extensive experiments have been conducted on five public datasets. The experimental results show that the proposed method outperforms the state‐of‐the‐art ones.
Zhuangzhuang Miao, Yong Zhang 0029, Xinglin Piao, Yi Chu
Comput. Animat. Virtual Worlds3
2023 Principal component analysis based on graph embedding
Fujiao Ju, Yaxiao Zhang, Xinglin Piao
Multim. Tools Appl.5
2023 Hypergraph Association Weakly Supervised Crowd Counting
abstract
Weakly supervised crowd counting involves the regression of the number of individuals present in an image, using only the total number as the label. However, this task is plagued by two primary challenges: the large variation of head size and uneven distribution of crowd density. To address these issues, we propose a novel Hypergraph Association Crowd Counting (HACC) framework. Our approach consists of a new multi-scale dilated pyramid module that can efficiently handle the large variation of head size. Further, we propose a novel hypergraph association module to solve the problem of uneven distribution of crowd density by encoding higher-order associations among features, which opens a new direction to solve this problem. Experimental results on multiple datasets demonstrate that our HACC model achieves new state-of-the-art results.
Bo Li 0128, Yong Zhang 0029, Xinglin Piao
ACM Trans. Multim. Comput. Commun. Appl.4
2022 SHCN: Self-supervised General Hypergraph Clustering Network
abstract
Clustering is a fundamental and hot issue in the unsupervised learning area. With the rapid development of deep learning and graph neural networks (GNNs) techniques, researchers have proposed a series of effective clustering methods. However, most existing approaches adopt a conventional graph to aggregate the neighborhood information, where only the pairwise relations are considered. Moreover, the redundancy/noise in the raw data samples may result in less accurate sample relations and inferior clustering results. In this paper, we proposed a new GNNs based clustering method, which adopts the hypergraph learning approach to explore the high-order relationship for accurate relation learning. Specifically, we first construct two hypergraph representations based on the topology feature and attribute feature from data samples. Then, a self-supervised structure is integrated to learn a cross-correlation matrix from the original hypergraph to act as a higher-order neighborhood with reduced redundancy and noise. Finally the embedding representation of the clustering space is learned in the graph convolution. The proposed method has been evaluated on six public datasets for clustering tasks. Experimental results show that our proposed method outperforms the state-of-the-art ones.
Mengran Li 0001, Xinglin Piao, Yong Zhang 0029, Yongli Hu
IEEE Big Data2
2022 Multitask Hypergraph Convolutional Networks: A Heterogeneous Traffic Prediction Framework
abstract
Traffic prediction methods on a single-source data have achieved excellent results in recent years, especially the Graph Convolutional Networks (GCN) based models with spatio-temporal dependency. In reality, various modes of urban transportation operate simultaneously. They influence and complement each other in common space-time occasions, constituting the transportation system dynamically. Thus, traffic data from multiple sources is ostensibly heterogeneous, but internally correlated. The typical single data driven models are, however, not universally applicable for heterogeneous traffic data. To address this issue, we propose a Multi-task Hypergraph Convolutional Neural Network (MT-HGCN) for the multi-source traffic prediction problem. The framework consists of a main task and a related task. Both tasks are based on Hypergraph Convolutional Neural Networks (HGCN) and are devoted to two prediction problems. Furthermore, the tasks are bridged by a feature compress unit, which models the correlation and shares the latent feature to improve the performance of the main task. The node-level forecasting has been evaluated on historical datasets of Beijing to verify the effectiveness of the proposed method. Compared with the state-of-the-arts, the superior performance of the proposed method can be obtained.
Yong Zhang 0029, Lixun Wang, Yongli Hu, Xinglin Piao
IEEE Trans. Intell. Transp. Syst.5
2021 Metro Passenger Flow Prediction via Dynamic Hypergraph Convolution Networks
abstract
Metro passenger flow prediction is a strategically necessary demand in an intelligent transportation system to alleviate traffic pressure, coordinate operation schedules, and plan future constructions. Graph-based neural networks have been widely used in traffic flow prediction problems. Graph Convolutional Neural Networks (GCN) captures spatial features according to established connections but ignores the high-order relationships between stations and the travel patterns of passengers. In this paper, we utilize a novel representation to tackle this issue - hypergraph. A dynamic spatio-temporal hypergraph neural network to forecast passenger flow is proposed. In the prediction framework, the primary hypergraph is constructed from metro system topology and then extended with advanced hyperedges discovered from pedestrian travel patterns of multiple time spans. Furthermore, hypergraph convolution and spatio-temporal blocks are proposed to extract spatial and temporal features to achieve node-level prediction. Experiments on historical datasets of Beijing and Hangzhou validate the effectiveness of the proposed method, and superior performance of prediction accuracy is achieved compared with the state-of-the-arts.
Yong Zhang 0029, Yongli Hu, Xinglin Piao
IEEE Trans. Intell. Transp. Syst.5
2020 Reweighted Non-convex Non-smooth Rank Minimization Based Spectral Clustering on Grassmann Manifold
Xinglin Piao, Yongli Hu, Junbin Gao, Xin Yang 0011
ACCV (5)1
2020 Kernel Clustering On Symmetric Positive Definite Manifolds Via Double Approximated Low Rank Representation
abstract
As an effective descriptor, Symmetric Positive Definite (SPD) matrix is widely used in several areas such as image clustering. Recently, researchers proposed some effective methods based on low rank theory for SPD data clustering with nonlinear metric. However, single nuclear norm is always adopted to formulate the low rank model in these methods, which would lead to suboptimal solution. In this paper, we proposed a novel double low rank representation method for SPD clustering problem, in which matrix factorization and nonconvex rank constraint are combined to reveal the intrinsic property of the data instead of employing the nuclear norm. Meanwhile, kernel method and Log-Euclidean metric are combined to better explore the intrinsic geometry within SPD data. The proposed method has been evaluated on several public datasets and the experimental results demonstrate that the proposed method outperforms the state-of-the-art ones.
Xinglin Piao, Yongli Hu, Junbin Gao, Xin Yang 0011, Wenwu Zhu 0001, Ge Li 0002
ICME1
2020 A Spectral Clustering on Grassmann Manifold via Double Low Rank Constraint
abstract
Data clustering is a fundamental topic in machine learning and data mining areas. In recent years, researchers have proposed a series of effective methods based on Low Rank Representation (LRR) which could explore low-dimension subspace structure embedded in original data effectively. The traditional LRR methods usually are designed for vectorial data from linear spaces with Euclidean distance. However, high-dimension data (such as video clip or imageset) are always considered as non-linear manifold data such as Grassmann manifold with non-linear metric. In addition, traditional LRR clustering method always adopt single nuclear norm as low rank constraint which would lead to suboptimal solution and decrease the clustering accuracy. In this paper, we proposed a new low rank method on Grassmann manifold for video or imageset data clustering task. In the proposed method, video or imageset data are formulated as sample data on Grassmann manifold first. And then a double low rank constraint is proposed by combining the nuclear norm and bilinear representation for better construct the representation matrix. The experimental results on several public datasets show that the proposed method outperforms the state-of-the-art clustering methods.
Xinglin Piao, Yongli Hu, Junbin Gao, Xin Yang 0011
ICPR1
2020 Point cloud semantic scene segmentation based on coordinate convolution
abstract
Abstract Point cloud semantic segmentation, a crucial research area in the 3D computer vision, lies at the core of many vision and robotics applications. Due to the irregular and disordered of the point cloud, however, the application of convolution on point clouds is challenging. In this article, we propose the “coordinate convolution,” which can effectively extract local structural information of the point cloud, to solve the inapplicability of conventional convolution neural network (CNN) structures on the 3D point cloud. The “coordinate convolution” is a projection operation of three planes based on the local coordinate system of each point. Specifically, we project the point cloud on three planes in the local coordinate system with a joint 2D convolution operation to extract its features. Additionally, we leverage a self‐encoding network based on image semantic segmentation U‐Net structure as the overall architecture of the point cloud semantic segmentation algorithm. The results demonstrate that the proposed method exhibited excellent performances for point cloud data sets corresponding to various scenes.
Zhaoxuan Zhang, Xuefeng Yin, Xinglin Piao, Yuxin Wang 0001, Xin Yang 0011
Comput. Animat. Virtual Worlds4
2019 Double Nuclear Norm Based Low Rank Representation on Grassmann Manifolds for Clustering
abstract
Unsupervised clustering for high-dimension data (such as imageset or video) is a hard issue in data processing and data mining area since these data always lie on a manifold (such as Grassmann manifold). Inspired of Low Rank representation theory, researchers proposed a series of effective clustering methods for high-dimension data with non-linear metric. However, most of these methods adopt the traditional single nuclear norm as the relaxation of the rank function, which would lead to suboptimal solution deviated from the original one. In this paper, we propose a new low rank model for high-dimension data clustering task on Grassmann manifold based on the Double Nuclear norm which is used to better approximate the rank minimization of matrix. Further, to consider the inner geometry or structure of data space, we integrated the adaptive Laplacian regularization to construct the local relationship of data samples. The proposed models have been assessed on several public datasets for imageset clustering. The experimental results show that the proposed models outperform the state-of-the-art clustering ones.
Xinglin Piao, Yongli Hu, Junbin Gao
CVPR1
2019 Real-virtual consistent traffic flow interaction
Xin Yang 0011, Shuai Li 0014, Xinglin Piao, Qiang Zhang 0008, Xiaopeng Wei
Graph. Model.4
2019 Cascaded network with deep intensity manipulation for scene understanding
abstract
Abstract Scene understanding is essential to robotic navigation and autonomous driving as it provides semantic information to their controlling system. However, it will fail when processing low‐light images/videos captured under adverse weather or at night use state‐of‐the‐art scene understanding methods. A naive way to directly infer semantics from low‐light images is ill posed because the low‐light condition distorts pixel intensities and buries details. In order to address this problem, we propose the Deep Intensity Manipulation Network (DIMNet), which could relight the input images and recover the details, and combine the DIMNet with a scene understanding network to get a cascaded network to learn the semantics from low‐light images. Through learning pixel intensity manipulation, our method can generate images not only visually pleasing but also practical for scene understanding. Qualitative and quantitative experiments demonstrate that the proposed method is effective and robust for both synthetic and real‐world images.
Xin Yang 0011, Shaozhe Chen, Xinglin Piao, Qiang Zhang 0008, Xiaopeng Wei
Comput. Animat. Virtual Worlds4
2019 Traffic Data Reconstruction via Adaptive Spatial-Temporal Correlations
abstract
Data missing remains a difficult and important problem in the transportation information system, which seriously restricts the application of the intelligent transportation system (ITS), dominatingly on traffic monitoring, e.g., traffic data collection, traffic state estimation, and traffic control. Numerous traffic data imputation methods had been proposed in the last decade. However, lacking of sufficient temporal variation characteristic analysis as well as spatial correlation measurements leads to limited completion precision, and poses a major challenge for an ITS. Leveraging the low-rank nature and the spatial-temporal correlation of traffic network data, this paper proposes a novel approach to reconstruct the missing traffic data based on low-rank matrix factorization, which elaborates the potential implications of the traffic matrix by decomposed factor matrices. To further exploit the temporal evolvement characteristics and the spatial similarity of road links, we design a time-series constraint and an adaptive Laplacian regularization spatial constraint to explore the local relationship with road links. The experimental results on six real-world traffic data sets show that our approach outperforms the other methods and can successfully reconstruct the road traffic data precisely for various structural loss modes.
Yang Wang 0068, Yong Zhang 0029, Xinglin Piao, Hao Liu 0040, Ke Zhang 0016
IEEE Trans. Intell. Transp. Syst.3
2016 Fisher discrimination-based l2, 1-norm sparse representation for face recognition
Yong Zhang 0029, Yongli Hu, Xinglin Piao, Qianjun Wu
Vis. Comput.6