Xinxin Hu

dblp:05/7267 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
12since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2024 MDIINet: A Few-Shot Semantic Segmentation Network by Exploiting Multi-dimensional Information Interaction
Hao Chen 0037, Xintao Qiu, Tao Wang 0085, Xinxin Hu, Yongyao Wen, Xinlin Zhang, Tong Tong 0001
ICIC (6)5
2024 Insider Threat Defense Strategies: Survey and Knowledge Integration
Chengyu Song, Jingjing Zhang 0005, Linru Ma, Xinxin Hu, Jianming Zheng, Lin Yang 0031
KSEM (5)4
2024 GCAT: graph calibration attention transformer for robust object tracking
Si Chen 0002, Xinxin Hu, Dahan Wang, Yan Yan 0001, Shunzhi Zhu
Neural Comput. Appl.2
2024 GAT-COBO: Cost-Sensitive Graph Neural Network for Telecom Fraud Detection
abstract
Along with the rapid evolution of mobile communication technologies, such as 5G, there has been a significant increase in telecom fraud, which severely dissipates individual fortune and social wealth. In recent years, graph mining techniques are gradually becoming a mainstream solution for detecting telecom fraud. However, the graph imbalance problem, caused by the Pareto principle, brings severe challenges to graph data mining. This emerging and complex issue has received limited attention in prior research. In this paper, we propose aGraphATtention network withCOst-sensitiveBOosting (GAT-COBO) for the graph imbalance problem. First, we design a GAT-based base classifier to learn the embeddings of all nodes in the graph. Then, we feed the embeddings into a well-designed cost-sensitive learner for imbalanced learning. Next, we update the weights according to the misclassification cost to make the model focus more on the minority class. Finally, we sum the node embeddings obtained by multiple cost-sensitive learners to obtain a comprehensive node representation, which is used for the downstream anomaly detection task. Extensive experiments on two real-world telecom fraud detection datasets demonstrate that our proposed method is effective for the graph imbalance problem, outperforming the state-of-the-art GNNs and GNN-based fraud detectors. In addition, our model is also helpful for solving the widespread over-smoothing problem in GNNs. The GAT-COBO code and datasets are available athttps://github.com/xxhu94/GAT-COBO.
Xinxin Hu, Haotian Chen 0003, Hongchang Chen, Xing Li 0013, Xiangyang Xue 0001
IEEE Trans. Big Data1
2024 Cost-Sensitive GNN-Based Imbalanced Learning for Mobile Social Network Fraud Detection
abstract
In recent years, the increasing prevalence of mobile social network fraud has led to significant distress and depletion of personal and social wealth, resulting in considerable economic harm. Graph neural networks (GNNs) have emerged as a popular approach to tackle this issue. However, the challenge of graph imbalance, which can greatly impede the effectiveness of GNN-based fraud detection methods, has received little attention in prior research. Thus, we are going to present a novel cost-sensitive graph neural network (CSGNN) in this article. Initially, reinforcement learning is utilized to train a suitable sampling threshold, followed by neighbor sampling based on node similarity, which helps to alleviate the graph imbalance issue preliminarily. Subsequently, message aggregation is executed on the sampled graph using GNN to obtain node embeddings. Concurrently, the optimization objective for the cost matrix is formulated using the sample histogram matrix, scatter matrix, and confusion matrix. The cost matrix and GNN are collaboratively optimized through the backpropagation algorithm. Ultimately, the derived cost-sensitive node embedding is employed for fraudulent node detection. Furthermore, this study provides a theoretical demonstration of the effectiveness of adaptive cost-sensitive learning in GNN. Extensive experiments are carried out on two publicly accessible real-world mobile network fraud datasets, revealing that the proposed CSGNN effectively addresses the graph imbalance issue while outperforming state-of-the-art algorithms in detection performance. The CSGNN code and datasets can be accessed at https://github.com/xxhu94/CSGNN.
Xinxin Hu, Haotian Chen 0003, Hongchang Chen, Xing Li 0013, Xiangyang Xue 0001
IEEE Trans. Comput. Soc. Syst.1
2023 Who are the evil backstage manipulators: Boosting graph attention networks against deep fraudsters
Xinxin Hu, Hongchang Chen, Haocong Jiang, Kai Wang 0066
Comput. Networks1
2023 CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation With Transformers
abstract
Scene understanding based on image segmentation is a crucial component of autonomous vehicles. Pixel-wise semantic segmentation of RGB images can be advanced by exploiting complementary features from the supplementary modality (${X}$-modality). However, covering a wide variety of sensors with a modality-agnostic model remains an unresolved problem due to variations in sensor characteristics among different modalities. Unlike previous modality-specific methods, in this work, we propose a unified fusion framework, CMX, for RGB-X semantic segmentation. To generalize well across different modalities, that often include supplements as well as uncertainties, a unified cross-modal interaction is crucial for modality fusion. Specifically, we design a Cross-Modal Feature Rectification Module (CM-FRM) to calibrate bi-modal features by leveraging the features from one modality to rectify the features of the other modality. With rectified feature pairs, we deploy a Feature Fusion Module (FFM) to perform sufficient exchange of long-range contexts before mixing. To verify CMX, for the first time, we unify five modalities complementary to RGB, i.e., depth, thermal, polarization, event, and LiDAR. Extensive experiments show that CMX generalizes well to diverse multi-modal fusion, achieving state-of-the-art performances on five RGB-Depth benchmarks, as well as RGB-Thermal, RGB-Polarization, and RGB-LiDAR datasets. Besides, to investigate the generalizability to dense-sparse data fusion, we establish an RGB-Event semantic segmentation benchmark based on the EventScape dataset, on which CMX sets the new state-of-the-art. The source code of CMX is publicly available athttps://github.com/huaaaliu/RGBX_Semantic_Segmentation.
Jiaming Zhang 0001, Huayao Liu, Kailun Yang 0001, Xinxin Hu, Ruiping Liu 0001, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.4
2022 BTG: A Bridge to Graph machine learning in telecommunications fraud detection
Xinxin Hu, Hongchang Chen, Haocong Jiang, Guanghan Chu
Future Gener. Comput. Syst.1
2022 Collusive anomalies detection based on collaborative markov random field
abstract
Abnormal collusive behavior, widely existing in various fields with concealment and synergy, is particularly harmful in user-generated online reviews and hard to detect by traditional methods. With the development of network science, this problem can be solved by analyzing structure features. As a graph-based anomaly detection method, the Markov random field (MRF)-based model has been widely used to identify the collusive anomalies and shown its effectiveness. However, existing methods are mostly unable to highlight the primary synergy relationship among nodes and consider much irrelevant information, which caused poor detectability. Therefore, this paper proposes a novel MRF-based method (ACEagle), considering node-level and community-level behavior features. Our method has several advantages: (1) based on the analysis of the nodes’ local structure, the community-level behavioral features are combined to calculate the nodes’ prior probability to close the ground truth, (2) it measured the behavior’s collaborative intensity between nodes by time and weight, constructing MRF by the synergic relationship exceeding the threshold to filter irrelevant structural information, (3) it operates in a completely unsupervised fashion requiring no labeled data, while still incorporating side information if available. Through experiments in user-reviewed datasets where abnormal collusive behavior is most typical, the results show that ACEagle is significantly outperforming state-of-the-art baselines in collusive anomalies detection.
Lixin Ji, Kai Wang 0066, Xinxin Hu
Intell. Data Anal.5
2022 Omnisupervised Omnidirectional Semantic Segmentation
abstract
Modern efficient Convolutional Neural Networks (CNNs) are able to perform semantic segmentation both swiftly and accurately, which covers typically separate detection tasks desired by Intelligent Vehicles (IV) in a unified way. Most of the current semantic perception frameworks are designed to work with pinhole cameras and benchmarked against public datasets with narrow Field-of-View (FoV) images. However, there is a large accuracy downgrade when a pinhole-yielded CNN is taken to omnidirectional imagery, causing it unreliable for surrounding perception. In this paper, we propose an omnisupervised learning framework for efficient CNNs, which bridges multiple heterogeneous data sources that are already available in the community, bypassing the labor-intensive process to have manually annotated panoramas, while improving their reliability in unseen omnidirectional domains. Being omnisupervised, the efficient CNN exploits both labeled pinhole images and unlabeled panoramas. The framework is based on our specialized ensemble method that considers the wide-angle and wrap-around features of omnidirectional images, to automatically generate panoramic labels for data distillation. A comprehensive variety of experiments demonstrates that the proposed solution helps to attain significant generalizability gains in panoramic imagery domains. Our approach outperforms state-of-the-art efficient segmenters on highly unconstrained IDD20K and PASS datasets.
Kailun Yang 0001, Xinxin Hu, Yicheng Fang, Kaiwei Wang, Rainer Stiefelhagen
IEEE Trans. Intell. Transp. Syst.2
2021 Capturing Omni-Range Context for Omnidirectional Segmentation
abstract
Convolutional Networks (ConvNets) excel at semantic segmentation and have become a vital component for perception in autonomous driving. Enabling an all-encompassing view of street-scenes, omnidirectional cameras present themselves as a perfect fit in such systems. Most segmentation models for parsing urban environments operate on common, narrow Field of View (FoV) images. Transferring these models from the domain they were designed for to 360° perception, their performance drops dramatically, e.g., by an absolute 30.0% (mIoU) on established test-beds. To bridge the gap in terms of FoV and structural distribution between the imaging domains, we introduce Efficient Concurrent Attention Networks (ECANets), directly capturing the inherent long-range dependencies in omnidirectional imagery. In addition to the learned attention-based contextual priors that can stretch across 360° images, we upgrade model training by leveraging multi-source and omni-supervised learning, taking advantage of both: Densely labeled and unlabeled data originating from multiple datasets. To foster progress in panoramic image segmentation, we put forward and extensively evaluate models on Wild PAnoramic Semantic Segmentation (WildPASS), a dataset designed to capture diverse scenes from all around the globe. Our novel model, training regimen and multi-source prediction fusion elevate the performance (mIoU) to new state-of-the-art results on the public PASS (60.2%) and the fresh WildPASS (69.0%) benchmarks.1
Kailun Yang 0001, Jiaming Zhang 0001, Simon Reiß, Xinxin Hu, Rainer Stiefelhagen
CVPR4
2021 Is Context-Aware CNN Ready for the Surroundings? Panoramic Semantic Segmentation in the Wild
abstract
Semantic segmentation, unifying most navigational perception tasks at the pixel level has catalyzed striking progress in the field of autonomous transportation. Modern Convolution Neural Networks (CNNs) are able to perform semantic segmentation both efficiently and accurately, particularly owing to their exploitation of wide context information. However, most segmentation CNNs are benchmarked against pinhole images with limited Field of View (FoV). Despite the growing popularity of panoramic cameras to sense the surroundings, semantic segmenters have not been comprehensively evaluated on omnidirectional wide-FoV data, which features rich and distinct contextual information. In this paper, we propose a concurrent horizontal and vertical attention module to leverage width-wise and height-wise contextual priors markedly available in the panoramas. To yield semantic segmenters suitable for wide-FoV images, we present a multi-source omni-supervised learning scheme with panoramic domain covered in the training via data distillation. To facilitate the evaluation of contemporary CNNs in panoramic imagery, we put forward the Wild PAnoramic Semantic Segmentation (WildPASS) dataset, comprising images from all around the globe, as well as adverse and unconstrained scenes, which further reflects perception challenges of navigation applications in the real world. A comprehensive variety of experiments demonstrates that the proposed methods enable our high-efficiency architecture to attain significant accuracy gains, outperforming the state of the art in panoramic imagery domains.
Kailun Yang 0001, Xinxin Hu, Rainer Stiefelhagen
IEEE Trans. Image Process.2
2020 DS-PASS: Detail-Sensitive Panoramic Annular Semantic Segmentation through SwaftNet for Surrounding Sensing
abstract
Semantically interpreting the traffic scene is crucial for autonomous transportation and robotics systems. However, state-of-the-art semantic segmentation pipelines are dominantly designed to work with pinhole cameras and train with narrow Field-of-View (FoV) images. In this sense, the perception capacity is severely limited to offer higher-level confidence for upstream navigation tasks. In this paper, we propose a network adaptation framework to achieve Panoramic Annular Semantic Segmentation (PASS), which allows to re-use conventional pinhole-view image datasets, enabling modern segmentation networks to comfortably adapt to panoramic images. Specifically, we adapt our proposed SwaftNet to enhance the sensitivity to details by implementing attention-based lateral connections between the detail-critical encoder layers and the context-critical decoder layers. We benchmark the performance of efficient segmenters on panoramic segmentation with our extended PASS dataset, demonstrating that the proposed real-time SwaftNet outperforms state-of-the-art efficient networks. Furthermore, we assess real-world performance when deploying the Detail-Sensitive PASS (DS-PASS) system on a mobile robot and an instrumented vehicle, as well as the benefit of panoramic semantics for visual odometry, showing the robustness and potential to support diverse navigational applications.
Kailun Yang 0001, Xinxin Hu, Kaite Xiang, Kaiwei Wang, Rainer Stiefelhagen
IV2
2020 In Defense of Multi-Source Omni-Supervised Efficient ConvNet for Robust Semantic Segmentation in Heterogeneous Unseen Domains
abstract
Semantic segmentation renders a unified way of surrounding perception, where most of driving scene detection tasks can be covered by running a single efficient ConvNet through a forward pass. However, current frameworks posit the closed-world paradigm expressed as a single source of distribution over a predetermined set of visual classes, forgetting that a deep model must be deployed in the wild facing unseen domains and unforeseen hazards. In spite of being accurate in its comfort zone, the segmentation model may not generalize well to a new domain. In addition, a model trained with single dataset is heavily limited in terms of recognizable classes. In this paper, we propose an omni-supervised learning framework for semantic segmentation which is able to leverage heterogeneous data sources. Our omni-supervised training framework incorporates all available labeled and unlabeled data, meanwhile bridges multiple training sets to be capable of recognizing more classes that are needed for autonomous navigation application at hand in the new domain. A comprehensive variety of experiments shows that with the proposed multi-source omni-supervised learning solution, an efficient ConvNet like our ERF-PSPNet attains significant robustness gains in open domains that are of critical relevance to real deployment of vision algorithms. Our approach surpasses the state of the art on the highly unconstrained PASS and IDD20K datasets.
Kailun Yang 0001, Xinxin Hu, Kaiwei Wang, Rainer Stiefelhagen
IV2
2020 PASS: Panoramic Annular Semantic Segmentation
abstract
Pixel-wise semantic segmentation is capable of unifying most of driving scene perception tasks, and has enabled striking progress in the context of navigation assistance, where an entire surrounding sensing is vital. However, current mainstream semantic segmenters are predominantly benchmarked against datasets featuring narrow Field of View (FoV), and a large part of vision-based intelligent vehicles use only a forward-facing camera. In this paper, we propose a Panoramic Annular Semantic Segmentation (PASS) framework to perceive the whole surrounding based on a compact panoramic annular lens system and an online panorama unfolding process. To facilitate the training of PASS models, we leverage conventional FoV imaging datasets, bypassing the efforts entailed to create fully dense panoramic annotations. To consistently exploit the rich contextual cues in the unfolded panorama, we adapt our real-time ERF-PSPNet to predict semantically meaningful feature maps in different segments, and fuse them to fulfill panoramic scene parsing. The innovation lies in the network adaptation to enable smooth and seamless segmentation, combined with an extended set of heterogeneous data augmentations to attain robustness in panoramic imagery. A comprehensive variety of experiments demonstrates the effectiveness for real-world surrounding perception in a single PASS, while the adaptation proposal is exceptionally positive for state-of-the-art efficient networks.
Kailun Yang 0001, Xinxin Hu, Luis Miguel Bergasa, Eduardo Romera, Kaiwei Wang
IEEE Trans. Intell. Transp. Syst.2
2019 ACNET: Attention Based Network to Exploit Complementary Features for RGBD Semantic Segmentation
abstract
Compared to RGB semantic segmentation, RGBD semantic segmentation can achieve better performance by taking depth information into consideration. However, it is still problematic for contemporary segmenters to effectively exploit RGBD information since the feature distributions of RGB and depth (D) images vary significantly in different scenes. In this paper, we propose an Attention Complementary Network (ACNet) that selectively gathers features from RGB and depth branches. The main contributions lie in the Attention Complementary Module (ACM) and the architecture with three parallel branches. More precisely, ACM is a channel attention-based module that extracts weighted features from RGB and depth branches. The architecture preserves the inference of the original RGB and depth branches, and enables the fusion branch at the same time. Based on the above structures, ACNet is capable of exploiting more high-quality features from different channels. We evaluate our model on SUN-RGBD and NYUDv2 datasets, and prove that our model outperforms state-of-the-art methods. In particular, a mIoU score of 48.3% on NYUDv2 test set is achieved with ResNet50. We will release our source code based on PyTorch and the trained segmentation model at https://github.com/anheidelonghu/ACNet.
Xinxin Hu, Kailun Yang 0001, Lei Fei, Kaiwei Wang
ICIP1
2019 Can we PASS beyond the Field of View? Panoramic Annular Semantic Segmentation for Real-World Surrounding Perception
abstract
Pixel-wise semantic segmentation unifies distinct scene perception tasks in a coherent way, and has catalyzed notable progress in autonomous and assisted navigation, where a whole surrounding perception is vital. However, current mainstream semantic segmenters are normally benchmarked against datasets with narrow Field of View (FoV), and most vision-based navigation systems use only a forward-view camera. In this paper, we propose a Panoramic Annular Semantic Segmentation (PASS) framework to perceive the entire surrounding based on a compact panoramic annular lens system and an online panorama unfolding process. To facilitate the training of PASS models, we leverage conventional FoV imaging datasets, bypassing the effort entailed to create dense panoramic annotations. To consistently exploit the rich contextual cues in the unfolded panorama, we adapt our real-time ERF-PSPNet to predict semantically meaningful feature maps in different segments and fuse them to fulfill smooth and seamless panoramic scene parsing. Beyond the enlarged FoV,we extend focal length-related and style transfer-based data augmentations, to robustify the semantic segmenter against distortions and blurs in panoramic imagery. A comprehensive variety of experiments demonstrates the qualified robustness of our proposal for realworld surrounding understanding.
Kailun Yang 0001, Xinxin Hu, Luis Miguel Bergasa, Eduardo Romera, Dongming Sun, Kaiwei Wang
IV2