Xiaoyu Tian

dblp:185/6745 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
11since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Federated learning via multi-attention guided UNet for thyroid nodule segmentation of ultrasound images
Zhuo Xiang, Xiaoyu Tian, Yiyao Liu, Minsi Chen, Cheng Zhao 0003, Li-Na Tang, En-Sheng Xue, Hong-Yuan Xue, Ying-Jia Li, Quan-Shui Li, Chang-Jun Wu, Tian-Tian Ren, Jin-Yu Wu, Tianfu Wang 0001, Wen-Ying Liu, Bo-Ji Liu, Li-Ping Sun, Chong-Ke Zhao, Hui-Xiong Xu, Bai Ying Lei
Neural Networks2
2025 MAGS: Max-Gap Loss-Guided Siamese-Reconstruction Network for Hyperspectral Image Partial Label Learning
abstract
Due to the powerful feature extraction capabilities of deep learning, a series of deep learning-based methods for hyperspectral image (HSI) classification have been proposed and achieved satisfactory performance. However, most of these methods require a large number of labeled data, and the collection of completely accurate pixel-level labeled HSI data is difficult, resulting from the intricate label ambiguity of HSI and incomplete prior knowledge of annotators. Simultaneously, a few researchers focus on label ambiguity for HSI classification. Partial label learning (PLL) is one of the strategies to solve the problem where each training instance is assigned a candidate label set, among which only one is the ground truth label, which can essentially alleviate labeling difficulties. In this article, a max-gap loss-guided Siamese reconstruction network (MAGS) is proposed to combine PLL with HSI classification. MAGS consists of three components, including a spatial-spectral encoder, a spatial-spectral decoder, and a Siamese spatial-spectral encoder for high-quality feature representation learning to facilitate label disambiguation. In the encoding process, MAGS introduces the cross-attention and max-matching fusion strategies to obtain more representative features. In addition, to improve label disambiguation, the maximum gap loss is designed to guide the model training. Quantitative and qualitative results indicate that the MAGS outperforms several state-of-the-art methods on three HSI datasets. The code is available athttps://github.com/Nemo96yu/MAGS.
Xiaoyu Tian, Fulin Luo, Chuan Fu, Tan Guo, Bo Du 0001, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.1
2025 A Node Deployment Strategy in Solar Insecticidal Lamps Internet of Things With Respect to Partial Coverage and Energy Harvesting Requirements
abstract
Coverage is a fundamental issue in the Solar Insecticidal Lamps IoTs (SIL-IoTs). Compared to complete coverage, partial coverage emerges as the preferred strategy for deploying SILs within a limited budget, as this deployment solution offers the highest cost-effectiveness. In this paper, we concentrate on studying the constrained SILs deployment problem, taking into account partial coverage and energy harvesting requirements, which we refer to as the cSILDP-PCEH problem. In this context, the positions for deploying SILs are restricted to a weighted set of candidate locations on the ridges. The weight assigned to each candidate location reflects the energy harvesting potential of the SIL deployed at that position. Our objective is to deploy a group of SILs in a subset of these candidate locations, ensuring a high overall energy harvesting potential, network connectivity, and achieving partial coverage. Due to the NP-hard nature of the problem, we introduce an approximation algorithm with a provable performance ratio tailored to our problem. Finally, we conduct a theoretical analysis of our proposed algorithm and perform extensive simulations. The simulation results demonstrate that the proposed algorithm achieves a minimum improvement of 16.45% in energy harvesting potential while preserving network connectivity and maintaining a comparable coverage level.
Fan Yang 0067, Xiaoyu Tian, Zhaojun Zhang, Lei Shu 0001, Xiaoyuan Jing
IEEE Trans. Sustain. Comput.2
2024 CopperTag: A Real-Time Occlusion-Resilient Fiducial Marker
abstract
Fiducial markers, like AprilTag and ArUco, are extensively utilized in robotics applications within industrial environments, encompassing navigation, docking, and object grasping tasks. However, in contrast to controlled laboratory conditions, markers installed in factory grounds or equipment surfaces, often face challenges like damage or contamination. These issues can lead to compromised marker integrity, resulting in reduced detection reliability. To address this challenge, we propose a novel fiducial marker called CopperTag, which incorporates circular and square elements to create a robust occlusion-resistant pattern. The CopperTag detection process relies on three fundamental steps: firstly, extracting all lines from the image; secondly, identifying corners; and lastly, searching for quadrilateral candidate regions using ellipses and nearby corners. The Reed-Solomon (RS) algorithm is utilized for both encoding and decoding the information content. This algorithm possesses the ability to recover corrupted messages in situations where CopperTag data is incomplete. The experimental results illustrate that CopperTag exhibits superior robustness and accuracy in detection when compared to other state-of-the-art fiducial markers, even in scenarios with heavy occlusion. Moreover, CopperTag maintains an average processing time of 10ms per frame on a standard laptop, effectively meeting the real-time demands of robotics applications.
Xu Bian, Wenzhao Chen, Xiaoyu Tian, Donglai Ran
ICRA3
2024 A Label Information Aware Model for Multi-label Text Classification
abstract
Multi-label text classification (MLTC) refers to that each document is associated with more than one label at the same time, which attract much attention from researchers in both academia and industry. Existing methods have difficulties in determining label-related components from documents, which cannot effectively establish the association between textual features and label information. In fact, there are some label information, such as label semantic information and co-occurrence relations among labels, could be used to improve the performance on multi-label text classification. In this paper, we propose a label information aware model to utilize these information. Our model makes use of label semantic information to determine label-related components from textual features for obtaining the label-specific textual presentation for each sample, and then take advantages of co-occurrence relations among labels to construct interaction among label-specific textual presentation. The superiority of our model has been proved through comparing our method with several existing models on two datasets.
Xiaoyu Tian, Yongbin Qin, Ruizhang Huang, Yanping Chen 0010
Neural Process. Lett.1
2024 CrCD: Multidirection-MLP-Based Cross-Contrastive Disambiguation for Hyperspectral Image Partial Label Learning
Xiaoyu Tian, Fulin Luo, Xiuwen Gong, Tan Guo, Bo Du 0001, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.1
2023 GeoMAE: Masked Geometric Target Prediction for Self-supervised Point Cloud Pre-Training
abstract
This paper tries to address a fundamental question in point cloud self-supervised learning: what is a good signal we should leverage to learn features from point clouds without annotations? To answer that, we introduce a point cloud representation learning framework, based on geometric feature reconstruction. In contrast to recent papers that directly adopt masked autoencoder (MAE) and only predict original coordinates or occupancy from masked point clouds, our method revisits differences between images and point clouds and identifies three self-supervised learning objectives peculiar to point clouds, namely centroid prediction, normal estimation, and curvature prediction. Combined, these three objectives yield an nontrivial self-supervised learning task and mutually facilitate models to better reason fine-grained geometry of point clouds. Our pipeline is conceptually simple and it consists of two major steps: first, it randomly masks out groups of points, followed by a Transformer-based point cloud encoder; second, a lightweight Transformer decoder predicts centroid, normal, and curvature for points in each voxel. We transfer the pre-trained Transformer encoder to a downstream peception model. On the nuScene Datset, our model achieves 3.38 mAP improvment for object detection, 2.1 mIoU gain for segmentation, and 1.7 AMOTA gain for multi-object tracking. We also conduct experiments on the Waymo Open Dataset and achieve significant performance improvements over baselines as well.11Our code is available at https://github.com/Tsinghua-MARS-Lab/GeoMAE.
Xiaoyu Tian, Haoxi Ran, Yue Wang 0041, Hang Zhao 0021
CVPR1
2023 Occ3D: A Large-Scale 3D Occupancy Prediction Benchmark for Autonomous Driving
abstract
Robotic perception requires the modeling of both 3D geometry and semantics. Existing methods typically focus on estimating 3D bounding boxes, neglecting finer geometric details and struggling to handle general, out-of-vocabulary objects. 3D occupancy prediction, which estimates the detailed occupancy states and semantics of a scene, is an emerging task to overcome these limitations.To support 3D occupancy prediction, we develop a label generation pipeline that produces dense, visibility-aware labels for any given scene. This pipeline comprises three stages: voxel densification, occlusion reasoning, and image-guided voxel refinement. We establish two benchmarks, derived from the Waymo Open Dataset and the nuScenes Dataset, namely Occ3D-Waymo and Occ3D-nuScenes benchmarks. Furthermore, we provide an extensive analysis of the proposed dataset with various baseline models. Lastly, we propose a new model, dubbed Coarse-to-Fine Occupancy (CTF-Occ) network, which demonstrates superior performance on the Occ3D benchmarks.The code, data, and benchmarks are released at \url{https://tsinghua-mars-lab.github.io/Occ3D/}.
Xiaoyu Tian, Longfei Yun, Yucheng Mao, Huitong Yang, Hang Zhao 0021
NeurIPS1
2023 MedoidsFormer: A Strong 3D Object Detection Backbone by Exploiting Interaction With Adjacent Medoid Tokens
abstract
In this paper, we propose MedoidsFormer, a novel transformer-based backbone equipped with a self-attention mechanism that is tailored explicitly to LiDAR-based 3D object detection. Unlike 2D object detection, the proportion of target objects to the input scene is much smaller, and their distribution is significantly sparser in 3D object detection. Given these observations, we introduce a new self-attention mechanism called Medoids Attention, focusing on exploiting interactions within surrounding regions, which not only reduces computation and memory costs but obtains discriminative context information. Instead of aggregating tokens from adjacent areas, we present a dynamic semantic-aware token mining process through k-Medoids clustering to direct select representative tokens for attention modeling. Our proposed method shows consistent improvement over existing 3D object detectors through extensive experiments and achieves state-of-the-art performance on the large-scale Waymo Open Dataset. We also conduct comprehensive ablation studies to verify the efficacy of the new self-attention mechanism and provide thorough insights.
Xiaoyu Tian, Qian Yu 0002, Jun-Hai Yong, Dong Xu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Overview of the NLPCC2022 Shared Task on Speech Entity Linking
Ruoyu Song 0002, Xiaoyu Tian, Yuhang Guo 0001
NLPCC (2)3
2021 Unsupervised Learning of 3D Scene Flow from Monocular Camera*
abstract
Scene flow represents the motion of points in the 3D space, which is the counterpart of the optical flow that represents the motion of pixels in the 2D image. However, it is difficult to obtain the ground truth of scene flow in the real scenes, and recent studies are based on synthetic data for training. Therefore, how to train a scene flow network with unsupervised methods based on real-world data shows crucial significance. A novel unsupervised learning method for scene flow is proposed in this paper, which utilizes the images of two consecutive frames taken by monocular camera without the ground truth of scene flow for training. Our method realizes the goal that training scene flow network with real-world data, which bridges the gap between training data and test data and broadens the scope of available data for training. Unsupervised learning of scene flow in this paper mainly consists of two parts: (i) depth estimation and camera pose estimation, and (ii) scene flow estimation based on four different loss functions. Depth estimation and camera pose estimation obtain the depth maps and camera pose between two consecutive frames, which provide further information for the next scene flow estimation. After that, we used depth consistency loss, dynamic-static consistency loss, Chamfer loss, and Laplacian regularization loss to carry out unsupervised training of the scene flow network. To our knowledge, this is the first paper that realizes the unsupervised learning of 3D scene flow from monocular camera. The experiment results on KITTI show that our method for unsupervised learning of scene flow meets great performance compared to traditional methods Iterative Closest Point (ICP) and Fast Global Registration (FGR). The source code is available at: https://github.com/IRMVLab/3DUnMonoFlow.
Guangming Wang 0001, Xiaoyu Tian, Ruiqi Ding, Hesheng Wang 0001
ICRA2