Xiaotian Yin

dblp:97/3083 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
8since 2021 · last 2026
0009-0005-4557-2138ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 Adaptive Agent Selection and Interaction Network for Image-to-Point Cloud Registration
abstract
Typical detection-free methods for image-to-point cloud registration leverage transformer-based architectures to aggregate cross-modal features and establish correspondences. However, they often struggle under challenging conditions, where noise disrupts similarity computation and leads to incorrect correspondences. Moreover, without dedicated designs, it remains difficult to effectively select informative and correlated representations across modalities, thereby limiting the robustness and accuracy of registration. To address these challenges, we propose a novel cross-modal registration framework composed of two key modules: the Iterative Agents Selection (IAS) module and the Reliable Agents Interaction (RAI) module. IAS enhances structural feature awareness with phase maps and employs reinforcement learning principles to efficiently select reliable agents. RAI then leverages these selected agents to guide cross-modal interactions, effectively reducing mismatches and improving overall robustness. Extensive experiments on the RGB-D Scenes v2 and 7-Scenes benchmarks demonstrate that our method consistently achieves state-of-the-art performance.
Zhixin Cheng, Xiaotian Yin, Jiacheng Deng 0002, Bohao Liao, Baoqun Yin, Tianzhu Zhang 0001
AAAI2
2026 GLASS: Geometry-Aware Local Alignment and Structure Synchronization Network for 2D-3D Registration
abstract
Image-to-point cloud registration methods typically follow a coarse-to-fine pipeline, extracting patch-level correspondences and refining them into dense pixel-to-point matches. However, in scenes with repetitive patterns, images often lack sufficient 3D structural cues and alignment with point clouds, leading to incorrect matches. Moreover, prior methods usually overlook structural consistency, limiting the full exploitation of correspondences. To address these issues, we propose two novel modules: the Local Geometry Enhancement (LGE) module and the Graph Distribution Consistency (GDC) module. LGE enhances both image and point cloud features with normal vectors, injecting geometric structure into image features to reduce mismatches. GDC constructs a graph from matched points to update features and explicitly constrain similarity distributions. Extensive experiments and ablations on two benchmarks, RGB-D Scenes v2 and 7-Scenes, demonstrate that our approach achieves state-of-the-art performance in image-to-point cloud registration.
Zhixin Cheng, Jiacheng Deng 0002, Xinjun Li, Bohao Liao, Li Liu 0067, Xiaotian Yin, Baoqun Yin, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Exploring the Better Multimodal Synergy Strategy for Vision-Language Models
abstract
Vision-Language models (VLMs) have shown great potential in enhancing open-world visual concept comprehension. Recent researches focus on an optimum multimodal collaboration strategy that significantly advances CLIP-based few-shot tasks. However, existing prompt-based solutions suffer from unidirectional information flow and increased parameters since they explicitly condition the vision prompts on textual prompts across different transformer layers using non-shareable coupling functions. To address this issue, we propose a Dual-shared mechanism based on LoRA (DsRA) that addresses VLM adaptation in low-data regimes. The proposed DsRA enjoys several merits. First, we design an inter-modal shared coefficient that focuses on capturing visual and textual shared patterns, ensuring effective mutual synergy between image and text features. Second, an intra-modal shared matrix is proposed to achieve efficient parameter fine-tuning by combining the different coefficients to generate layer-wise adapters placed in encoder layers. Our extensive experiments demonstrate that DsRA improves the generalizability under few-shot classification, base-to-new generalization, and domain generalization settings. Our code will be released soon.
Xiaotian Yin, Xin Liu 0089, Yuan Wang 0064, Yuwen Pan, Tianzhu Zhang 0001
AAAI1
2025 CA-I2P: Channel-Adaptive Registration Network with Global Optimal Selection
abstract
Detection-free methods typically follow a coarse-to-fine pipeline, extracting image and point cloud features for patch-level matching and refining dense pixel-to-point correspondences. However, differences in feature channel attention between images and point clouds may lead to degraded matching results, ultimately impairing registration accuracy. Furthermore, similar structures in the scene could lead to redundant correspondences in cross-modal matching. To address these issues, we propose Channel Adaptive Adjustment Module (CAA) and Global Optimal Selection Module (GOS). CAA enhances intra-modal features and suppresses cross-modal sensitivity, while GOS replaces local selection with global optimization. Experiments on RGB-D Scenes V2 and 7-Scenes demonstrate the superiority of our method, achieving state-of-the-art performance in image-to-point cloud registration.
Zhixin Cheng, Jiacheng Deng 0002, Xinjun Li, Xiaotian Yin, Bohao Liao, Baoqun Yin, Wenfei Yang, Tianzhu Zhang 0001
ICCV4
2025 PAF: Prototype Adaptive Fusion for Test-Time Adaptation of Vision-Language Models
abstract
Leveraging Vision-Language Models (VLMs) like CLIP for various downstream tasks has emerged as a significant research trend. Recently, researchers have introduced Test-Time Adaptation (TTA) as a technique for models to learn online from unlabeled samples at test time, improving the generalization performance of VLMs to target domains. However, existing TTA methods either require expensive backpropagating gradient computations for each test sample or only extract knowledge from a limited number of historical test samples in the cache model, resulting in suboptimal adaptation performance. To address these limitations, we propose a Prototype Adaptive Fusion (PAF) framework, a novel TTA approach that makes full use of historical knowledge from test samples. Unlike traditional cache-based methods, which store only a few low-entropy samples per class, PAF introduces a prototype fusion mechanism that constructs class prototype representations through cumulatively merging features from qualified test samples. Furthermore, we propose an enhanced version, Easy-Hard PAF (EH-PAF), which adaptively applies a category-specific strategy based on CLIP prediction to improve performance. Extensive experiments across 15 diverse datasets demonstrate that our method consistently outperforms previous state-of-the-art approaches.
Xiaotian Yin, Xin Liu 0089, Huakai Lai, Tianzhu Zhang 0001
ACM Multimedia3
2024 Task-Adaptive Prompted Transformer for Cross-Domain Few-Shot Learning
abstract
Cross-Domain Few-Shot Learning (CD-FSL) aims at recognizing samples in novel classes from unseen domains that are vastly different from training classes, with few labeled samples. However, the large domain gap between training and novel classes makes previous FSL methods perform poorly. To address this issue, we propose MetaPrompt, a Task-adaptive Prompted Transformer model for CD-FSL, by jointly exploiting prompt learning and the parameter generation framework. The proposed MetaPrompt enjoys several merits. First, a task-conditioned prompt generator is established upon attention mechanisms. It can flexibly produce a task-adaptive prompt with arbitrary length for unseen tasks, by selectively gathering task characteristics from the contextualized support embeddings. Second, the task-adaptive prompt is attached to Vision Transformer to facilitate fast task adaptation, steering the task-agnostic representation to incorporate task knowledge. To our best knowledge, this is the first work to exploit a prompt-based parameter generation mechanism for CD-FSL. Extensive experimental results on the Meta-Dataset benchmark demonstrate that our method achieves superior results against state-of-the-art methods.
Xin Liu 0089, Xiaotian Yin, Tianzhu Zhang 0001, Yongdong Zhang 0001
AAAI3
2024 Hierarchy-Aware Interactive Prompt Learning for Few-Shot Classification
abstract
Few-Shot Learning (FSL) leverages prior knowledge and generalization strategies to quickly adapt to new tasks or recognize new objects with minimal input. Recently, CLIP-based methods, aided by contrastive language-image pre-training, have demonstrated impressive few-shot performance. However, these methods solely employ fixed-length uni-modal prompts at the initial encoder layer, neglecting the multi-level adaptation and cross-modal interaction for the intermediate features. To address this issue, we propose Hierarchy-Aware Interactive Prompt Learning (HIPL), by jointly exploring hierarchical prompt learning and cross-modal prompt interaction for CLIP-based FSC. The proposed HIPL enjoys several merits. First, we design a hierarchical prompt aggregation module to progressively generate higher-level prompts via the attention mechanisms, equipping the CLIP with hierarchical adaptation capability. Second, a cross-modal prompt interaction module is proposed to facilitate deep interaction between stage-wise prompts, ensuring mutual synergy between vision and textual features. To the best of our knowledge, this is the first work to learn multi-level prompts by progressive aggregation. Our extensive experiments demonstrate that HIPL outperforms previous methods in few-shot classification and base-to-new generalization. Our code is available athttps://github.com/Yxt1212/HIPL
Xiaotian Yin, Wenfei Yang, Tianzhu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 A Discriminative Multi-task Learning for Autism Classification Based on Speech Signals
abstract
About 70 million people around the world are suffering from autism, which is about one in every 160 children. The causes of autism are complex, and there is no specific drug treatment. However, for an individual with autism, the earlier the age of treatment, the greater the improvement. In this paper, we collected an autism speech dataset and conducted a study on speech feature classification of autistic and normal children. We built a deep neural network, using Convolutional Recurrent Neural Network as the front-end encoder, and added Convolutional Block Attention Module to it. Integrate local features using recurrent neural networks. To prevent overfitting, we add Connectionist Temporal Classification based speech recognition auxiliary task during training. After introducing the loss function in the field of face recognition, the best classification accuracy reached 94.76%.
Xiaotian Yin, Chao Zhang 0047
ISCC1
2016 Inter-calibration of satellite passive microwave land observations from MWRI and AMSR-E over the bare soil and grassland
abstract
Evaluating and calibrating MWRI Tb record based on AMSR-E Tb record are very meaningful since these two sensors have similar configurations and the majority of geophysical algorithms have been developed based on AMSR-E Tb record. While previous studies are focus on the ocean and polar region ice sheet, this study aimed to evaluate and calibrate MWRI Tb record over bare soil and grassland and analyze the effects of these two land cover types on Tb record quality. The results show that the quality of MWRI Tb record over bare soil performance better than that over grassland and the quality at different frequency over bare soil is more stable than that over grassland. In summary, grassland had more significantly effects on MWRI Tb record than bare soil.
Xiaoran Lv, Adu Gong, Xiaotian Yin, Jing Li 0018, Jingmei Wang
IGARSS3
2015 Decentralized human trajectories tracking using hodge decomposition in sensor networks
abstract
With the recent development of localization and tracking systems for both indoor and outdoor settings, we consider the problem of analyzing and representing the huge amount of natural trajectories from human movements that we expect to gather in the near future. In this paper we argue the topological representation, which records how a target moves around the natural obstacles in the underlying environment, can be sufficiently descriptive for many applications and efficient enough for both storing, comparing and classifying these natural human trajectories. Technically, the representation uses the homotopy type of the trajectory. By using harmonic one-forms and Hodge decomposition, we pre-process the sensor network with a purely decentralized algorithm such that the homology class of a trajectory can be obtained by a simple integration along the trajectory. This supports real-time classification of trajectories up to the homology accuracy with minimum communication cost. We test the effectiveness of our approach by showing how to classify randomly generated trajectories in a multi-level arts museum layout as well as how to distinguish real world taxi trajectories in a large city.
Xiaotian Yin, Chien-Chun Ni, Jiaxin Ding 0001, Dengpan Zhou, Jie Gao 0001, Xianfeng Gu
SIGSPATIAL/GIS1
2012 Scalable routing in 3D high genus sensor networks using graph embedding
abstract
We study scalable routing for a sensor network deployed in complicated 3D settings such as underground tunnels in gas system or water system. The nodes are in general 3D space but they are very sparsely located and the network has complex topology. We propose a routing scheme by first embdding the network on a surface with possibly non-zero genus. Then we compute a canonical hyperbolic metric of the embedded surface, and use geodesics to decompose the network into canonical components called pairs of `pants' whose topology is simpler (with genus zero). The adjacency of the pants components is extracted as a high level routing map and stored at every node. With the hyperbolic metric one can use greedy routing to navigate within and across pants. Altogether this leads to a two-level routing scheme by first finding a sequence of pants and then realizing the route with greedy steps. We show by simulation that the number of pants is closely related to the true `genus' of the network and that the routing scheme is efficient and scalable.
Xiaokang Yu, Xiaotian Yin, Jie Gao 0001, Xianfeng Gu
INFOCOM2
2012 Brain Surface Conformal Parameterization With the Ricci Flow
abstract
In brain mapping research, parameterized 3-D surface models are of great interest for statistical comparisons of anatomy, surface-based registration, and signal processing. Here, we introduce the theories of continuous and discrete surface Ricci flow, which can create Riemannian metrics on surfaces with arbitrary topologies with user-defined Gaussian curvatures. The resulting conformal parameterizations have no singularities and they are intrinsic and stable. First, we convert a cortical surface model into a multiple boundary surface by cutting along selected anatomical landmark curves. Secondly, we conformally parameterize each cortical surface to a parameter domain with a user-designed Gaussian curvature arrangement. In the parameter domain, a shape index based on conformal invariants is computed, and inter-subject cortical surface matching is performed by solving a constrained harmonic map. We illustrate various target curvature arrangements and demonstrate the stability of the method using longitudinal data. To map statistical differences in cortical morphometry, we studied brain asymmetry in 14 healthy control subjects. We used a manifold version of Hotelling's T(2) test, applied to the Jacobian matrices of the surface parameterizations. A permutation test, along with the cumulative distribution of p-values, were used to estimate the overall statistical significance of differences. The results show our algorithm's power to detect subtle group differences in cortical surfaces.
Yalin Wang 0001, Jie Shi 0001, Xiaotian Yin, Xianfeng Gu, Tony F. Chan, Shing-Tung Yau, Arthur W. Toga, Paul M. Thompson
IEEE Trans. Medical Imaging3
2011 Deterministic greedy routing with guaranteed delivery in 3D wireless sensor networks
abstract
With both computational complexity and storage space bounded by a small constant, greedy routing is recognized as an appealing approach to support scalable routing in wireless sensor networks. However, significant challenges have been encountered in extending greedy routing from 2D to 3D space. In this research we develop decentralized solutions to achieve greedy routing in 3D sensor networks. Our proposed approach is based on a unit tetrahedron cell (UTC) mesh structure. We propose a distributed algorithm to realize volumetric harmonic mapping of the UTC mesh under spherical boundary condition. It is a one-to-one map that yields virtual coordinates for each node in the network. Since a boundary has been mapped to a sphere, node-based greedy routing is always successful thereon. At the same time, we exploit the UTC mesh to develop a face-based greedy routing algorithm, and prove its success at internal nodes. To deliver a data packet to its destination, face-based and node-based greedy routing algorithms are employed alternately at internal and boundary UTCs, respectively. As far as we know, this is the first work that realizes truly deterministic greedy routing with constant-bounded storage and computation in 3D wireless sensor networks.
Su Xia, Xiaotian Yin, Hongyi Wu, Miao Jin, Xianfeng Gu
MobiHoc2
2011 Computing shortest words via shortest loops on hyperbolic surfaces
Xiaotian Yin, Feng Luo 0002, Xianfeng Gu, Shing-Tung Yau
Comput. Aided Des.1
2011 GPU-Assisted Computation of Centroidal Voronoi Tessellation
abstract
Centroidal Voronoi tessellations (CVT) are widely used in computational science and engineering. The most commonly used method is Lloyd's method, and recently the L-BFGS method is shown to be faster than Lloyd's method for computing the CVT. However, these methods run on the CPU and are still too slow for many practical applications. We present techniques to implement these methods on the GPU for computing the CVT on 2D planes and on surfaces, and demonstrate significant speedup of these GPU-based methods over their CPU counterparts. For CVT computation on a surface, we use a geometry image stored in the GPU to represent the surface for computing the Voronoi diagram on it. In our implementation a new technique is proposed for parallel regional reduction on the GPU for evaluating integrals over Voronoi cells.
Guodong Rong 0001, Yang Liu 0014, Wenping Wang 0001, Xiaotian Yin, Xianfeng Gu, Xiaohu Guo
IEEE Trans. Vis. Comput. Graph.4
2010 Direct-Product Volumetric Parameterization of Handlebodies via Harmonic Fields
abstract
Volumetric parameterization plays an important role for geometric modeling. Due to the complicated topological nature of volumes, it is much more challenging than the surface case. This work focuses on the parameterization of volumes with a boundary surface embedded in 3D space. The intuition is to decompose the volume as the direct product of a two dimensional surface and a one dimensional curve. We first partition the boundary surface into ceiling, floor and walls. Then we compute the harmonic field in the volume with a Dirichlet boundary condition. By tracing the integral curve along the gradient of the harmonic function, we can parameterize the volume to the parametric domain. The method is guaranteed to produce bijection for handle bodies with complex topology, including topological balls as a degenerate case. Furthermore, the parameterization is regular everywhere. We apply the proposed parameterization method to construct hexahedral mesh.
Jiazhi Xia, Ying He 0001, Xiaotian Yin, Shuchu Han, Xianfeng Gu
Shape Modeling International3
2009 Greedy routing with guaranteed delivery using Ricci flows
Rik Sarkar, Xiaotian Yin, Jie Gao 0001, Feng Luo 0002, Xianfeng Gu
IPSN2
2009 Generalized Koebe's method for conformal mapping multiply connected domains
abstract
Surface parameterization refers to the process of mapping the surface to canonical planar domains, which plays crucial roles in texture mapping and shape analysis purposes. Most existing techniques focus on simply connected surfaces. It is a challenging problem for multiply connected genus zero surfaces. This work generalizes conventional Koebe's method for multiply connected planar domains. According to Koebe's uniformization theory, all genus zero multiply connected surfaces can be mapped to a planar disk with multiply circular holes. Furthermore, this kind of mappings are angle preserving and differ by Möbius transformations. We introduce a practical algorithm to explicitly construct such a circular conformal mapping. Our algorithm pipeline is as follows: suppose the input surface has n boundaries, first we choose 2 boundaries, and fill the other n -- 2 boundaries to get a topological annulus; then we apply discrete Yamabe flow method to conformally map the topological annulus to a planar annulus; then we remove the filled patches to get a planar multiply connected domain. We repeat this step for the planar domain iteratively. The two chosen boundaries differ from step to step. The iterative construction leads to the desired conformal mapping, such that all the boundaries are mapped to circles. In theory, this method converges quadratically faster than conventional Koebe's method. We give theoretic proof and estimation for the converging rate. In practice, it is much more robust and efficient than conventional non-linear methods based on curvature flow. Experimental results demonstrate the robustness and efficiency of the method.
Wei Zeng 0002, Xiaotian Yin, Min Zhang 0069, Feng Luo 0002, Xianfeng Gu
Symposium on Solid and Physical Modeling2
2008 3D Non-rigid Surface Matching and Registration Based on Holomorphic Differentials
Wei Zeng 0002, Yang Wang 0001, Xiaotian Yin, Xianfeng Gu, Dimitris Samaras
ECCV (3)4
2008 Slit Map: Conformal Parameterization for Multiply Connected Surfaces
Xiaotian Yin, Junfei Dai, Shing-Tung Yau, Xianfeng Gu
GMP1
2007 Computing Shortest Cycles Using Universal Covering Space
abstract
Summary form only given. In this paper we generalize the shortest path algorithm to the shortest cycles in each homotopy class on a surface with arbitrary topology, utilizing the universal covering space (UCS) in algebraic topology. In order to store and handle the UCS, we propose a two-level data structure which is efficient for storage and easy to process. We also pointed several practical applications for our shortest cycle algorithms and the UCS data structure.
Xiaotian Yin, Miao Jin, Xianfeng Gu
CAD/Graphics1
2007 Focal surfaces of discrete geometry
Jingyi Yu 0001, Xiaotian Yin, Xianfeng Gu, Leonard McMillan, Steven J. Gortler
Symposium on Geometry Processing2
2007 Computing shortest cycles using universal covering space
Xiaotian Yin, Miao Jin, Xianfeng Gu
Vis. Comput.1