Xiao Teng

dblp:69/10238 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hyperspectral anomaly detection based on spatial-spectral feature fusion autoencoder
Zhimin Zhang 0005, Chengzhen Ma, Xiao Teng, Huansheng Ning, Lingfeng Mao 0001
Neurocomputing3
2026 YOLO_NMV: intelligent pedestrian detection for non-motorized vehicles
abstract
Non-motorized vehicles are widely used for short-distance travel and logistics, but distractions or illegal behavior often lead to accidents, posing a threat to pedestrians and drivers. To enhance the intelligence and safety of non-motorized vehicles, efficient pedestrian detection algorithms are needed. Based on the YOLOv8 architecture, the YOLO_NMV algorithm is proposed. The improvements include introducing StarNet_ADown in the backbone to enhance feature extraction. In the Neck layer, the C2f operation is improved using StarBlock, while LDTConv (Linear Deformable T-Convolutions) is employed in a different module to reduce channel redundancy and improve model robustness. Compared to the YOLOv8n model, YOLO_NMV reduces parameters, computation, and weight by 40.9%, 31.7%, and 38.1%, respectively, while improving precision, recall, mAP50, and mAP50-95 by 5.4%, 1.4%, 2.4%, and 2.3%. YOLO_NMV is more lightweight, accurate, and robust, making it suitable for deployment in non-motorized vehicles with limited computing power.
Xiao Teng, Zhengjiang Shen
Neural Comput. Appl.2
2026 Distilling structural knowledge from CNNs to vision transformers for data-efficient visual recognition
Dingyao Chen, Xiao Teng, Xun Yang 0001, Long Lan
Neural Networks2
2026 RefSAM: Efficiently adapting segmenting anything model for referring video object segmentation
Yonglin Li, Jing Zhang 0037, Xiao Teng, Xinwang Liu 0002, Long Lan
Neural Networks3
2026 Reliable Exploration Strategy for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification addresses the task of associating individuals across disjoint camera views in the absence of annotated training data. While current methods often focus on hard samples to learn discriminative features, these hard samples are more prone to label noise, which can negatively impact model performance. To address this issue, we propose an information-complementary hybrid contrastive learning framework that consists of two branches for extracting richer semantic information. Specifically, the first branch performs cluster-level contrastive learning to capture general semantic patterns, while the second branch employs relation-guided instance-level contrastive learning, which leverages local structures to mine hard samples for more discriminative representations. To mitigate label noise during hard sample mining, we introduce a temporal-guided dynamic weighting module that uses temporal clustering results to evaluate the reliability of the mined hard samples. This module generates instance-level weights to reduce the impact of unreliable samples, ensuring more stable training. Extensive experiments on four popular benchmarks demonstrate the superiority of our method compared to state-of-the-art approaches.
Xiao Teng, Long Lan
IEEE Signal Process. Lett.2
2026 Exploring Direction Alignment and Discrepancy Standardization for Knowledge Distillation
abstract
Knowledge Distillation (KD) is a widely popular model compression technique that can effectively transfer knowledge from a pre-trained, large-scale teacher model to a more compact and lightweight student model. Traditional KD methods aim to improve the student’s representation capability by mimicking the teacher’s features, e.g., minimizing the \(\mathcal{L}_{2}\) distance between their intermediate features. However, due to the capacity gap between the student and the teacher, student often struggles to precisely mimic the features of the teacher. To address this challenge, we propose to boost the knowledge distillation for the visual recognition tasks via Direction Alignment and Discrepancy Standardization ( DADS) , which exploits the feature scaling technique to distill from both the feature direction and feature discrepancy. To this end, we devise an efficient feature alignment module to align the dimensions of teacher and student features. Moreover, we align the direction of student features and teacher features, which are pre-processed by normalization. Furthermore, we leverage the Kullback–Leibler (KL) divergence to refine the features alignment, minimizing discrepancy in the distribution of features across samples, which is pre-processed by \(\mathcal{Z}\) -score standardization. In this way, our proposed approach can effectively transfer the knowledge from the teacher to the student, facilitating the downstream visual recognition applications, such as image classification and semantic segmentation. Extensive experimental analyses clearly validate the effectiveness of DADS . Compared with previous KD methods, our approach sets a new benchmark, achieving state-of-the-art results on visual recognition tasks.
Dingyao Chen, Xiao Teng, Xiang Zhang 0008, Xun Yang 0001, Long Lan
ACM Trans. Knowl. Discov. Data2
2025 Relieving Universal Label Noise for Unsupervised Visible-Infrared Person Re-Identification by Inferring from Neighbors
abstract
Unsupervised visible-infrared person re-identification (USL-VI-ReID) is of great research and practical significance yet remains challenging due to the absence of annotations. Existing approaches aim to learn modality-invariant representations in an unsupervised setting. However, these methods often encounter label noise within and across modalities due to suboptimal clustering results and considerable modality discrepancies, which impedes effective training. To address these challenges, we propose a straightforward yet effective solution for USL-VI-ReID by mitigating universal label noise using neighbor information. Specifically, we introduce the Neighbor-guided Universal Label Calibration (N-ULC) module, which replaces explicit hard pseudo labels in both homogeneous and heterogeneous spaces with soft labels derived from neighboring samples to reduce label noise. Additionally, we present the Neighbor-guided Dynamic Weighting (N-DW) module to enhance training stability by minimizing the influence of unreliable samples. Extensive experiments on the RegDB and SYSU-MM01 datasets demonstrate that our method outperforms existing USL-VI-ReID approaches, despite its simplicity.
Xiao Teng, Long Lan, Dingyao Chen, Kele Xu
AAAI1
2025 Coupling Category Alignment for Graph Domain Adaptation
abstract
Graph domain adaptation (GDA), which transfers knowledge from a labeled source domain to an unlabeled target graph domain, attracts considerable attention in numerous fields. However, existing methods commonly employ message-passing neural networks (MPNNs) to learn domain-invariant representations by aligning the entire domain distribution, inadvertently neglecting category-level distribution alignment and potentially causing category confusion. To address the problem, we propose an effective framework named Coupling Category Alignment (CoCA) for GDA, which effectively addresses the category alignment issue with theoretical guarantees. CoCA incorporates a graph convolutional network branch and a graph kernel network branch, which explore graph topology in implicit and explicit manners. To mitigate category-level domain shifts, we leverage knowledge from both branches, iteratively filtering highly reliable samples from the target domain using one branch and fine-tuning the other accordingly. Furthermore, with these reliable target domain samples, we incorporate the coupled branches into a holistic contrastive learning framework. This framework includes multi-view contrastive learning to ensure consistent representations across the dual branches, as well as cross-domain contrastive learning to achieve category-level domain consistency. Theoretically, we establish a sharper generalization bound, which ensures the effectiveness of category alignment. Extensive experiments on benchmark datasets validate the superiority of the proposed CoCA compared with baselines.
Xiao Teng, Zhiguang Cao, Mengzhu Wang
IJCAI2
2024 VinT-6D: A Large-Scale Object-in-hand Dataset from Vision, Touch and Proprioception
abstract
This paper addresses the scarcity of large-scale datasets for accurate object-in-hand pose estimation, which is crucial for robotic in-hand manipulation within the "Perception-Planning-Control" paradigm. Specifically, we introduce VinT-6D, the first extensive multi-modal dataset integrating vision, touch, and proprioception, to enhance robotic manipulation. VinT-6D comprises 2 million VinT-Sim and 0.1 million VinT-Real entries, collected via simulations in Mujoco and Blender and a custom-designed real-world platform. This dataset is tailored for robotic hands, offering models with whole-hand tactile perception and high-quality, well-aligned data. To the best of our knowledge, the VinT-Real is the largest considering the collection difficulties in the real-world environment so it can bridge the gap of simulation to real compared to the previous works. Built upon VinT-6D, we present a benchmark method that shows significant improvements in performance by fusing multi-modal information. The project is available at https://VinT-6D.github.io/.
Zhaoliang Wan, Yonggen Ling, Senlin Yi, Lu Qi 0001, Wang Wei Lee, Minglei Lu, Xiao Teng, Xu Yang 0004, Ming-Hsuan Yang 0001, Hui Cheng 0002
ICML8
2024 A Robust Model Predictive Controller for Tactile Servoing
abstract
Tactile servoing is an effective approach to enabling robots to safely interact with unknown environments. One of the core problems in tactile servoing is to robustly converge the contact features to the desired ones via a dedicated controller. This paper proposes a Data-Driven Model Predictive Controller (DDMPC) to compute the motion command given the previous interaction experience and feature deviations in tactile space. Compared with the manually designed PID-based controller, the proposed controller depends on the sound control theory and its convergence is guaranteed from a computational perspective. It is applied to the balancing control of a rolling bottle on a robotic forearm covered by a custom tactile sensor array. The real experiment demonstrates the superior robustness of the proposed approach and shows its great potential for other tactile servoing scenarios with measurement noise, which is inevitable for current tactile sensors.
Yihao Huang 0006, Wang Wei Lee, Tianliang Liu, Xiao Teng, Yu Zheng 0001, Qiang Li 0001
ICRA5
2024 A High-Performance Anthropomorphic Robotic Arm for Household Applications
abstract
Anthropomorphic robotic arms, mimicking the structure and function of human arms, show great potential for helping people in various tedious and repetitive household tasks. However, such arms mostly consist of multiple serial links controlled independently by actuators at joints with high reduction ratios, posing challenges in household services in terms of load capacity, responsiveness, and safety. In this paper, we propose a high-performance anthropomorphic arm called TRX-Arm based on differential cable transmission, characterized by features of high dynamics, high load capacity, and inherent compliance. TRX-Arm is composed of three deferential cable-driven coupling joints and one independent roll joint. Thanks to the cable differential transmission, the joints are capable of achieving doubled torque and stiffness without replacing motors. To enhance safety in human-robot interaction, the actuators including motors, reducer, belt, and pulley are mounted at the shoulder near the base and drive the joints remotely using cables, thereby minimizing the inertia of the whole arm. The workspace of TRX-Arm has a volume of 1.56 m3, much larger than that of the human arm. Real experiments show its capabilities including high repeatability and load capacity as well as high dynamic behavior of a dual-arm robot platform built with TRX-Arms.
Tianliang Liu, Jingchen Li 0001, Xiangchi Chen, Shuai Wang 0007, Xiao Teng, Wang Wei Lee, Xiong Li 0001, Yu Zheng 0001
IROS6
2024 Enhancing Unsupervised Visible-Infrared Person Re-Identification with Bidirectional-Consistency Gradual Matching
abstract
Unsupervised visible-infrared person re-identification (USL-VI-ReID) is of great research and practical significance yet remains challenging due to significant modality discrepancy and lack of annotations. Many existing approaches utilize variants of bipartite graph global matching algorithms to address this issue, aiming to establish cross-modality correspondences. However, these methods may encounter mismatches due to significant modality gaps and limited model representation. To mitigate this, we propose a simple yet effective framework for USL-VI-ReID, which gradually establishes associations between different modalities. To measure the confidence whether samples from different modalities belong to the same identity, we introduce a bidirectional-consistency criterion, which not only considers direct relationships between samples from different modalities but also incorporates potential hard negative samples from the same modality. Additionally, we propose a cross-modality correlation preserving module to further enhance the semantic representation of the model by maintaining consistency in correlations across modalities. Extensive experiments conducted on the public SYSU-MM01 and RegDB datasets demonstrate the superiority of our method over existing USL-VI-ReID approaches across various settings, despite the simplicity of our method.
Xiao Teng, Kele Xu, Long Lan
ACM Multimedia1
2024 Self-distillation Enhanced Vertical Wavelet Spatial Attention for Person Re-identification
Huibin Tan, Long Lan, Xiao Teng
MMM (2)4
2024 Instance-Level Scaling and Dynamic Margin-Alignment Knowledge Distillation
Xiao Teng, Zheng Qin 0002, Long Lan, Jing Zhang 0037
PRCV (11)2
2024 TIG-CL: Teacher-Guided Individual- and Group-Aware Contrastive Learning for Unsupervised Person Reidentification in Internet of Things
abstract
Unsupervised person reidentification (Re-ID) has attracted widespread due to its potential in Internet of Things applications, such as intelligent visual surveillance, it refers to retrieving the same individual across different camera views without using labeled data. To tackle the problem, a prevalent technique adopted by existing methods involves generating pseudo labels through clustering algorithms. However, this approach can result in merging individuals with different identities into the same group (i.e., cluster) during the training process. As a result, the resulting group centers may obscure the inherent characteristics of individual identities, thereby hindering the model from learning discriminative representations. To address the issue, we present a teacher-guided individual- and group-aware contrastive learning framework. Specifically, we propose a departure from the traditional approach of relying solely on contrastive learning between individual features and their corresponding group centers. Instead, we also exploit the relationship among individuals to construct contrast pairs and facilitate the learning of more discriminative features. This strategy enables the model to learn more about the individual characteristics that distinguish different persons, thus enhancing its ability to reidentify individuals accurately. Moreover, our method introduces a novel hybrid distillation module that enables simultaneous probability distillation at the group level and relationship distillation at the individual level. Guided by the teacher model, this module leads to improved feature representations of the student model. Extensive experimental results verify the effectiveness of our approach on four popular Re-ID data sets. The code will be made publicly available.
Xiao Teng, Xueqiong Li, Xinwang Liu 0002, Long Lan
IEEE Internet Things J.1
2024 Highly Efficient Active Learning With Tracklet-Aware Co-Cooperative Annotators for Person Re-Identification
abstract
Supervised person re-identification (ReID) has attracted widespread attentions in the computer vision community due to its great potential in real-world applications. However, the demand of human annotation heavily limits the application as it is costly to annotate identical pedestrians appearing from different cameras. Thus, how to reduce the annotation cost while preserving the performance remains challenging and has been studied extensively. In this article, we propose a tracklet-aware co-cooperative annotators' framework to reduce the demand of human annotation. Specifically, we partition the training samples into different clusters and associate adjacent images in each cluster to produce the robust tracklet which decreases the annotation requirements significantly. Besides, to further reduce the cost, we introduce a powerful teacher model in our framework to implement the active learning strategy and select the most informative tracklets for human annotator, the teacher model itself, in our setting, also acts as an annotator to label the relatively certain tracklets. Thus, our final model could be well-trained with both confident pseudo-labels and human-given annotations. Extensive experiments on three popular person ReID datasets demonstrate that our approach could achieve competitive performance compared with state-of-the-art methods in both active learning and unsupervised learning (USL) settings.
Xiao Teng, Long Lan, Xueqiong Li, Yuhua Tang
IEEE Trans. Neural Networks Learn. Syst.1
2023 Learning to Purification for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification is a challenging and promising task in computer vision. Nowadays unsupervised person re-identification methods have achieved great progress by training with pseudo labels. However, how to purify feature and label noise is less explicitly studied in the unsupervised manner. To purify the feature, we take into account two types of additional features from different local views to enrich the feature representation. The proposed multi-view features are carefully integrated into our cluster contrast learning to leverage more discriminative cues that the global feature easily ignored and biased. To purify the label noise, we propose to take advantage of the knowledge of teacher model in an offline scheme. Specifically, we first train a teacher model from noisy pseudo labels, and then use the teacher model to guide the learning of our student model. In our setting, the student model could converge fast with the supervision of the teacher model thus reduce the interference of noisy labels as the teacher model greatly suffered. After carefully handling the noise and bias in the feature learning, our purification modules are proven to be very effective for unsupervised person re-identification. Extensive experiments on two popular person re-identification datasets demonstrate the superiority of our method. Especially, our approach achieves a state-of-the-art accuracy 85.8% @mAP and 94.5% @Rank-1 on the challenging Market-1501 benchmark with ResNet-50 under the fully unsupervised setting. Code has been available at: https://github.com/tengxiao14/Purification_ReID.
Long Lan, Xiao Teng, Jing Zhang 0037, Xiang Zhang 0008, Dacheng Tao
IEEE Trans. Image Process.2
2022 Multi-scale local cues and hierarchical attention-based LSTM for stock price trend prediction
Xiao Teng, Xiang Zhang 0008, Zhigang Luo
Neurocomputing1