Xianlin Zhang

dblp:202/0480 · DBLP profile ↗
← Back
16ranked-venue papers
4as first author
14since 2021 · last 2026
0000-0003-3905-2062ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RUCLIP: Robust concept unlearning in CLIP via semantic anchors
Yue Zhang 0016, Qinghong Yin, Xianlin Zhang, Xueming Li 0002
Expert Syst. Appl.4
2026 CRColor: Cycle reference learning for exemplar-based image colorization
Mingdao Wang, Xianlin Zhang, Xueming Li 0002, Yue Zhang 0016
Neurocomputing3
2026 Chain-of-Evidence Multimodal Reasoning for Few-Shot Temporal Action Localization
Mengshi Qi, Hongwei Ji, Wulian Yun, Xianlin Zhang, Huadong Ma
IEEE Trans. Image Process.4
2026 Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning
abstract
Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challenging in real-world scenarios. However, current video understanding methods are mainly concerned with what and where the action is, which is unable to meet the requirements. Meanwhile, most of the existing datasets lack the labels indicating the degree of action standardization, and the action quality assessment datasets lack explainability and detailed feedback. Therefore, we define a new Human Action Form Assessment (AFA) task, and introduce a new diverse dataset CoT-AFA, which contains a large scale of fitness and martial arts videos with multi-level annotations for comprehensive video analysis. We enrich the CoT-AFA dataset with a novel Chain-of-Thought explanation paradigm. Instead of offering isolated feedback, our explanations provide a complete reasoning process-from identifying an action step to analyzing its outcome and proposing a concrete solution. Furthermore, we propose a framework named Explainable Fitness Assessor, which can not only judge an action but also explain why and provide a solution. This framework employs two parallel processing streams and a dynamic gating mechanism to fuse visual and semantic information, thereby boosting its analytical capabilities. The experimental results demonstrate that our method has achieved improvements in explanation generation (e.g., + 16.0% in CIDEr),action classification (+ 2.7% in accuracy) and quality assessment (+ 2.1% in accuracy), revealing great potential of CoT-AFA for future studies. Our dataset and source code are available at https://github.com/MICLAB-BUPT/EFA.
Mengshi Qi, Yeteng Wu, Wulian Yun, Xianlin Zhang, Huadong Ma
IEEE Trans. Image Process.4
2025 Sketch-1-to-3: One Single Sketch to 3D Detailed Face Reconstruction
abstract
3D face reconstruction from a single sketch is a critical yet underexplored task with significant practical applications. The primary challenges stem from the substantial modality gap between 2D sketches and 3D facial structures, including: (1) accurately extracting facial keypoints from 2D sketches; (2) preserving diverse facial expressions and fine-grained texture details; and (3) training a high-performing model with limited data. In this paper, we propose Sketch-1-to-3, a novel framework for realistic 3D face reconstruction from a single sketch, to address these challenges. Specifically, we first introduce the Geometric Contour and Texture Detail (GCTD) module, which enhances the extraction of geometric contours and texture details from facial sketches. Additionally, we design a deep learning architecture with a domain adaptation module and a tailored loss function to align sketches with the 3D facial space, enabling high-fidelity expression and texture reconstruction. To facilitate evaluation and further research, we construct SketchFaces, a real hand-drawn facial sketch dataset, and Syn-SketchFaces, a synthetic facial sketch dataset. Extensive experiments demonstrate that Sketch-1-to-3 achieves state-of-the-art performance in sketch-based 3D face reconstruction.
Liting Wen, Zimo Yang, Xianlin Zhang, Chi Ding, Mingdao Wang, Xueming Li 0002
MMAsia3
2025 Spcolor: Semantic prior guided exemplar-based image colorization
Xianlin Zhang, Mingdao Wang, Xueming Li 0002, Yue Zhang 0016
Pattern Recognit.2
2024 Modeling the skeleton-language uncertainty for 3D action recognition
Mingdao Wang, Xianlin Zhang, Xueming Li 0002, Yue Zhang 0016
Neurocomputing2
2024 Exemplar-based video colorization with long-term spatiotemporal dependency
Xueming Li 0002, Xianlin Zhang, Mingdao Wang, Jiatong Han, Yue Zhang 0016
Knowl. Based Syst.3
2024 Learning Representations by Contrastive Spatio-Temporal Clustering for Skeleton-Based Action Recognition
abstract
Self-supervised representation learning has proven constructive for skeleton-based action recognition. For better performance, existing methods mainly focus on 1) multi-modal data augmentations and 2) triplet contrastive samples construction. However, designing these strategies is always heuristics and hard. Instead of exploring more similar strategies, this paper addresses this issue with a different view and proposes a novel Contrastive Spatio-Temporal Clustering (CSTC) module. CSTC constructs a supervised signal (pseudo-label) of action sequences in an online clustering manner, and it is complementary to the recent data augmentations or triplet contrastive samples construction strategies. Specifically, CSTC can be formulated as an optimal transport problem. we introduce the spatio-temporal regularizations into the original optimal transport term to guide the pseudo-label generation, i.e., a semantic regularization learned by frame index is proposed to constrain the frame order, and a prior normal distribution regularization based on sampling characteristics of samples is proposed to maintain the dependability of spatial cluster assignments. Furthermore, to enhance the learning of latent features, we propose a Bidirectional Cross-modal Clustering Consistency Objective (B3CO) to enforce cluster assignments consistency for different modalities of the same sample. Last, since fusing spatial and temporal clustering losses directly during back-propagation will confuse the learned dimension-specific semantics, we propose a simple yet effective training strategy to fix it by training the model using these two losses alternately. By integrating the above designs into the MoCo framework, we propose a Contrastive Spatio-Temporal Clustering Network (CSTCN), which can excavate cross-modal discriminative spatio-temporal features in the clustering space. Experimental results on NTU RGB+D 60, NTU RGB+D 120, and PKU-MMD II datasets show that CSTCN achieves state-of-the-art performance in both single- and multi-modal models, especially in the KNN and semi-supervised evaluation protocols. Besides, the key module CSTC shows good generalization capability, and achieves consistent performance improvement on the basis of several state-of-the-art methods which focus on data augmentations and triplet contrastive samples construction.
Mingdao Wang, Xueming Li 0002, Xianlin Zhang, Lei Ma 0003, Yue Zhang 0016
IEEE Trans. Multim.4
2023 AABLSTM: A Novel Multi-task Based CNN-RNN Deep Model for Fashion Analysis
abstract
With the rapid growth of online commerce and fashion-related applications, visual clothing analysis and recognition has become a hotspot in computer vision. In this paper, we propose a novel AABLSTM network, which is based on deep CNN-RNN, to solve the visual fashion analysis of clothing category classification, attribute detection, and landmark localization. The designed fashion model is leveraged with the multi-task driven mechanism as follows: firstly, a bidirectional LSTM (Bi-LSTM) branch is proposed for efficiently mining the semantic association between related attributes so as to improve the precision of clothing category classification and attribute detection; then, an imitated hourglass sub-network of “down-up sampling” is constructed for boosting the accuracy of fashion landmark localization; and finally, a specially designed multi-loss function is constructed to better optimize the network training. Extensive experimental results on large-scale fashion datasets demonstrate the superior performance of our approach.
Xianlin Zhang, Mengling Shen, Xueming Li 0002, Xiaojie Wang 0006
ACM Trans. Multim. Comput. Commun. Appl.1
2022 TIR: A Two-Stage Insect Recognition Method for Convolutional Neural Network
Yunqi Feng 0002, Yang Liu 0132, Xianlin Zhang, Xueming Li 0002
PRCV (2)3
2022 Hierarchical graph attention network with pseudo-metapath for skeleton-based action recognition
Mingdao Wang, Xueming Li 0002, Xianlin Zhang, Yue Zhang 0016
Neurocomputing3
2022 A deformable CNN-based triplet model for fine-grained sketch-based image retrieval
Xianlin Zhang, Mengling Shen, Xueming Li 0002, Fangxiang Feng
Pattern Recognit.1
2021 A method of inpainting moles and acne on the high-resolution face photos
abstract
Abstract With the rapid development of mobile phones, more and more high‐resolution photos are taken. The demand for high‐resolution image inpainting is becoming increasingly urgent. In order to repair high‐resolution face images automatically and quickly, this paper proposes an improved generative adversarial networks method. Firstly, we made a high‐resolution dataset for training and testing, and abandoned the traditional 256256 size data. Secondly, since the existing methods can only repair the mask with fixed size and shape on the image, when the global average pooling layer is used in the network, the improved network can repair the moles and acne with arbitrary sizes and shapes on the human face photos. Finally, in order to achieve optimal performance of the network, a mixed loss function is used in training. The experimental results prove that our method has not only achieved good results in qualitative results, but also achieved excellent results in quantitative results.
Xuewei Li 0005, Xueming Li 0002, Xianlin Zhang, Yang Liu 0132, Jiayi Liang, Ziliang Guo, Keyu Zhai
IET Image Process.3
2019 A survey on freehand sketch recognition and retrieval
Xianlin Zhang, Xueming Li 0002, Yang Liu 0132, Fangxiang Feng
Image Vis. Comput.1
2018 Better freehand sketch synthesis for sketch-based image retrieval: Beyond image edges
Xianlin Zhang, Xueming Li 0002, Xuewei Li 0005, Mengling Shen
Neurocomputing1