Xiang Tian 0002

dblp:16/318-2 · DBLP profile ↗
← Back
23ranked-venue papers
0as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DDPT: Enhancing complex reasoning in large language models via distillation and dynamic prompt tuning
Ge Teng, Chen Shen 0003, Wenxiao Wang 0001, Sinan Fan, Liang Xie 0003, Xiang Tian 0002, Peng Chen 0008, Yaowu Chen, Jieping Ye
Neurocomputing6
2025 Tracing and Dissecting How LLMs Recall Factual Knowledge for Real World Questions
abstract
Yiqun Wang, Chaoqun Wan, Sile Hu, Yonggang Zhang, Xiang Tian, Yaowu Chen, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chaoqun Wan, Sile Hu, Yonggang Zhang 0003, Xiang Tian 0002, Yaowu Chen, Xu Shen 0001, Jieping Ye
ACL (1)5
2025 Targeted Knowledge Enhancement: A Systematic Continual Pre-Training Approach for Effective Domain Adaptation
Chaoqun Wan, Xiang Tian 0002, Yaowu Chen
IEEE Big Data3
2025 Multi-Label Zero-Shot Learning Via Contrastive Label-Based Attention
abstract
Multi-label zero-shot learning (ML-ZSL) strives to recognize all objects in an image, regardless of whether they are present in the training data. Recent methods incorporate an attention mechanism to locate labels in the image and generate class-specific semantic information. However, the attention mechanism built on visual features treats label embeddings equally in the prediction score, leading to severe semantic ambiguity. This study focuses on efficiently utilizing semantic information in the attention mechanism. We propose a contrastive label-based attention method (CLA) to associate each label with the most relevant image regions. Specifically, our label-based attention, guided by the latent label embedding, captures discriminative image details. To distinguish region-wise correlations, we implement a region-level contrastive loss. In addition, we utilize a global feature alignment module to identify labels with general information. Extensive experiments on two benchmarks, NUS-WIDE and Open Images, demonstrate that our CLA outperforms the state-of-the-art methods. Especially under the ZSL setting, our method achieves 2.0% improvements in mean Average Precision (mAP) for NUS-WIDE and 4.0% for Open Images compared with recent methods.
Shixuan Meng, Rongxin Jiang 0001, Xiang Tian 0002, Fan Zhou 0007, Yaowu Chen, Junjie Liu 0002, Chen Shen 0003
Int. J. Neural Syst.3
2025 Addressing task conflicts in LLMs multi-task fine-tuning with task-specific subnetwork refinement
Chaoqun Wan, Xiang Tian 0002, Yaowu Chen
Mach. Learn.3
2025 SonarPoint: Weak-Heterogeneity Awareness Object Detection Network for 3D Sonar Point Cloud
abstract
Underwater target detection is primarily achieved through two methods: optical imaging and underwater sonar. 3D sonar, as the most advanced underwater detection technology, is characterized by strong penetration and long scanning distance, making it more suitable for tasks such as deep-sea exploration, murky water detection, and long-distance target identification. However, acquiring underwater sonar images is challenging, and there is no open-source 3D sonar dataset. Traditional three-dimensional target detection methods typically require highquality data and face significant challenges when dealing with weak heterogeneous sonar point clouds caused by high noise, low resolution, and occlusions. To address the aforementioned issues, we first propose a novel fuzzy decoupling module that differs from traditional foreground-background segmentation. This module simultaneously extracts valuable information about the target and its surrounding environment, mitigating the reduction in heterogeneity caused by noise and sonar side lobes. To achieve efficient fusion and capture global information after fuzzy decoupling, a multi-hop Mamba seamless adaptive decoupling point is introduced. It effectively enhances the connection between the two decoupled parts. To address missing and occlusion problems, a second-stage refinement based on Markov prediction is proposed. This low-cost approach, in contrast to using the original point cloud for contour and detail completion, enriches target boundary information. To validate our method, we have designed a practical 3D sonar imaging system and tested it through lake-based experiments. We have collected extensive raw data from Qiandao Lake and conducted annotation work. Through qualitative and quantitative experiments, our method outperforms the most advanced methods by 11.4%.
Tiancheng Cai, Peng Chen 0008, Weibo Mao, Yingtian Hu, Yilong Zhang 0001, Yuanjie Dang, Ronghua Liang, Xiang Tian 0002
IEEE Trans. Circuits Syst. Video Technol.9
2025 DGL-GAN: discriminator-guided GAN compression
Yuesong Tian, Li Shen 0008, Xiang Tian 0002, Dacheng Tao, Zhifeng Li 0001, Wei Liu 0005, Yaowu Chen
Vis. Comput.3
2024 URRL-IMVC: Unified and Robust Representation Learning for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) aims to cluster multi-view data that are only partially available. This poses two main challenges: effectively leveraging multi-view information and mitigating the impact of missing views. Prevailing solutions employ cross-view contrastive learning and missing view recovery techniques. However, they either neglect valuable complementary information by focusing only on consensus between views or provide unreliable recovered views due to the absence of supervision. To address these limitations, we propose a novel Unified and Robust Representation Learning for Incomplete Multi-View Clustering (URRL-IMVC). URRL-IMVC directly learns a unified embedding that is robust to view missing conditions by integrating information from multiple views and neighboring samples. Firstly, to overcome the limitations of cross-view contrastive learning, URRL-IMVC incorporates an attention-based auto-encoder framework to fuse multi-view information and generate unified embeddings. Secondly, URRL-IMVC directly enhances the robustness of the unified embedding against view-missing conditions through KNN imputation and data augmentation techniques, eliminating the need for explicit missing view recovery. Finally, incremental improvements are introduced to further enhance the overall performance, such as the Clustering Module and the customization of the Encoder. We extensively evaluate the proposed URRL-IMVC framework on various benchmark datasets, demonstrating its state-of-the-art performance. Furthermore, comprehensive ablation studies are performed to validate the effectiveness of our design.
Ge Teng, Ting Mao, Chen Shen 0003, Xiang Tian 0002, Yaowu Chen, Jieping Ye
KDD4
2024 Rethinking Out-of-Distribution Detection From a Human-Centric Perspective
Yao Zhu 0003, Yuefeng Chen, Rong Zhang 0006, Hui Xue 0001, Xiang Tian 0002, Rongxin Jiang 0001, Bolun Zheng, Yaowu Chen
Int. J. Comput. Vis.6
2024 A Unified Asymmetric Knowledge Distillation Framework for Image Classification
abstract
Abstract Knowledge distillation is a model compression technique that transfers knowledge learned by teacher networks to student networks. Existing knowledge distillation methods greatly expand the forms of knowledge, but also make the distillation models complex and symmetric. However, few studies have explored the commonalities among these methods. In this study, we propose a concise distillation framework to unify these methods and a method to construct asymmetric knowledge distillation under the framework. Asymmetric distillation aims to enable differentiated knowledge transfers for different distillation objects. We designed a multi-stage shallow-wide branch bifurcation method to distill different knowledge representations and a grouping ensemble strategy to supervise the network to teach and learn selectively. Consequently, we conducted experiments using image classification benchmarks to verify the proposed method. Experimental results show that our implementation can achieve considerable improvements over existing methods, demonstrating the effectiveness of the method and the potential of the framework.
Xin Ye 0010, Xiang Tian 0002, Bolun Zheng, Fan Zhou 0007, Yaowu Chen
Neural Process. Lett.2
2024 Knowledge Distillation via Multi-Teacher Feature Ensemble
abstract
This letter proposes a novel method for effectively utilizing multiple teachers in feature-based knowledge distillation. Our method involves a multi-teacher feature ensemble module for generating a robust feature ensemble and a student-teacher mapping module for bridging the student feature and ensemble feature. In addition, we utilize separate optimization, where the student's feature extractor is optimized under distillation supervision while its classifier is obtained through classifier reconstruction. We evaluate our method on the CIFAR-100, ImageNet and MS-COCO datasets, and the experimental results demonstrate its effectiveness.
Xin Ye 0010, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen
IEEE Signal Process. Lett.3
2023 Information-Containing Adversarial Perturbation for Combating Facial Manipulation Systems
abstract
With the development of deep learning technology, the facial manipulation system has become powerful and easy to use. Such systems can modify the attributes of the given facial images, such as hair color, gender, and age. Malicious applications of such systems pose a serious threat to individuals’ privacy and reputation. Existing studies have proposed various approaches to protect images against facial manipulations. Passive defense methods aim to detect whether the face is real or fake, which works for posterior forensics but can not prevent malicious manipulation. Initiative defense methods protect images upfront by injecting adversarial perturbations into images to disrupt facial manipulation systems but can not identify whether the image is fake. To address the limitation of existing methods, we propose a novel two-tier protection method named Information-containing Adversarial Perturbation (IAP), which provides more comprehensive protection for facial images. We use an encoder to map a facial image and its identity message to a cross-model adversarial example which can disrupt multiple facial manipulation systems to achieve initiative protection. Recovering the message in adversarial examples with a decoder serves passive protection, contributing to provenance tracking and fake image detection. We introduce a feature-level correlation measurement that is more suitable to measure the difference between the facial images than the commonly used mean squared error. Moreover, we propose a spectral diffusion method to spread messages to different frequency channels, thereby improving the robustness of the message against facial manipulation. Extensive experimental results demonstrate that our proposed IAP can recover the messages from the adversarial examples with high average accuracy and effectively disrupt the facial manipulation systems.
Yao Zhu 0003, Yuefeng Chen, Rong Zhang 0006, Xiang Tian 0002, Bolun Zheng, Yaowu Chen
IEEE Trans. Inf. Forensics Secur.5
2022 MPC: Multi-view Probabilistic Clustering
abstract
Despite the promising progress having been made, the two challenges of multi-view clustering (MVC) are still waiting for better solutions: i) Most existing methods are either not qualified or require additional steps for incomplete multi-view clustering and ii) noise or outliers might significantly degrade the overall clustering performance. In this paper, we propose a novel unified framework for incomplete and complete MVC named multi-view probabilistic clustering (MPC). MPC equivalently transforms multi-view pairwise posterior matching probability into composition of each view's individual distribution, which tolerates data missing and might extend to any number of views. Then graph-context-aware refinement with path propagation and co-neighbor propagation is used to refine pairwise probability, which alleviates the impact of noise and outliers. Finally, MPC also equivalently transforms probabilistic clustering's objective to avoid complete pairwise computation and adjusts clustering assignments by maximizing joint probability iteratively. Extensive experiments on multiple benchmarks for incomplete and complete MVC show that MPC significantly outperforms previous state-of-the-art methods in both effectiveness and efficiency.
Junjie Liu 0002, Junlong Liu, Shaotian Yan, Rongxin Jiang 0001, Xiang Tian 0002, Boxuan Gu, Yaowu Chen, Chen Shen 0003, Jianqiang Huang 0001
CVPR5
2022 Boosting Out-of-distribution Detection with Typical Features
abstract
Out-of-distribution (OOD) detection is a critical task for ensuring the reliability and safety of deep neural networks in real-world scenarios. Different from most previous OOD detection methods that focus on designing OOD scores or introducing diverse outlier examples to retrain the model, we delve into the obstacle factors in OOD detection from the perspective of typicality and regard the feature's high-probability region of the deep model as the feature's typical set. We propose to rectify the feature into its typical set and calculate the OOD score with the typical features to achieve reliable uncertainty estimation. The feature rectification can be conducted as a plug-and-play module with various OOD scores. We evaluate the superiority of our method on both the commonly used benchmark (CIFAR) and the more challenging high-resolution benchmark with large label space (ImageNet). Notably, our approach outperforms state-of-the-art methods by up to 5.11% in the average FPR95 on the ImageNet benchmark.
Yao Zhu 0003, Yuefeng Chen, Chuanlong Xie, Rong Zhang 0006, Hui Xue 0001, Xiang Tian 0002, Bolun Zheng, Yaowu Chen
NeurIPS7
2022 Dynamic supervisor for cross-dataset object detection
Ze Chen 0001, Zhihang Fu, Jianqiang Huang 0001, Mingyuan Tao, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen, Xian-Sheng Hua 0001
Neurocomputing7
2022 Learning Frequency Domain Priors for Image Demoireing
abstract
Image demoireing is a multi-faceted image restoration task involving both moire pattern removal and color restoration. In this paper, we raise a general degradation model to describe an image contaminated by moire patterns, and propose a novel multi-scale bandpass convolutional neural network (MBCNN) for single image demoireing. For moire pattern removal, we propose a multi-block-size learnable bandpass filters (M-LBFs), based on a block-wise frequency domain transform, to learn the frequency domain priors of moire patterns. We also introduce a new loss function named Dilated Advanced Sobel loss (D-ASL) to better sense the frequency information. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, and then performs local fine tuning of the color per pixel. To determine the most appropriate frequency domain transform, we investigate several transforms including DCT, DFT, DWT, learnable non-linear transform and learnable orthogonal transform. We finally adopt the DCT. Our basic model won the AIM2019 demoireing challenge. Experimental results on three public datasets show that our method outperforms state-of-the-art methods by a large margin.
Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Xiang Tian 0002, Jiyong Zhang 0001, Yaoqi Sun, Lin Liu 0016, Ales Leonardis, Gregory Slabaugh
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Toward Understanding and Boosting Adversarial Transferability From a Distribution Perspective
abstract
Transferable adversarial attacks against Deep neural networks (DNNs) have received broad attention in recent years. An adversarial example can be crafted by a surrogate model and then attack the unknown target model successfully, which brings a severe threat to DNNs. The exact underlying reasons for the transferability are still not completely understood. Previous work mostly explores the causes from the model perspective, e.g., decision boundary, model architecture, and model capacity. Here, we investigate the transferability from the data distribution perspective and hypothesize that pushing the image away from its original distribution can enhance the adversarial transferability. To be specific, moving the image out of its original distribution makes different models hardly classify the image correctly, which benefits the untargeted attack, and dragging the image into the target distribution misleads the models to classify the image as the target class, which benefits the targeted attack. Towards this end, we propose a novel method that crafts adversarial examples by manipulating the distribution of the image. We conduct comprehensive transferable attacks against multiple DNNs to demonstrate the effectiveness of the proposed method. Our method can significantly improve the transferability of the crafted attacks and achieves state-of-the-art performance in both untargeted and targeted scenarios, surpassing the previous best method by up to 40% in some cases. In summary, our work provides new insight into studying adversarial transferability and provides a strong counterpart for future research on adversarial defense.
Yao Zhu 0003, Yuefeng Chen, Kejiang Chen, Yuan He 0011, Xiang Tian 0002, Bolun Zheng, Yaowu Chen, Qingming Huang
IEEE Trans. Image Process.6
2021 Spatial likelihood voting with self-knowledge distillation for weakly supervised object detection
Ze Chen 0001, Zhihang Fu, Jianqiang Huang 0001, Mingyuan Tao, Rongxin Jiang 0001, Xiang Tian 0002, Yaowu Chen, Xian-Sheng Hua 0001
Image Vis. Comput.6
2020 Implicit Dual-Domain Convolutional Network for Robust Color Image Compression Artifact Reduction
abstract
Several dual-domain convolutional neural network-based methods show outstanding performance in reducing image compression artifacts. However, they are unable to handle color images as the compression processes for gray scale and color images are different. Moreover, these methods train a specific model for each compression quality, and they require multiple models to achieve different compression qualities. To address these problems, we proposed an implicit dual-domain convolutional network (IDCN) with a pixel position labeling map and quantization tables as inputs. We proposed an extractor-corrector framework-based dual-domain correction unit (DCU) as the basic component to formulate the IDCN; the implicit dual-domain translation allows the IDCN to handle color images with discrete cosine transform (DCT)-domain priors. A flexible version of IDCN (IDCN-f) was also developed to handle a wide range of compression qualities. Experiments for both objective and subjective evaluations on benchmark datasets show that IDCN is superior to the state-of-the-art methods and IDCN-f exhibits excellent abilities to handle a wide range of compression qualities with a little trade-off against performance; further, it demonstrates great potential for practical applications.
Bolun Zheng, Yaowu Chen, Xiang Tian 0002, Fan Zhou 0007
IEEE Trans. Circuits Syst. Video Technol.3
2019 Frame Interpolation Using Phase and Amplitude Feature Pyramids
abstract
This paper presents a compact neural network for video frame interpolation using phase and amplitude feature pyramids. We design a set of one-dimensional separable complex Gabor filters to extract phase and amplitude feature pyramids for each input image, which is efficient and effective for motion representation. The pyramids are fused and fed into a decoder network to estimate bi-directional optical flow. The interpolated frame is refined by a context-aware synthesis module. We train our model on quintets of frames using motion linear regularization. The proposed network contains much fewer parameters than state-of-the-art approaches. The experiments show that our method outperforms the competing methods. Moreover, our method achieves marked visual improvement in the challenging scenario with lighting changes.
Lunan Zhou, Yaowu Chen, Xiang Tian 0002, Rongxin Jiang 0001
ICIP3
2017 Illumination insensitive efficient second-order minimization for planar object tracking
abstract
Tracking for planar objects is an important issue to vision-based robotic applications. In direct visual tracking (DVT) methods, the similarity between two images is often measured through the sum of squared differences (SSD) especially with the efficient second-order minimization (ESM) due to its simplicity and efficiency. However, SSD-based ESM is not robust to illumination changes since it is usually built upon the brightness constancy assumption. Contrast to image brightness, gradient orientations (GO) are invariant to both linear and non-linear illumination changes as verified in practice. Based on GO, we propose an illumination insensitive ESM method for planar object tracking in this paper. In order to introduce GO into the ESM, we generalized the original ESM formulas for multi-dimensional features. In addition, a denoising method based on the Perona-Malik function and a mask image were suggested to improve GO's robustness against image noise and low texture. Our experimental results on dataset for planar objects with illumination changes and a benchmark dataset confirm the proposed method is robust to illumination variations and capable to deal with the general tracking challenges.
Lin Chen 0030, Fan Zhou 0007, Xiang Tian 0002, Haibin Ling, Yaowu Chen
ICRA4
2014 Fast transcoding from H.264 to HEVC based on region feature analysis
Yaowu Chen, Xiang Tian 0002
Multim. Tools Appl.3
2011 Video image assessment with a distortion-weighing spatiotemporal visual attention model
Xiang Tian 0002, Yaowu Chen
Multim. Tools Appl.2