Xu Shen 0001

dblp:09/10130-1 · DBLP profile ↗
← Back
38ranked-venue papers
5as first author
26since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 3 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 4 first-author · 12 since 2021
YearPublicationVenuePosition
2026 Bridging the Language Gap: Uncovering and Aligning Shared Circuits for Multi-Hop Reasoning in Multilingual LLMs
abstract
Large language models (LLMs) present a paradox: they can correctly answer a multi-hop factual query in a high-resource language like English, yet fail on the identical query in another language. This raises a fundamental question about the nature of multilingual knowledge: are facts missing, or merely inaccessible? The underlying mechanisms for this knowledge gap have remained largely unexplored. In this work, we resolve this question by introducing a mechanistic interpretability framework that traces the causal pathways of multi-hop knowledge reasoning. Our analysis reveals a core, non-obvious finding: cross-lingual inconsistencies do not stem from a knowledge deficit. Instead, factual knowledge is robustly stored in a set of **shared, language-agnostic semantic neurons**. The failure originates from **misaligned attention pathways**, where a common set of critical attention heads fails to correctly route information along the reasoning chain to the appropriate knowledge neurons in lower-resource languages. This mechanistic diagnosis motivates a targeted alignment strategy: a surgical fine-tuning of only these critical heads. Experiments demonstrate that our method achieves significant improvements in multilingual multi-hop factuality—with positive cross-lingual transfer—while uniquely preserving general model capabilities, offering a scalable and mechanistically-grounded approach to building more reliable multilingual models.
Zhen Huang 0007, Yonggang Zhang 0003, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
AAAI5
2025 Visual Evidence Prompting Mitigates Hallucinations in Large Vision-Language Models
abstract
Wei Li, Zhen Huang, Houqiang Li, Le Lu, Yang Lu, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wei Li 0317, Zhen Huang 0007, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
ACL (1)7
2025 Interpret and Improve In-Context Learning via the Lens of Input-Label Mappings
abstract
Chenghao Sun, Zhen Huang, Yonggang Zhang, Le Lu, Houqiang Li, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhen Huang 0007, Yonggang Zhang 0003, Le Lu 0001, Houqiang Li, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
ACL (1)7
2025 Tracing and Dissecting How LLMs Recall Factual Knowledge for Real World Questions
abstract
Yiqun Wang, Chaoqun Wan, Sile Hu, Yonggang Zhang, Xiang Tian, Yaowu Chen, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Chaoqun Wan, Sile Hu, Yonggang Zhang 0003, Xiang Tian 0002, Yaowu Chen, Xu Shen 0001, Jieping Ye
ACL (1)7
2025 Knowledge Graph Finetuning Enhances Knowledge Manipulation in Large Language Models
abstract
Despite the impressive performance of general large language models(LLMs), many of their applications in specific domains (e.g., low-data and knowledge-intensive) still confront significant challenges. Supervised fine-tuning (SFT)---where a general LLM is further trained on a small labeled dataset to adapt for specific tasks or domains---has shown great power for developing domain-specific LLMs. However, existing SFT data primarily consist of Question and Answer (Q&A) pairs, which poses a significant challenge for LLMs to comprehend the correlation and logic of knowledge underlying the Q&A. To address this challenge, we propose a conceptually flexible and general framework to boost SFT, namely Knowledge Graph-Driven Supervised Fine-Tuning (KG-SFT). The key idea of KG-SFT is to generate high-quality explanations for each Q&A pair via a structured knowledge graph to enhance the knowledge comprehension and manipulation of LLMs. Specifically, KG-SFT consists of three components: Extractor, Generator, and Detector. For a given Q&A pair, (i) Extractor first identifies entities within Q&A pairs and extracts relevant reasoning subgraphs from external KGs, (ii) Generator then produces corresponding fluent explanations utilizing these reasoning subgraphs, and (iii) finally, Detector performs sentence-level knowledge conflicts detection on these explanations to guarantee the reliability. KG-SFT focuses on generating high-quality explanations to improve the quality of Q&A pair, which reveals a promising direction for supplementing existing data augmentation methods. Extensive experiments on fifteen different domains and six different languages demonstrate the effectiveness of KG-SFT, leading to an accuracy improvement of up to 18% and an average of 8.7% in low-data scenarios.
Hanzhu Chen, Xu Shen 0001, Jie Wang 0005, Qitan Lv, Feng Wu 0001, Jieping Ye
ICLR2
2025 Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMs
abstract
Task arithmetic is a straightforward yet highly effective strategy for model merging, enabling the resultant model to exhibit multi-task capabilities. Recent research indicates that models demonstrating linearity enhance the performance of task arithmetic. In contrast to existing methods that rely on the global linearization of the model, we argue that this linearity already exists within the model's submodules. In particular, we present a statistical analysis and show that submodules (e.g., layers, self-attentions, and MLPs) exhibit significantly higher linearity than the overall model. Based on these findings, we propose an innovative model merging strategy that independently merges these submodules. Especially, we derive a closed-form solution for optimal merging weights grounded in the linear properties of these submodules. Experimental results demonstrate that our method consistently outperforms the standard task arithmetic approach and other established baselines across different model scales and various tasks. This result highlights the benefits of leveraging the linearity of submodules and provides a new perspective for exploring solutions for effective and practical multi-task model merging.
Rui Dai 0005, Sile Hu, Xu Shen 0001, Yonggang Zhang 0003, Xinmei Tian 0001, Jieping Ye
ICLR3
2025 Improving the Instance-Dependent Transition Matrix Estimation by Exploiting Self-Supervised Learning
abstract
The transition matrix reveals the transition relationship between clean labels and noisy labels. It plays an important role in building statistically consistent classifiers for learning with noisy labels. However, in real-world applications, the transition matrix is usually unknown and has to be estimated. It is a challenging task to accurately estimate the transition matrix which usually depends on the instance. With both instances and noisy labels at hand, the major difficulty of estimating the transition matrix comes from the absence of clean label information. Recent work suggests that self-supervised learning methods can effectively infer clean label information. These methods could even achieve comparable performance with supervised learning on many benchmark datasets but without requiring any labels. Motivated by this, our paper presents a practical approach that harnesses self-supervised learning to extract clean label information, which reduces the estimation error of the instance-dependent transition matrix. By exploiting the estimated transition matrix, the performance of classifiers is improved. Empirical results on different datasets illustrate that our proposed methodology outperforms existing state-of-the-art methods in terms of both classification accuracy and transition matrix estimation.
Yexiong Lin, Yu Yao 0005, Zhaoqing Wang, Xu Shen 0001, Jun Yu 0001, Bo Han 0003, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 SAC-KG: Exploiting Large Language Models as Skilled Automatic Constructors for Domain Knowledge Graph
abstract
Knowledge graphs (KGs) play a pivotal role in knowledge-intensive tasks across specialized domains, where the acquisition of precise and dependable knowledge is crucial.However, existing KG construction methods heavily rely on human intervention to attain qualified KGs, which severely hinders the practical applicability in real-world scenarios.To address this challenge, we propose a general KG construction framework, named SAC-KG, to exploit large language models (LLMs) as Skilled Automatic Constructors for domain Knowledge Graph.SAC-KG effectively involves LLMs as domain experts to generate specialized and precise multi-level KGs.Specifically, SAC-KG consists of three components: Generator, Verifier, and Pruner.For a given entity, Generator produces its relations and tails from raw domain corpora, to construct a specialized single-level KG.Verifier and Pruner then work together to ensure precision by correcting generation errors and determining whether newly produced tails require further iteration for the next-level KG.Experiments demonstrate that SAC-KG automatically constructs a domain KG at the scale of over one million nodes and achieves a precision of 89.32%, leading to a superior performance with over 20% increase in precision rate compared to existing state-of-the-art methods for the KG construction task.
Hanzhu Chen, Xu Shen 0001, Qitan Lv, Jie Wang 0005, Xiaoqi Ni, Jieping Ye
ACL (1)2
2024 Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning
abstract
Extending large image-text pre-trained models (e.g., CLIP) for video understanding has made significant advancements. To enable the capability of CLIP to perceive dynamic information in videos, existing works are dedicated to equipping the visual encoder with various temporal modules. However, these methods exhibit “asymmetry” between the visual and textual sides, with neither temporal descriptions in input texts nor temporal modules in text encoder. This limitation hinders the potential of language supervision emphasized in CLIP, and restricts the learning of temporal features, as the text encoder has demonstrated limited proficiency in motion understanding. To address this issue, we propose leveraging “MoTion-Enhanced Descriptions” (MoTED) to facilitate the extraction of distinctive temporal features in videos. Specifically, we first generate discriminative motion-related descriptions via querying GPT-4 to compare easy-confusing action categories. Then, we incorporate both the visual and textual encoders with additional perception modules to process the video frames and generated descriptions, respectively. Finally, we adopt a contrastive loss to align the visual and textual motion features. Extensive experiments on five benchmarks show that MoTED surpasses state-of-the-art methods with convincing gaps, laying a solid foundation for empowering CLIP with strong temporal modeling.
Chaoqun Wan, Tongliang Liu, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
CVPR5
2024 Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional Understanding
abstract
Contrastively trained vision-language models such as CLIP have achieved remarkable progress in vision and language representation learning.Despite the promising progress, their proficiency in compositional reasoning over attributes and relations (e.g., distinguishing between "the car is underneath the person" and "the person is underneath the car") remains notably inadequate.We investigate the cause for this deficient behavior is the composition attribution issue, where the attribution scores (e.g., attention scores or GradCAM scores) for relations (e.g., underneath) or attributes (e.g., red) in the text are substantially lower than those for object terms.In this work, we show such issue is mitigated via a novel framework called CAE (Composition Attribution Enhancement).This generic framework incorporates various interpretable attribution methods to encourage the model to pay greater attention to composition words denoting relationships and attributes within the text.Detailed analysis shows that our approach enables the models to adjust and rectify the attribution of the texts.Extensive experiments across seven benchmarks reveal that our framework significantly enhances the ability to discern intricate details and construct more sophisticated interpretations of combined visual and linguistic elements.
Wei Li 0317, Zhen Huang 0007, Xinmei Tian 0001, Le Lu 0001, Houqiang Li, Xu Shen 0001, Jieping Ye
EMNLP6
2024 From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning
abstract
Large Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend to admit mistakes and provide inaccurate responses even if they initially provided the correct answer. Recent works propose to employ supervised fine-tuning (SFT) to mitigate the sycophancy issue, while it typically leads to the degeneration of LLMs' general capability. To address the challenge, we propose a novel supervised pinpoint tuning (SPT), where the region-of-interest modules are tuned for a given objective. Specifically, SPT first reveals and verifies a small percentage (<5%) of the basic modules, which significantly affect a particular behavior of LLMs. i.e., sycophancy. Subsequently, SPT merely fine-tunes these identified modules while freezing the rest. To verify the effectiveness of the proposed SPT, we conduct comprehensive experiments, demonstrating that SPT significantly mitigates the sycophancy issue of LLMs (even better than SFT). Moreover, SPT introduces limited or even no side effects on the general capability of LLMs. Our results shed light on how to precisely, effectively, and efficiently explain and improve the targeted ability of LLMs.
Wei Chen 0005, Zhen Huang 0007, Liang Xie 0003, Binbin Lin 0001, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Deng Cai 0001, Yonggang Zhang 0003, Wenxiao Wang 0001, Xu Shen 0001, Jieping Ye
ICML11
2024 Interpreting and Improving Large Language Models in Arithmetic Calculation
abstract
Large language models (LLMs) have demonstrated remarkable potential across numerous applications and have shown an emergent ability to tackle complex reasoning tasks, such as mathematical computations. However, even for the simplest arithmetic calculations, the intrinsic mechanisms behind LLMs remains mysterious, making it challenging to ensure reliability. In this work, we delve into uncovering a specific mechanism by which LLMs execute calculations. Through comprehensive experiments, we find that LLMs frequently involve a small fraction ($<$5%) of attention heads, which play a pivotal role in focusing on operands and operators during calculation processes. Subsequently, the information from these operands is processed through multi-layer perceptrons (MLPs), progressively leading to the final solution. These pivotal heads/MLPs, though identified on a specific dataset, exhibit transferability across different datasets and even distinct tasks. This insight prompted us to investigate the potential benefits of selectively fine-tuning these essential heads/MLPs to boost the LLMs’ computational performance. We empirically find that such precise tuning can yield notable enhancements on mathematical prowess, without compromising the performance on non-mathematical tasks. Our work serves as a preliminary exploration into the arithmetic calculation abilities inherent in LLMs, laying a solid foundation to reveal more intricate mathematical tasks.
Chaoqun Wan, Yonggang Zhang 0003, Yiu-Ming Cheung, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye
ICML6
2024 Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control
abstract
As the development and application of Large Language Models (LLMs) continue to advance rapidly, enhancing their trustworthiness and aligning them with human preferences has become a critical area of research. Traditional methods rely heavily on extensive data for Reinforcement Learning from Human Feedback (RLHF), but representation engineering offers a new, training-free approach. This technique leverages semantic features to control the representation of LLM's intermediate hidden states, enabling the model to meet specific requirements such as increased honesty or heightened safety awareness. However, a significant challenge arises when attempting to fulfill multiple requirements simultaneously. It proves difficult to encode various semantic contents, like honesty and safety, into a singular semantic feature, restricting its practicality. In this work, we address this challenge through Sparse Activation Control. By delving into the intrinsic mechanisms of LLMs, we manage to identify and pinpoint modules that are closely related to specific tasks within the model, i.e. attention heads. These heads display sparse characteristics that allow for near-independent control over different tasks. Our experiments, conducted on the open-source Llama series models, have yielded encouraging results. The models were able to align with human preferences on issues of safety, factualness, and bias concurrently.
Yuxin Xiao, Chaoqun Wan, Yonggang Zhang 0003, Wenxiao Wang 0001, Binbin Lin 0001, Xiaofei He 0001, Xu Shen 0001, Jieping Ye
NeurIPS7
2024 Conditional Consistency Regularization for Semi-Supervised Multi-Label Image Classification
abstract
Consistency regularization has achieved great successes in Semi-Supervised Single-Label Image Classification (SS-SLC) with deep learning models, while few effort has been devoted to Semi-Supervised Multi-Label Image Classification (SS-MLC) with deep learning models. One intuitive solution for introducing consistency regularization to SS-MLC is to regularize model predictions to be invariant to different augmented data of the same input image. However, the solution lacks the consideration of label relations, which are key elements in multi-label image classification. In this article, we go beyond the consistency regularization for multi-view input images, and propose Conditional Consistency Regularization (CCR) that is tailored for SS-MLC. Specifically, for two augmented input images, we make the two model predictions conditioned on different label states (i.e., positive, negative, or unknown for each class). By encouraging the two predictions to be consistent, the model is able to build relations between the given two different label states, which helps to make use of label relations for boosting image classification. The experiments on large-scale real-world SS-MLC benchmarks demonstrate that the proposed method can surpass state-of-the-art methods by a large margin.
Zhengning Wu, Tianyu He, Xiaobo Xia, Jun Yu 0001, Xu Shen 0001, Tongliang Liu
IEEE Trans. Multim.5
2023 Contextual Convolutional Networks
Shuxian Liang, Xu Shen 0001, Tongliang Liu, Xian-Sheng Hua 0001
ICLR2
2023 Moderate Coreset: A Universal Method of Data Selection for Real-world Data-efficient Deep Learning
Xiaobo Xia, Jun Yu 0001, Xu Shen 0001, Bo Han 0003, Tongliang Liu
ICLR4
2023 CS-Isolate: Extracting Hard Confident Examples by Content and Style Isolation
abstract
Label noise widely exists in large-scale image datasets. To mitigate the side effects of label noise, state-of-the-art methods focus on selecting confident examples by leveraging semi-supervised learning. Existing research shows that the ability to extract hard confident examples, which are close to the decision boundary, significantly influences the generalization ability of the learned classifier. In this paper, we find that a key reason for some hard examples being close to the decision boundary is due to the entanglement of style factors with content factors. The hard examples become more discriminative when we focus solely on content factors, such as semantic information, while ignoring style factors. Nonetheless, given only noisy data, content factors are not directly observed and have to be inferred. To tackle the problem of inferring content factors for classification when learning with noisy labels, our objective is to ensure that the content factors of all examples in the same underlying clean class remain unchanged as their style information changes. To achieve this, we utilize different data augmentation techniques to alter the styles while regularizing content factors based on some confident examples. By training existing methods with our inferred content factors, CS-Isolate proves their effectiveness in learning hard examples on benchmark datasets. The implementation is available at https://github.com/tmllab/2023_NeurIPS_CS-isolate.
Yexiong Lin, Yu Yao 0005, Mingming Gong, Xu Shen 0001, Dong Xu 0001, Tongliang Liu
NeurIPS5
2022 Cloth-Changing Person Re-identification from A Single Image with Gait Prediction and Regularization
abstract
Cloth-Changing person re-identification (CC-ReID) aims at matching the same person across different locations over a long-duration, e.g., over days, and therefore inevitably has cases of changing clothing. In this paper, we focus on handling well the CC-ReID problem under a more challenging setting, i.e., just from a single image, which enables an efficient and latency-free person identity matching for surveillance. Specifically, we introduce Gait recognition as an auxiliary task to drive the Image ReID model to learn cloth-agnostic representations by leveraging personal unique and cloth-independent gait information, we name this framework as GI-ReID. GI-ReID adopts a two-stream architecture that consists of an image ReID-Stream and an auxiliary gait recognition stream (Gait-Stream). The Gait-Stream, that is discarded in the inference for high efficiency, acts as a regulator to encourage the ReID-Stream to capture cloth-invariant biometric motion features during the training. To get temporal continuous motion cues from a single image, we design a Gait Sequence Prediction (GSP) module for Gait-Stream to enrich gait information. Finally, a semantics consistency constraint over two streams is enforced for effective knowledge regularization. Extensive experiments on multiple image-based Cloth-Changing ReID benchmarks, e.g., LTCC, PRCC, Real28, and VC-Clothes, demonstrate that GI-ReID performs favorably against the state-of-the-art methods.
Xin Jin 0014, Tianyu He, Kecheng Zheng, Zhiheng Yin, Xu Shen 0001, Zhen Huang 0007, Ruoyu Feng 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
CVPR5
2022 Meta Convolutional Neural Networks for Single Domain Generalization
abstract
In single domain generalization, models trained with data from only one domain are required to perform well on many unseen domains. In this paper, we propose a new model, termed meta convolutional neural network, to solve the single domain generalization problem in image recognition. The key idea is to decompose the convolutional features of images into meta features. Acting as “visual words”, meta features are defined as universal and basic visual elements for image representations (like words for documents in language). Taking meta features as reference, we propose compositional operations to eliminate irrelevant features of local convolutional features by an addressing process and then to reformulate the convolutional feature maps as a composition of related meta features. In this way, images are universally coded without biased information from the unseen domain, which can be processed by following modules trained in the source domain. The compositional operations adopt a regression analysis technique to learn the meta features in an online batch learning manner. Extensive experiments on multiple benchmark datasets verify the superiority of the proposed model in improving single domain generalization ability.
Chaoqun Wan, Xu Shen 0001, Yonggang Zhang 0003, Zhiheng Yin, Xinmei Tian 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR2
2022 Delving into Details: Synopsis-to-Detail Networks for Video Recognition
Shuxian Liang, Xu Shen 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ECCV (4)2
2022 Meta Clustering Learning for Large-scale Unsupervised Person Re-identification
abstract
Unsupervised Person Re-identification (U-ReID) with pseudo labeling recently reaches a competitive performance compared to fully-supervised ReID methods based on modern clustering algorithms. However, such clustering-based scheme becomes computationally prohibitive for large-scale datasets, making it infeasible to be applied in real-world application. How to efficiently leverage endless unlabeled data with limited computing resources for better U-ReID is under-explored. In this paper, we make the first attempt to the large-scale U-ReID and propose a "small data for big task" paradigm dubbed Meta Clustering Learning (MCL). MCL only pseudo-labels a subset of the entire unlabeled data via clustering to save computing for the first-phase training. After that, the learned cluster centroids, termed as meta-prototypes in our MCL, are regarded as a proxy annotator to softly annotate the rest unlabeled data for further polishing the model. To alleviate the potential noisy labeling issue in the polishment phase, we enforce two well-designed loss constraints to promise intra-identity consistency and inter-identity strong correlation. For multiple widely-used U-ReID benchmarks, our method significantly saves computational cost while achieving a comparable or even better performance compared to prior works.
Xin Jin 0014, Tianyu He, Xu Shen 0001, Tongliang Liu, Xinchao Wang, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
ACM Multimedia3
2021 Partial Person Re-Identification With Part-Part Correspondence Learning
abstract
Driven by the success of deep learning, the last decade has seen rapid advances in person re-identification (re-ID). Nonetheless, most of approaches assume that the input is given with the fulfillment of expectations, while imperfect input remains rarely explored to date, which is a non-trivial problem since directly apply existing methods without adjustment can cause significant performance degradation. In this paper, we focus on recognizing partial (flawed) input with the assistance of proposed Part-Part Correspondence Learning (PPCL), a self-supervised learning framework that learns correspondence between image patches without any additional part-level supervision. Accordingly, we propose Part-Part Cycle (PP-Cycle) constraint and Part-Part Triplet (PP-Triplet) constraint that exploit the duality and uniqueness between corresponding image patches respectively. We verify our proposed PPCL on several partial person re-ID benchmarks. Experimental results demonstrate that our approach can surpass previous methods in terms of the standard evaluation metric.
Tianyu He, Xu Shen 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
CVPR2
2021 Revisiting Knowledge Distillation: An Inheritance and Exploration Framework
abstract
Knowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the teacher model and the student model. However, directly pushing the student model to mimic the probabilities/features of the teacher model to a large extent limits the student model in learning undiscovered knowledge/features. In this paper, we propose a novel inheritance and exploration knowledge distillation framework (IE-KD), in which a student model is split into two parts - inheritance and exploration. The inheritance part is learned with a similarity loss to transfer the existing learned knowledge from the teacher model to the student model, while the exploration part is encouraged to learn representations different from the inherited ones with a dis-similarity loss. Our IE-KD framework is generic and can be easily combined with existing distillation or mutual learning methods for training deep neural networks. Extensive experiments demonstrate that these two parts can jointly push the student model to learn more diversified and effective representations, and our IE-KD can be a general technique to improve the student network to achieve SOTA performance. Furthermore, by applying our IE-KD to the training of two networks, the performance of both can be improved w.r.t. deep mutual learning.
Zhen Huang 0007, Xu Shen 0001, Jun Xing, Tongliang Liu, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR2
2021 Dense Interaction Learning for Video-based Person Re-identification
abstract
Video-based person re-identification (re-ID) aims at matching the same person across video clips. Efficiently exploiting multi-scale fine-grained features while building the structural interaction among them is pivotal for its success. In this paper, we propose a hybrid framework, Dense Interaction Learning (DenseIL), that takes the principal advantages of both CNN-based and Attention-based architectures to tackle video-based person re-ID difficulties. DenseIL contains a CNN encoder and a Dense Interaction (DI) decoder. The CNN encoder is responsible for efficiently extracting discriminative spatial features while the DI decoder is designed to densely model spatial-temporal inherent interaction across frames. Different from previous works, we additionally let the DI decoder densely attends to intermediate fine-grained CNN features and that naturally yields multi-grained spatial-temporal representation for each video clip. Moreover, we introduce Spatio-TEmporal Positional Embedding (STEP-Emb) into the DI decoder to investigate the positional relation among the spatial-temporal inputs. Our experiments consistently and significantly outperform all the state-of-the-art methods on multiple standard video-based person re-ID datasets.
Tianyu He, Xin Jin 0014, Xu Shen 0001, Jianqiang Huang 0001, Zhibo Chen 0001, Xian-Sheng Hua 0001
ICCV3
2021 3D Local Convolutional Neural Networks for Gait Recognition
abstract
The goal of gait recognition is to learn the unique spatiotemporal pattern about the human body shape from its temporal changing characteristics. As different body parts behave differently during walking, it is intuitive to model the spatio-temporal patterns of each part separately. However, existing part-based methods equally divide the feature maps of each frame into fixed horizontal stripes to get local parts. It is obvious that these stripe partition-based methods cannot accurately locate the body parts. First, different body parts can appear at the same stripe (e.g., arms and the torso), and one part can appear at different stripes in different frames (e.g., hands). Second, different body parts possess different scales, and even the same part in different frames can appear at different locations and scales. Third, different parts also exhibit distinct movement patterns (e.g., at which frame the movement starts, the position change frequency, how long it lasts). To overcome these issues, we propose novel 3D local operations as a generic family of building blocks for 3D gait recognition backbones. The proposed 3D local operations support the extraction of local 3D volumes of body parts in a sequence with adaptive spatial and temporal scales, locations and lengths. In this way, the spatio-temporal patterns of the body parts are well learned from the 3D local neighborhood in partspecific scales, locations, frequencies and lengths. Experiments demonstrate that our 3D local convolutional neural networks achieve state-of-the-art performance on popular gait datasets. Code is available at: https://github.com/yellowtownhz/3DLocalCNN.
Zhen Huang 0007, Dixiu Xue, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ICCV3
2021 Video Object Segmentation with Dynamic Memory Networks and Adaptive Object Alignment
abstract
In this paper, we propose a novel solution for object-matching based semi-supervised video object segmentation, where the target object masks in the first frame are provided. Existing object-matching based methods focus on the matching between the raw object features of the current frame and the first/previous frames. However, two issues are still not solved by these object-matching based methods. As the appearance of the video object changes drastically over time, 1) unseen parts/details of the object present in the current frame, resulting in incomplete annotation in the first annotated frame (e.g. view/scale changes). 2) even for the seen parts/details of the object in the current frame, their positions change relatively (e.g. pose changes/camera motion), leading to a misalignment for the object matching. To obtain the complete information of the target object, we propose a novel object-based dynamic memory network that exploits visual contents of all the past frames. To solve the misalignment problem caused by position changes of visual contents, we propose an adaptive object alignment module by incorporating a region translation function that aligns object proposals towards templates in the feature space. Our method achieves state-of-the-art results on latest benchmark datasets DAVIS 2017 ($\mathcal{J}$ of 81.4% and $\mathcal{F}$ of 87.5% on the validation set) and YouTube-VOS (the overall score of 82.7% on the validation set) with a very efficient inference time (0.16 second/frame on DAVIS 2017 validation set). Code is available at: https://github.com/liang4sx/DMN-AOA.
Shuxian Liang, Xu Shen 0001, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ICCV2
2020 Spatio-Temporal Inception Graph Convolutional Networks for Skeleton-Based Action Recognition
abstract
Skeleton-based human action recognition has attracted much attention with the prevalence of accessible depth sensors. Recently, graph convolutional networks (GCNs) have been widely used for this task due to their powerful capability to model graph data. The topology of the adjacency graph is a key factor for modeling the correlations of the input skeletons. Thus, previous methods mainly focus on the design/learning of the graph topology. But once the topology is learned, only a single-scale feature and one transformation exist in each layer of the networks. Many insights, such as multi-scale information and multiple sets of transformations, that have been proven to be very effective in convolutional neural networks (CNNs), have not been investigated in GCNs. The reason is that, due to the gap between graph-structured skeleton data and conventional image/video data, it is very challenging to embed these insights into GCNs. To overcome this gap, we reinvent the split-transform-merge strategy in GCNs for skeleton sequence processing. Specifically, we design a simple and highly modularized graph convolutional network architecture for skeleton-based action recognition. Our network is constructed by repeating a building block that aggregates multi-granularity information from both the spatial and temporal paths. Extensive experiments demonstrate that our network outperforms state-of-the-art methods by a significant margin with only 1/5 of the parameters and 1/10 of the FLOPs.
Zhen Huang 0007, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Multimedia2
2019 Quantization Networks
abstract
Although deep neural networks are highly effective, their high computational and memory costs severely hinder their applications to portable devices. As a consequence, lowbit quantization, which converts a full-precision neural network into a low-bitwidth integer version, has been an active and promising research topic. Existing methods formulate the low-bit quantization of networks as an approximation or optimization problem. Approximation-based methods confront the gradient mismatch problem, while optimizationbased methods are only suitable for quantizing weights and can introduce high computational cost during the training stage. In this paper, we provide a simple and uniform way for weights and activations quantization by formulating it as a differentiable non-linear function. The quantization function is represented as a linear combination of several Sigmoid functions with learnable biases and scales that could be learned in a lossless and end-to-end manner via continuous relaxation of the steepness of Sigmoid functions. Extensive experiments on image classification and object detection tasks show that our quantization networks outperform state-of-the-art methods. We believe that the proposed method will shed new lights on the interpretation of neural network quantization.
Jiwei Yang, Xu Shen 0001, Jun Xing, Xinmei Tian 0001, Houqiang Li, Bing Deng, Jianqiang Huang 0001, Xian-Sheng Hua 0001
CVPR2
2019 Attribute-Driven Feature Disentangling and Temporal Aggregation for Video Person Re-Identification
abstract
Video-based person re-identification plays an important role in surveillance video analysis, expanding image-based methods by learning features of multiple frames. Most existing methods fuse features by temporal average-pooling, without exploring the different frame weights caused by various viewpoints, poses, and occlusions. In this paper, we propose an attribute-driven method for feature disentangling and frame re-weighting. The features of single frames are disentangled into groups of sub-features, each corresponds to specific semantic attributes. The sub-features are re-weighted by the confidence of attribute recognition and then aggregated at the temporal dimension as the final representation. By means of this strategy, the most informative regions of each frame are enhanced and contributes to a more discriminative sequence representation. Extensive ablation studies demonstrate the effectiveness of feature disentangling as well as temporal re-weighting. The experimental results on the iLIDS-VID, PRID-2011 and MARS datasets demonstrate that our proposed method outperforms existing state-of-the-art approaches.
Yiru Zhao, Xu Shen 0001, Zhongming Jin 0001, Hongtao Lu 0001, Xian-Sheng Hua 0001
CVPR2
2018 Sequence-to-Sequence Learning via Shared Latent Representation
abstract
Sequence-to-sequence learning is a popular research area in deep learning, such as video captioning and speech recognition. Existing methods model this learning as a mapping process by first encoding the input sequence to a fixed-sized vector, followed by decoding the target sequence from the vector. Although simple and intuitive, such mapping model is task-specific, unable to be directly used for different tasks. In this paper, we propose a star-like framework for general and flexible sequence-to-sequence learning, where different types of media contents (the peripheral nodes) could be encoded to and decoded from a shared latent representation (SLR) (the central node). This is inspired by the fact that human brain could learn and express an abstract concept in different ways. The media-invariant property of SLR could be seen as a high-level regularization on the intermediate vector, enforcing it to not only capture the latent representation intra each individual media like the auto-encoders, but also their transitions like the mapping models. Moreover, the SLR model is content-specific, which means it only needs to be trained once for a dataset, while used for different tasks. We show how to train a SLR model via dropout and use it for different sequence-to-sequence tasks. Our SLR model is validated on the Youtube2Text and MSR-VTT datasets, achieving superior performance on video-to-sentence task, and the first sentence-to-video results.
Xu Shen 0001, Xinmei Tian 0001, Jun Xing, Yong Rui, Dacheng Tao
AAAI1
2018 Local Convolutional Neural Networks for Person Re-Identification
abstract
Recent works have shown that person re-identification can be substantially improved by introducing attention mechanisms, which allow learning both global and local representations. However, all these works learn global and local features in separate branches. As a consequence, the interaction/boosting of global and local information are not allowed, except in the final feature embedding layer. In this paper, we propose local operations as a generic family of building blocks for synthesizing global and local information in any layer. This building block can be inserted into any convolutional networks with only a small amount of prior knowledge about the approximate locations of local parts. For the task of person re-identification, even with only one local block inserted, our local convolutional neural networks (Local CNN) can outperform state-of-the-art methods consistently on three large-scale benchmarks, including Market-1501, CUHK03, and DukeMTMC-ReID.
Jiwei Yang, Xu Shen 0001, Xinmei Tian 0001, Houqiang Li, Jianqiang Huang 0001, Xian-Sheng Hua 0001
ACM Multimedia2
2018 Continuous Dropout
abstract
Dropout has been proven to be an effective algorithm for training robust deep networks because of its ability to prevent overfitting by avoiding the co-adaptation of feature detectors. Current explanations of dropout include bagging, naive Bayes, regularization, and sex in evolution. According to the activation patterns of neurons in the human brain, when faced with different situations, the firing rates of neurons are random and continuous, not binary as current dropout does. Inspired by this phenomenon, we extend the traditional binary dropout to continuous dropout. On the one hand, continuous dropout is considerably closer to the activation characteristics of neurons in the human brain than traditional binary dropout. On the other hand, we demonstrate that continuous dropout has the property of avoiding the co-adaptation of feature detectors, which suggests that we can extract more independent feature detectors for model averaging in the test stage. We introduce the proposed continuous dropout to a feedforward neural network and comprehensively compare it with binary dropout, adaptive dropout, and DropConnect on Modified National Institute of Standards and Technology, Canadian Institute for Advanced Research-10, Street View House Numbers, NORB, and ImageNet large scale visual recognition competition-12. Thorough experiments demonstrate that our method performs better in preventing the co-adaptation of feature detectors and improves test performance.
Xu Shen 0001, Xinmei Tian 0001, Tongliang Liu, Dacheng Tao
IEEE Trans. Neural Networks Learn. Syst.1
2017 Patch Reordering: A NovelWay to Achieve Rotation and Translation Invariance in Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have demonstrated state-of-the-art performance on many visual recognition tasks. However, the combination of convolution and pooling operations only shows invariance to small local location changes in meaningful objects in input. Sometimes, such networks are trained using data augmentation to encode this invariance into the parameters, which restricts the capacity of the model to learn the content of these objects. A more efficient use of the parameter budget is to encode rotation or translation invariance into the model architecture, which relieves the model from the need to learn them. To enable the model to focus on learning the content of objects other than their locations, we propose to conduct patch ranking of the feature maps before feeding them into the next layer. When patch ranking is combined with convolution and pooling operations, we obtain consistent representations despite the location of meaningful objects in input. We show that the patch ranking module improves the performance of the CNN on many benchmark tasks, including MNIST digit recognition, large-scale image recognition, and image retrieval.
Xu Shen 0001, Xinmei Tian 0001, Shaoyan Sun, Dacheng Tao
AAAI1
2017 Classification and Representation Joint Learning via Deep Networks
abstract
Deep learning has been proven to be effective for classification problems. However, the majority of previous works trained classifiers by considering only class label information and ignoring the local information from the spatial distribution of training samples. In this paper, we propose a deep learning framework that considers both class label information and local spatial distribution information between training samples. A two-channel network with shared weights is used to measure the local distribution. The classification performance can be improved with more detailed information provided by the local distribution, particularly when the training samples are insufficient. Additionally, the class label information can help to learn better feature representations compared with other feature learning methods that use only local distribution information between samples. The local distribution constraint between sample pairs can also be viewed as a regularization of the network, which can efficiently prevent the overfitting problem. Extensive experiments are conducted on several benchmark image classification datasets, and the results demonstrate the effectiveness of our proposed method.
Xinmei Tian 0001, Xu Shen 0001, Dacheng Tao
IJCAI3
2016 Transform-Invariant Convolutional Neural Networks for Image Classification and Search
abstract
Convolutional neural networks (CNNs) have achieved state-of-the-art results on many visual recognition tasks. However, current CNN models still exhibit a poor ability to be invariant to spatial transformations of images. Intuitively, with sufficient layers and parameters, hierarchical combinations of convolution (matrix multiplication and non-linear activation) and pooling operations should be able to learn a robust mapping from transformed input images to transform-invariant representations. In this paper, we propose randomly transforming (rotation, scale, and translation) feature maps of CNNs during the training stage. This prevents complex dependencies of specific rotation, scale, and translation levels of training images in CNN models. Rather, each convolutional kernel learns to detect a feature that is generally helpful for producing the transform-invariant answer given the combinatorially large variety of transform levels of its input feature maps. In this way, we do not require any extra training supervision or modification to the optimization process and training images. We show that random transformation provides significant improvements of CNNs on many benchmark tasks, including small-scale image recognition, large-scale image recognition, and image retrieval.
Xu Shen 0001, Xinmei Tian 0001, Anfeng He, Shaoyan Sun, Dacheng Tao
ACM Multimedia1
2016 Multi-modal and multi-scale photo collection summarization
Xu Shen 0001, Xinmei Tian 0001
Multim. Tools Appl.1
2016 Monet: A System for Reliving Your Memories by Theme-Based Photo Storytelling
abstract
With the ever-increasing use of smartphones and digital cameras, people are now able to take photos anywhere and anytime. Most of these photos simply end up stored in the cloud without further interaction. This occurs because we lack intelligent services to organize these personal photos well. Therefore, there is an urgent need for such a system to enable people to relive their memories by turning their photos into stories. This paper presents a storytelling system named Monet, which automatically creates interesting stories from personal photos by mimicking cinematic knowledge based on a set of predesigned editing styles. The system consists of two stages: photo summarization, which selects a subset of the “best” photos to represent a photo collection, and story remixing, which generates a stylish music video from the selected photos. During photo summarization, photos are grouped into events based on multimodal features (time and location). The “best” photos are then selected according to visual quality, event representativeness, and diversity. The second stage, story remixing, automatically selects an appropriate theme-dependent editing style based on the photo content. Each selected photo is converted to a video clip by applying a virtual camera with appropriate motions. A series of video effects, color filters, shapes, and transitions are then applied to the video clips according to cinematic rules. The generated video is finally multiplexed with a music clip to generate the story. Evaluations show that our system achieves superior performance to state-of-the-art photo event detection and story generation systems.
Xu Shen 0001, Tao Mei 0001, Xinmei Tian 0001, Nenghai Yu, Yong Rui
IEEE Trans. Multim.2
2015 Photo Quality Assessment with DCNN that Understands Image Well
Xu Shen 0001, Houqiang Li, Xinmei Tian 0001
MMM (2)2