Xiangyuan Lan

dblp:151/8902 · DBLP profile ↗
← Back
82ranked-venue papers
8as first author
52since 2021 · last 2026
0000-0001-8564-0346ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 49 · 7 first-author · 27 since 2021Artificial intelligence and machine learning · 48 · 5 first-author · 32 since 2021Security and privacy · 7 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 CATAL: Causally Disentangled Task Representation Learning for Offline Meta-Reinforcement Learning
abstract
Context-based Offline Meta Reinforcement Learning (COMRL) has shown promising results in improving the cross-task generalization ability of meta-policies. However, current methods often lead to entangled task representations, in which each latent dimension is influenced by multiple causal factors that govern variations in environment dynamics and reward mechanisms. This entanglement can degrade generalization performance, particularly when multiple causal factors vary simultaneously across tasks. To address this limitation, we propose CAusally disentangled TAsk representation Learning (CATAL) method for COMRL that aims to improve the generalization ability of the meta-policy, where each latent dimension in the task representations aligns to a single causal factor.Theoretically, we show that under mild conditions, the task representations learned by CATAL are causally disentangled. Empirically, extensive results on multi-task MuJoCo benchmarks show that CATAL consistently outperforms existing COMRL baselines in both in-distribution and out-of-distribution generalization.
Shan Cong, Chao Yu 0004, Xiangyuan Lan
AAAI3
2026 Bolster Hallucination Detection via Prompt-Guided Data Augmentation
abstract
Large language models (LLMs) have garnered significant interest in AI community. Despite their impressive generation capabilities, they have been found to produce misleading or fabricated information, a phenomenon known as hallucinations. Consequently, hallucination detection has become critical to ensure the reliability of LLM-generated content. One primary challenge in hallucination detection is the scarcity of well-labeled datasets containing both truthful and hallucinated outputs. To address this issue, we introduce Prompt-guided data Augmented haLlucination dEtection (PALE), a novel framework that leverages prompt-guided responses from LLMs as data augmentation for hallucination detection. This strategy can generate both truthful and hallucinated data under prompt guidance at a relatively low cost. To more effectively evaluate the truthfulness of the sparse intermediate embeddings produced by LLMs, we introduce an estimation metric called the Contrastive Mahalanobis Score (CM Score). This score is based on modeling the distributions of truthful and hallucinated data in the activation space. CM Score employs a matrix decomposition approach to more accurately capture the underlying structure of these distributions. Importantly, our framework does not require additional human annotations, offering strong generalizability and practicality for real-world applications. Extensive experiments demonstrate that PALE achieves superior hallucination detection performance, outperforming the competitive baseline by a significant margin of 6.55%.
Wenyun Li 0001, Zheng Zhang 0006, Dongmei Jiang, Xiangyuan Lan
AAAI4
2026 X-SAM: From Segment Anything to Any Segmentation
abstract
Large Language Models (LLMs) demonstrate strong capabilities in broad knowledge representation, yet they are inherently deficient in pixel-level perceptual understanding. Although the Segment Anything Model (SAM) represents a significant advancement in visual-prompt-driven image segmentation, it exhibits notable limitations in multi-mask prediction and category-specific segmentation tasks, and it cannot integrate all segmentation tasks within a unified model architecture. To address these limitations, we present X-SAM, a streamlined Multimodal Large Language Model (MLLM) framework that extends the segmentation paradigm from segment anything to any segmentation. Specifically, we introduce a novel unified framework that enables more advanced pixel-level perceptual comprehension for MLLMs. Furthermore, we propose a new segmentation task, termed Visual GrounDed (VGD) segmentation, which segments all instance objects with interactive visual prompts and empowers MLLMs with visual grounded, pixel-wise interpretative capabilities. To enable effective training on diverse data sources, we present a unified training strategy that supports co-training across multiple datasets. Experimental results demonstrate that X-SAM achieves state-of-the-art performance on a wide range of image segmentation benchmarks, highlighting its efficiency for multimodal, pixel-level visual understanding.
Hao Wang 0050, Limeng Qiao, Zequn Jie, Chengjian Feng, Lin Ma 0002, Xiangyuan Lan, Xiaodan Liang
AAAI8
2026 DyToS: Budget-aware dynamic token scheduling for efficient multi-modal large language models
Yifei Xing 0001, Ruiping Wang 0001, Dongmei Jiang, Xiangyuan Lan
Neurocomputing7
2026 SelaVPR++: Towards Seamless Adaptation of Foundation Models for Efficient Place Recognition
abstract
Recent studies show that the visual place recognition (VPR) method using pre-trained visual foundation models can achieve promising performance. In our previous work, we propose a novel method to realize seamless adaptation of foundation models to VPR (SelaVPR). This method can produce both global and local features that focus on discriminative landmarks to recognize places for two-stage VPR by a parameter-efficient adaptation approach. Although SelaVPR has achieved competitive results, we argue that the previous adaptation is inefficient in training time and GPU memory usage, and the re-ranking paradigm is also costly in retrieval latency and storage usage. In pursuit of higher efficiency and better performance, we propose an extension of the SelaVPR, called SelaVPR++. Concretely, we first design a parameter-, time-, and memory-efficient adaptation method that uses lightweight multi-scale convolution (MultiConv) adapters to refine intermediate features from the frozen foundation backbone. This adaptation method does not back-propagate gradients through the backbone during training, and the MultiConv adapter facilitates feature interactions along the spatial axes and introduces proper local priors, thus achieving higher efficiency and better performance. Moreover, we propose an innovative re-ranking paradigm for more efficient VPR. Instead of relying on local features for re-ranking, which incurs huge overhead in latency and storage, we employ compact binary features for initial retrieval and robust floating-point (global) features for re-ranking. To obtain such binary features, we propose a similarity-constrained deep hashing method, which can be easily integrated into the VPR pipeline. Finally, we improve our training strategy and unify the training protocol of several common training datasets to merge them for better training of VPR models. Extensive experiments show that SelaVPR++ is highly efficient in training time, GPU memory usage, and retrieval latency (6000× faster than TransVPR), as well as outperforms the state-of-the-art methods by a large margin (ranks 1st on MSLS challenge leaderboard).
Xiangyuan Lan, Yunpeng Liu 0001, Yaowei Wang 0001, Chun Yuan 0003
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Toward Visual Grounding: A Survey
abstract
Visual Grounding, also known as Referring Expression Comprehension and Phrase Grounding, aims to ground the specific region(s) within the image(s) based on the given expression text. This task simulates the common referential relationships between visual and linguistic modalities, enabling machines to develop human-like multimodal comprehension capabilities. Consequently, it has extensive applications in various domains. However, since 2021, visual grounding has witnessed significant advancements, with emerging new concepts such as grounded pre-training, grounding multimodal LLMs, generalized visual grounding, and giga-pixel grounding, which have brought numerous new challenges. In this survey, we first examine the developmental history of visual grounding and provide an overview of essential background knowledge, including fundamental concepts and evaluation metrics. We systematically track and summarize the advancements, and then meticulously define and organize the various settings to standardize future research and ensure a fair comparison. In the dataset section, we compile a comprehensive list of current relevant datasets, conduct a fair comparative analysis, and provide ultimate performance prediction to inspire the development of new standard benchmarks. Additionally, we delve into numerous applications and highlight several advanced topics. Finally, we outline the challenges confronting visual grounding and propose valuable directions for future research, which may serve as inspiration for subsequent researchers. By extracting common technical details, this survey encompasses the representative work in each subtopic over the past decade. To the best of our knowledge, this paper represents the most comprehensive overview currently available in the field of visual grounding. This survey is designed to be suitable for both beginners and experienced researchers, serving as an invaluable resource for understanding key concepts and tracking the latest research developments.
Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang 0001, Changsheng Xu
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 AlignMamba-2: Enhancing multimodal fusion and sentiment analysis with modality-aware Mamba
Yan Li 0121, Yifei Xing 0001, Xiangyuan Lan, Xin Li 0034, Dongmei Jiang
Pattern Recognit.3
2026 HMVformer++: Hierarchical multi-view fusion transformer for efficient 3D human pose estimation
Kangkang Zhou, Xiangyuan Lan, Yu Shi 0003
Pattern Recognit.5
2026 Decoupled gradient-guided stratification for resource-efficient multi-modal data pruning
Yifei Xing 0001, Ruiping Wang 0001, Xiangyuan Lan, Yaowei Wang 0001
Pattern Recognit. Lett.5
2025 Transferable Adversarial Face Attack with Text Controlled Attribute
abstract
Traditional adversarial attacks typically produce adversarial examples under norm-constrained conditions, whereas unrestricted adversarial examples are free-form with semantically meaningful perturbations. Current unrestricted adversarial impersonation attacks exhibit limited control over adversarial face attributes and often suffer from low transferability. In this paper, we propose a novel Text Controlled Attribute Attack (TCA2) to generate photorealistic adversarial impersonation faces guided by natural language. Specifically, the category-level personal softmax vector is employed to precisely guide the impersonation attacks. Additionally, we propose both data and model augmentation strategies to achieve transferable attacks on unknown target models. Finally, a generative model, i.e, Style-GAN, is utilized to synthesize impersonated faces with desired attributes. Extensive experiments on two high-resolution face recognition datasets validate that our TCA2 method can generate natural text-guided adversarial impersonation faces with high transferability. We also evaluate our method on real-world face recognition systems, i.e, Face++ and Aliyun, further demonstrating the practical potential of our approach.
Wenyun Li 0001, Zheng Zhang 0006, Xiangyuan Lan, Dongmei Jiang
AAAI3
2025 DM-Adapter: Domain-Aware Mixture-of-Adapters for Text-Based Person Retrieval
abstract
Text-based person retrieval (TPR) has gained significant attention as a fine-grained and challenging task that closely aligns with practical applications. Tailoring CLIP to person domain is now a emerging research topic due to the abundant knowledge of vision-language pretraining, but challenges still remain during fine-tuning: (i) Previous full-model fine-tuning in TPR is computationally expensive and prone to overfitting.(ii) Existing parameter-efficient transfer learning (PETL) for TPR lacks of fine-grained feature extraction. To address these issues, we propose Domain-Aware Mixture-of-Adapters (DM-Adapter), which unifies Mixture-of-Experts (MOE) and PETL to enhance fine-grained feature representations while maintaining efficiency. Specifically, Sparse Mixture-of-Adapters is designed in parallel to MLP layers in both vision and language branches, where different experts specialize in distinct aspects of person knowledge to handle features more finely. To promote the router to exploit domain information effectively and alleviate the routing imbalance, Domain-Aware Router is then developed by building a novel gating function and injecting learnable domain-aware prompts. Extensive experiments show that our DM-Adapter achieves state-of-the-art performance, outperforming previous methods by a significant margin.
Zimo Liu, Xiangyuan Lan, Wenming Yang, Yaowei Li 0001, Qingmin Liao
AAAI3
2025 QER: Quantized Low-Rank Error Reconstructor for LLM Low-Bitwidth Quantization
abstract
Large Language Models (LLMs) have achieved remarkable success but face significant deployment challenges in cloud and edge environments due to their massive computational and storage requirements. Model quantization serves as a key solution to enhance the scalability and efficiency of LLMs within distributed cloud platforms. Existing Post-Training Quantization (PTQ) methods often exhibit suboptimal performance in low-bit settings. To further improve their precision, Quantization-Aware Training (QAT) combined with Low-Rank Adaptation (LoRA) has been explored for error correction. However, a critical issue is that the quantized base model and full-precision LoRA parameters suffer from precision mismatch, introducing additional errors during weight merging. To address these challenges, we propose a Quantized Low-rank Error Reconstructor (QER) for LLM low-bitwidth quantization. QER first enables lossless merging in low-bitwidth format by aligning the bitwidth of its low-rank parameters with the quantized base parameters, eliminating dequantization and requantization steps. Through this process, QER reconstructs original errors into two components: the quantization errors of QER parameters (i.e., quantized low-rank parameters) and potential overflow errors during low-bitwidth merging. These two errors are directly related to QER parameters, making them easier to optimize via gradient-based updates within an error-aware training framework. Requiring only 128 samples and 1 training epoch, QER demonstrates superior performance on LLaMA-1/2 families. In 4-bit quantization, compared to QLLM with error correction, QER reduces average perplexity by 13.8% (from 10.97 to 9.45) and improves average accuracy by 3.01 percentage points (from 51.84% to 54.85%) on LLaMA-1-7B. QER bridges the gap between quantization and low-rank adaptation, enabling efficient and accurate low-precision LLM deployment.
Shoukai Xu, Runhao Zeng, Xiangyuan Lan, Yaowei Wang 0001, Mingkui Tan
CloudCom6
2025 AlignMamba: Enhancing Multimodal Mamba with Local and Global Cross-modal Alignment
abstract
Cross-Modal alignment is crucial for multimodal representation fusion due to the inherent heterogeneity between modalities. While Transformer-Based methods have shown promising results in modeling inter-modal relationships, their quadratic computational complexity limits their applicability to long-sequence or large-scale data. Although recent Mamba-Based approaches achieve linear complexity, their sequential scanning mechanism poses fundamental challenges in comprehensively modeling cross-modal relationships. To address this limitation, we propose Align-Mamba, an efficient and effective method for multimodal fusion. Specifically, grounded in Optimal Transport, we introduce a local cross-modal alignment module that explicitly learns token-level correspondences between different modalities. Moreover, we propose a global cross-modal alignment loss based on Maximum Mean Discrepancy to implicitly enforce the consistency between different modal distributions. Finally, the unimodal representations after local and global alignment are passed to the Mamba backbone for further cross-modal interaction and multimodal fusion. Extensive experiments on complete and incomplete multimodal fusion tasks demonstrate the effectiveness and efficiency of the proposed method. For instance, on the CMU-MOSI dataset, AlignMamba improves classification accuracy by 0.9%, reduces GPU memory usage by 20.3%, and decreases inference time by 83.3%.
Yan Li 0121, Yifei Xing 0001, Xiangyuan Lan, Xin Li 0034, Dongmei Jiang
CVPR3
2025 Contrast Memory for Unsupervised Anomaly Detection
abstract
Real-world multivariate time series unsupervised anomaly detection is a challenging problem due to intricate temporal correlations. Recently, impressive progress have been made in tackling this issue through the design of large-scale models, facilitated by the growing model parameters. However, in resource-constrained scenarios such as ubiquitous computing and edge computing, the large-scale models suffer from issues like high parameter complexity and expensive training overheads. Existing methods can only strive for a direct tradeoff between model size and performance. To address this challenge, we propose DiMER (Diminutive Memory-Enhanced Reconstruction), a model with parameters of the order of 0.1M. In DiMER, we introduce a novel contrast memory mechanism to learn normal patterns with diminutive network and propose a temporal reconstruction loss nto add the autocorrelation information. In addition, we introduce a multi-space composite detection criterion, an anomaly score calculation that takes into account both memory space and data space. Extensive experiments on real-world datasets across various domains demonstrate that the proposed model achieves comparable or even superior performance to large-scale models while maintaining lightweight.
Jiahao Li 0007, Yiqiang Chen 0001, Yunbing Xing, Yang Gu 0001, Xiangyuan Lan
ICASSP5
2025 HYMAN: Hybrid Memory and Attention Network for Unsupervised Anomaly Detection
abstract
Detecting anomalies in unsupervised multivariate time series is challenging due to the intricate temporal patterns present in both local short-term and global long-term dependencies. Long short-term memory has achieved impressive results in this domain, yet it is gradually being supplemented by Transformers, due to limitations such as non-parallelization, gradient vanishing, and difficulty in focusing on local information. Leveraging the attention mechanism, Transformers can attend to all time steps simultaneously, effectively capturing local dependencies. However, they may often face challenges in efficiently modeling long-term dependencies, particularly in real-world scenarios. To address these issues, we propose the HYbrid Memory and Attention Network (HYMAN), which integrates memory and attention mechanisms together to model both global and local information. The attention captures short-term dependencies by focusing on temporal autocorrelation, while the memory stores and updates key historical patterns in global information, facilitating the learning of long-term dependencies. In contrast to previous approaches, HYMAN eliminates the need for auxiliary loss, simplifying the training by reducing the effort for coefficients tuning. In the inference phase, HYMAN introduces a novel anomaly scoring method that fuses features from both the temporal and latent spaces, offering high-performance detection compared to traditional methods that rely solely on reconstruction. Extensive experiments on real-world benchmark datasets demonstrate that HYMAN achieves state-of-the-art performance by leveraging the complementary strengths of memory and attention mechanisms.
Jiahao Li 0007, Yiqiang Chen 0001, Yunbing Xing, Yang Gu 0001, Xiangyuan Lan
ICASSP5
2025 EMMA: Empowering Multi-modal Mamba with Structural and Hierarchical Alignment
abstract
Mamba-based architectures have shown to be a promising new direction for deep learning models owing to their competitive performance and sub-quadratic deployment speed. However, current Mamba multi-modal large language models (MLLM) are insufficient in extracting visual features, leading to imbalanced cross-modal alignment between visual and textural latents, negatively impacting performance on multi-modal tasks. In this work, we propose Empowering Multi-modal Mamba with Structural and Hierarchical Alignment (EMMA), which enables the MLLM to extract fine-grained visual information. Specifically, we propose a pixel-wise alignment module to autoregressively optimize the learning and processing of spatial image-level features along with textual tokens, enabling structural alignment at the image level. In addition, to prevent the degradation of visual information during the cross-model alignment process, we propose a multi-scale feature fusion (MFF) module to combine multi-scale visual features from intermediate layers, enabling hierarchical alignment at the feature level. Extensive experiments are conducted across a variety of multi-modal benchmarks. Our model shows lower latency than other Mamba-based MLLMs and is nearly four times faster than transformer-based MLLMs of similar scale during inference. Due to better cross-modal alignment, our model exhibits lower degrees of hallucination and enhanced sensitivity to visual details, which manifests in superior performance across diverse multi-modal benchmarks. Code provided at https://github.com/xingyifei2016/EMMA.
Yifei Xing 0001, Xiangyuan Lan, Ruiping Wang 0001, Dongmei Jiang, Yaowei Wang 0001
ICLR2
2025 Open-Det: An Efficient Learning Framework for Open-Ended Detection
abstract
Open-Ended object Detection (OED) is a novel and challenging task that detects objects and generates their category names in a free-form manner, without requiring additional vocabularies during inference. However, the existing OED models, such as GenerateU, require large-scale datasets for training, suffer from slow convergence, and exhibit limited performance. To address these issues, we present a novel and efficient Open-Det framework, consisting of four collaborative parts. Specifically, Open-Det accelerates model training in both the bounding box and object name generation process by reconstructing the Object Detector and the Object Name Generator. To bridge the semantic gap between Vision and Language modalities, we propose a Vision-Language Aligner with V-to-L and L-to-V alignment mechanisms, incorporating with the Prompts Distiller to transfer knowledge from the VLM into VL-prompts, enabling accurate object name generation for the LLM. In addition, we design a Masked Alignment Loss to eliminate contradictory supervision and introduce a Joint Loss to enhance classification, resulting in more efficient training. Compared to GenerateU, Open-Det, using only 1.5% of the training data (0.077M vs. 5.077M), 20.8% of the training epochs (31 vs. 149), and fewer GPU resources (4 V100 vs. 16 A100), achieves even higher performance (+1.0% in APr). The source codes are available at: https://github.com/Med-Process/Open-Det.
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang
ICML4
2025 Online Preference-based Reinforcement Learning with Self-augmented Feedback from Large Language Model
Songjun Tu, Jingbo Sun 0001, Xiangyuan Lan, Dongbin Zhao
AAMAS4
2025 K-Space Bispectrum Steganography for Robust Unlearnable Data
Jiahao Li 0007, Yiqiang Chen 0001, Yunbing Xing, Yang Gu 0001, Xiangyuan Lan
ACM Multimedia5
2025 DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
abstract
Popular transformer detectors have achieved promising performance through query-based learning using attention mechanisms. However, the roles of existing decoder query types (e.g., content query and positional query) are still underexplored. These queries are generally predefined with a fixed number (fixed-query), which limits their flexibility. We find that the learning of these fixed-query is impaired by Recurrent Opposing in Teractions (ROT) between two attention operations: Self-Attention (query-to-query) and Cross-Attention (query-to-encoder), thereby degrading decoder efficiency. Furthermore, "query ambiguity" arises when shared-weight decoder layers are processed with both one-to-one and one-to-many label assignments during training, violating DETR's one-to-one matching principle. To address these challenges, we propose DS-Det, a more efficient detector capable of detecting a flexible number of objects in images. Specifically, we reformulate and introduce a new unified Single-Query paradigm for decoder modeling, transforming the fixed-query into flexible. Furthermore, we propose a simplified decoder framework through attention disentangled learning: locating boxes with Cross-Attention (one-to-many process), deduplicating predictions with Self-Attention (one-to-one process), addressing ''query ambiguity'' and ''ROT'' issues directly, and enhancing decoder efficiency. We further introduce a unified PoCoo loss that leverages box size priors to prioritize query learning on hard samples such as small objects. Extensive experiments across five different backbone models on COCO2017 and WiderPerson datasets demonstrate the general effectiveness and superiority of DS-Det. The source codes are available at https://github.com/Med-Process/DS-Det/.
Guiping Cao, Xiangyuan Lan, Wenjian Huang 0001, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001
ACM Multimedia2
2025 VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
abstract
With the rapid development of MLLMs, evaluating their visual capabilities has become increasingly crucial. Current benchmarks primarily fall into two main types: basic perception benchmarks,which focus on local details but lack deep reasoning (e.g., ''what is in the image?''), and mainstream reasoning benchmarks, which concentrate on prominent image elements but may fail to assess subtle clues requiring intricate analysis. However, profound visual understanding and complex reasoning depend more on interpreting subtle, inconspicuous local details than on perceiving salient, macro-level objects. These details, though occupying minimal image area, often contain richer, more critical information for robust analysis. To bridge this gap, we introduce the VER-Bench, a novel framework to evaluate MLLMs' ability to: 1) identify fine-grained visual clues, often occupying, on average, just 0.25% of the image area; 2) integrate these clues with world knowledge for complex reasoning. Comprising 374 carefully designed questions across Geospatial, Temporal, Situational, Intent, System State, and Symbolic reasoning, each question in VER-Bench is accompanied by structured evidence: visual clues and question-related reasoning derived from them. VER-Bench reveals current models' limitations in extracting subtle visual evidence and constructing evidence-based reasoning chains, highlighting the need to enhance models' capabilities in fine-grained visual evidence extraction, integration, and reasoning for genuine visual understanding and human-like analysis. The dataset is available at https://github.com/verbta/ACMMM-25-Materials.
Chenhui Qiang, Zhaoyang Wei, Xumeng Han, Siyao Li, Xiangyuan Lan, Jianbin Jiao, Zhenjun Han
ACM Multimedia6
2025 Towards Implicit Aggregation: Robust Image Representation for Place Recognition in the Transformer Era
abstract
Visual place recognition (VPR) is typically regarded as a specific image retrieval task, whose core lies in representing images as global descriptors. Over the past decade, dominant VPR methods (e.g., NetVLAD) have followed a paradigm that first extracts the patch features/tokens of the input image using a backbone, and then aggregates these patch features into a global descriptor via an aggregator. This backbone-plus-aggregator paradigm has achieved overwhelming dominance in the CNN era and remains widely used in transformer-based models. In this paper, however, we argue that a dedicated aggregator is not necessary in the transformer era, that is, we can obtain robust global descriptors only with the backbone. Specifically, we introduce some learnable aggregation tokens, which are prepended to the patch tokens before a particular transformer block. All these tokens will be jointly processed and interact globally via the intrinsic self-attention mechanism, implicitly aggregating useful information within the patch tokens to the aggregation tokens. Finally, we only take these aggregation tokens from the last output tokens and concatenate them as the global representation. Although implicit aggregation can provide robust global descriptors in an extremely simple manner, where and how to insert additional tokens, as well as the initialization of tokens, remains an open issue worthy of further exploration. To this end, we also propose the optimal token insertion strategy and token initialization method derived from empirical studies. Experimental results show that our method outperforms state-of-the-art methods on several VPR datasets with higher efficiency and ranks 1st on the MSLS challenge leaderboard. The code is available at https://github.com/lu-feng/image.
Canming Ye, Xiangyuan Lan, Yunpeng Liu 0001, Chun Yuan 0003
NeurIPS4
2025 Learning When to Think: Shaping Adaptive Reasoning in R1-Style Models via Multi-Stage RL
abstract
Large reasoning models (LRMs) are proficient at generating explicit, step-by-step reasoning sequences before producing final answers. However, such detailed reasoning can introduce substantial computational overhead and latency, particularly for simple problems. To address this over-thinking problem, we explore how to equip LRMs with adaptive thinking capabilities—enabling them to dynamically decide whether or not to engage in explicit reasoning based on problem complexity. Building on R1-style distilled models, we observe that inserting a simple ellipsis ("...") into the prompt can stochastically trigger either a thinking or no-thinking mode, revealing a latent controllability in the reasoning behavior. Leveraging this property, we propose AutoThink, a multi-stage reinforcement learning (RL) framework that progressively optimizes reasoning policies via stage-wise reward shaping. AutoThink learns to invoke explicit reasoning only when necessary, while defaulting to succinct responses for simpler tasks. Experiments on five mainstream mathematical benchmarks demonstrate that AutoThink achieves favorable accuracy–efficiency trade-offs compared to recent prompting and RL-based pruning methods. It can be seamlessly integrated into any R1-style model, including both distilled and further fine-tuned variants. Notably, AutoThink improves relative accuracy by 6.4\% while reducing token usage by 52\% on DeepSeek-R1-Distill-Qwen-1.5B, establishing a scalable and adaptive reasoning paradigm for LRMs. Project Page: https://github.com/ScienceOne-AI/AutoThink.
Songjun Tu, Xiangyu Tian, Linjing Li, Xiangyuan Lan, Dongbin Zhao
NeurIPS6
2025 Generic Scene Graph Generation Model with Hierarchical Prompt Learning
Xuhan Zhu, Yifei Xing 0001, Ruiping Wang 0001, Yaowei Wang 0001, Xiangyuan Lan
Int. J. Comput. Vis.5
2025 ClickTrack: Towards real-time interactive single object tracking
Kuiran Wang, Xuehui Yu, Wenwen Yu, Guorong Li, Xiangyuan Lan, Qixiang Ye, Jianbin Jiao, Zhenjun Han
Pattern Recognit.5
2025 UP-Person: Unified Parameter-Efficient Transfer Learning for Text-Based Person Retrieval
abstract
Text-based Person Retrieval (TPR) as a multi-modal task, which aims to retrieve the target person from a pool of candidate images given a text description, has recently garnered considerable attention due to the progress of contrastive visual-language pre-trained model. Prior works leverage pre-trained CLIP to extract person visual and textual features and fully fine-tune the entire network, which have shown notable performance improvements compared to uni-modal pre-training models. However, full-tuning a large model is prone to overfitting and hinders the generalization ability. In this paper, we propose a novelUnifiedParameter-Efficient Transfer Learning (PETL) method for Text-basedPersonRetrieval (UP-Person) to thoroughly transfer the multi-modal knowledge from CLIP. Specifically, UP-Person simultaneously integrates three lightweight PETL components including Prefix, LoRA and Adapter, where Prefix and LoRA are devised together to mine local information with task-specific information prompts, and Adapter is designed to adjust global feature representations. Additionally, two vanilla submodules are optimized to adapt to the unified architecture of TPR. For one thing, S-Prefix is proposed to boost attention of prefix and enhance the gradient propagation of prefix tokens, which improves the flexibility and performance of the vanilla prefix. For another thing, L-Adapter is designed in parallel with layer normalization to adjust the overall distribution, which can resolve conflicts caused by overlap and interaction among multiple submodules. Extensive experimental results demonstrate that our UP-Person achieves state-of-the-art results across various person retrieval datasets, including CUHK-PEDES, ICFG-PEDES and RSTPReid while merely fine-tuning 4.7% parameters. Code is available at https://github.com/Liu-Yating/UP-Person.
Yaowei Li 0001, Xiangyuan Lan, Wenming Yang, Zimo Liu, Qingmin Liao
IEEE Trans. Circuits Syst. Video Technol.3
2025 Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
abstract
Small Object Detection (SOD) poses significant challenges due to limited information and the model's low class prediction score. While Transformer-based detectors have shown promising performance, their potential for SOD remains largely unexplored. In typical DETR-like frameworks, the CNN backbone network, specialized in aggregating local information, struggles to capture the necessary contextual information for SOD. The multiple attention layers in the Transformer Encoder face difficulties in effectively attending to small objects and can also lead to blurring of features. Furthermore, the model's lower class prediction score of small objects compared to large objects further increases the difficulty of SOD. To address these challenges, we introduce a novel approach calledCross-DINO. This approach incorporates the deep MLP network to aggregate initial feature representations with both short and long range information for SOD. Then, a new Cross Coding Twice Module (CCTM) is applied to integrate these initial representations to the Transformer Encoder feature, enhancing the details of small objects. Additionally, we introduce a new kind of soft label named Category-Size (CS), integrating the Category and Size of objects. By treating CS as new ground truth, we propose a new loss function called Boost Loss to improve the class prediction score of the model. Extensive experimental results on COCO, WiderPerson, VisDrone, AI-TOD, and SODA-D datasets demonstrate that Cross-DINO efficiently improves the performance of DETR-like models on SOD. Specifically, our model achieves36.4%AP$_{S}$on COCO for SOD with only 45M parameters, outperforming the DINO by+4.4%AP$_{S}$(36.4% vs. 32.0%) with fewer parameters and FLOPs, under 12 epochs training setting.
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001
IEEE Trans. Multim.3
2024 Deep Homography Estimation for Visual Place Recognition
abstract
Visual place recognition (VPR) is a fundamental task for many applications such as robot localization and augmented reality. Recently, the hierarchical VPR methods have received considerable attention due to the trade-off between accuracy and efficiency. They usually first use global features to retrieve the candidate images, then verify the spatial consistency of matched local features for re-ranking. However, the latter typically relies on the RANSAC algorithm for fitting homography, which is time-consuming and non-differentiable. This makes existing methods compromise to train the network only in global feature extraction. Here, we propose a transformer-based deep homography estimation (DHE) network that takes the dense feature map extracted by a backbone network as input and fits homography for fast and learnable geometric verification. Moreover, we design a re-projection error of inliers loss to train the DHE network without additional homography labels, which can also be jointly trained with the backbone network to help it extract the features that are more suitable for local matching. Extensive experiments on benchmark datasets show that our method can outperform several state-of-the-art methods. And it is more than one order of magnitude faster than the mainstream hierarchical VPR methods using RANSAC. The code is released at https://github.com/Lu-Feng/DHE-VPR.
Shuting Dong, Bingxi Liu 0001, Xiangyuan Lan, Dongmei Jiang, Chun Yuan 0003
AAAI5
2024 Hierarchical Prompt Learning for Scene Graph Generation
Xuhan Zhu, Yifei Xing 0001, Ruiping Wang 0001, Yaowei Wang 0001, Xiangyuan Lan
BMVC5
2024 CricaVPR: Cross-Image Correlation-Aware Representation Learning for Visual Place Recognition
abstract
Over the past decade, most methods in visual place recognition (VPR) have used neural networks to produce feature representations. These networks typically produce a global representation of a place image using only this image itself and neglect the cross-image variations (e.g. viewpoint and illumination), which limits their robustness in challenging scenes. In this paper, we propose a robust global representation method with cross-image correlation awareness for VPR, named CricaVPR. Our method uses the attention mechanism to correlate multiple images within a batch. These images can be taken in the same place with different conditions or viewpoints, or even captured from different places. Therefore, our method can utilize the cross-image variations as a cue to guide the representation learning, which ensures more robust features are produced. To further facilitate the robustness, we propose a multi-scale convolution-enhanced adaptation method to adapt pre-trained visual foundation models to the VPR task, which introduces the multi-scale local information to further enhance the cross-image correlation-aware representation. Experimental results show that our method out-performs state-of-the-art methods by a large margin with significantly less training time. The code is released at https://github.com/Lu-Feng/CricaVPR.
Xiangyuan Lan, Dongmei Jiang, Yaowei Wang 0001, Chun Yuan 0003
CVPR2
2024 Cascade Memory for Unsupervised Anomaly Detection
abstract
Unsupervised anomaly detection is to detect previously unseen rare samples without any prior knowledge about them. With the emergence of deep learning, many methods employ normal data reconstruction to train detection models, which is expected to yield relatively large errors when reconstructing anomalies. However, recent studies find that anomalies can be overgeneralized, resulting in reconstruction errors as small as normal samples. In this paper, we examine the anomaly overgeneralization problem and propose global semantic information learning. Normal and anomalous samples may share the same local feature such as textures, edges, and corners, but have separability at the global semantic level. To address this, we propose a novel cascade memory architecture designed to capture global semantic information in the latent space and introduce a configurable sparsification and random forgetting mechanism. Our proposed method achieves state-of-the-art experimental results on different public benchmarks, without the introduction of any additional auxiliary loss terms. The code is available at https://github.com/LiJiahao-Alex/Cascade-Memory.
Jiahao Li 0007, Yiqiang Chen 0001, Yunbing Xing, Yang Gu 0001, Xiangyuan Lan
ECAI5
2024 Towards Seamless Adaptation of Pre-trained Models for Visual Place Recognition
abstract
Recent studies show that vision models pre-trained in generic visual learning tasks with large-scale data can provide useful feature representations for a wide range of visual perception problems. However, few attempts have been made to exploit pre-trained foundation models in visual place recognition (VPR). Due to the inherent difference in training objectives and data between the tasks of model pre-training and VPR, how to bridge the gap and fully unleash the capability of pre-trained models for VPR is still a key issue to address. To this end, we propose a novel method to realize seamless adaptation of pre-trained models for VPR. Specifically, to obtain both global and local features that focus on salient landmarks for discriminating places, we design a hybrid adaptation method to achieve both global and local adaptation efficiently, in which only lightweight adapters are tuned without adjusting the pre-trained model. Besides, to guide effective adaptation, we propose a mutual nearest neighbor local feature loss, which ensures proper dense local features are produced for local matching and avoids time-consuming spatial verification in re-ranking. Experimental results show that our method outperforms the state-of-the-art methods with less training data and training time, and uses about only 3% retrieval runtime of the two-stage VPR methods with RANSAC-based spatial verification. It ranks 1st on the MSLS challenge leaderboard (at the time of submission). The code is released at https://github.com/Lu-Feng/SelaVPR.
Xiangyuan Lan, Shuting Dong, Yaowei Wang 0001, Chun Yuan 0003
ICLR3
2024 Revisiting Context Aggregation for Image Matting
abstract
Traditional studies emphasize the significance of context information in improving matting performance. Consequently, deep learning-based matting methods delve into designing pooling or affinity-based context aggregation modules to achieve superior results. However, these modules cannot well handle the context scale shift caused by the difference in image size during training and inference, resulting in matting performance degradation. In this paper, we revisit the context aggregation mechanisms of matting networks and find that a basic encoder-decoder network without any context aggregation modules can actually learn more universal context aggregation, thereby achieving higher matting performance compared to existing methods. Building on this insight, we present AEMatter, a matting network that is straightforward yet very effective. AEMatter adopts a Hybrid-Transformer backbone with appearance-enhanced axis-wise learning (AEAL) blocks to build a basic network with strong context aggregation learning capability. Furthermore, AEMatter leverages a large image training strategy to assist the network in learning context aggregation from data. Extensive experiments on five popular matting datasets demonstrate that the proposed AEMatter outperforms state-of-the-art matting methods by a large margin. The source code is available at https://github.com/aipixel/AEMatter.
Qinglin Liu, Xiaoqian Lv, Quanling Meng, Zonglin Li 0004, Xiangyuan Lan, Shuo Yang 0006, Shengping Zhang, Liqiang Nie
ICML5
2024 MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001
IJCAI3
2024 Calibration for Long-tailed Scene Graph Generation
abstract
Miscalibrated models tend to be unreliable and insecure for downstream applications. In this work, we attempt to highlight and remedy miscalibration in current scene graph generation (SGG) models, which has been overlooked by previous works. We discover that obtaining well-calibrated models for SGG is more challenging than conventional calibration settings, as long-tailed SGG training data exacerbates miscalibration with overconfidence in head classes and underconfidence in tail classes. We further analyze which components are explicitly impacted by the long-tailed data during optimization, thereby exacerbating miscalibration and unbalanced learning, including biased parameters, deviated boundaries, and distorted target distribution. To address the above issues, we propose the Compositional Optimization Calibration (COC) method, comprising three modules: i. A parameter calibration module that utilizes a hyperspherical classifier to eliminate the bias introduced by biased parameters. ii. A boundary calibration module that disperses features of majority classes to consolidate the decision boundaries of minority classes and mitigate deviated boundaries. iii. A target distribution calibration module that addresses distorted target distribution, leverages within-triplet prior to guide confidence-aware and label-aware target calibration, and applies curriculum regulation to constrain learning focus from easy to hard classes. Extensive evaluation on popular benchmarks demonstrates the effectiveness of our proposed method in improving model calibration and resolving unbalanced learning for long-tailed SGG. Finally, our proposed method performs best on model calibration compared to different types of calibration methods and achieves state-of-the-art trade-off performance on balanced SGG learning.
Xuhan Zhu, Yifei Xing 0001, Ruiping Wang 0001, Yaowei Wang 0001, Xiangyuan Lan
ACM Multimedia5
2024 SuperVLAD: Compact and Robust Image Descriptors for Visual Place Recognition
abstract
Visual place recognition (VPR) is an essential task for multiple applications such as augmented reality and robot localization. Over the past decade, mainstream methods in the VPR area have been to use feature representation based on global aggregation, as exemplified by NetVLAD. These features are suitable for large-scale VPR and robust against viewpoint changes. However, the VLAD-based aggregation methods usually learn a large number of (e.g., 64) clusters and their corresponding cluster centers, which directly leads to a high dimension of the yielded global features. More importantly, when there is a domain gap between the data in training and inference, the cluster centers determined on the training set are usually improper for inference, resulting in a performance drop. To this end, we first attempt to improve NetVLAD by removing the cluster center and setting only a small number of (e.g., only 4) clusters. The proposed method not only simplifies NetVLAD but also enhances the generalizability across different domains. We name this method SuperVLAD. In addition, by introducing ghost clusters that will not be retained in the final output, we further propose a very low-dimensional 1-Cluster VLAD descriptor, which has the same dimension as the output of GeM pooling but performs notably better. Experimental results suggest that, when paired with a transformer-based backbone, our SuperVLAD shows better domain generalization performance than NetVLAD with significantly fewer parameters. The proposed method also surpasses state-of-the-art methods with lower feature dimensions on several benchmark datasets. The code is available at https://github.com/lu-feng/SuperVLAD.
Xinyao Zhang 0001, Canming Ye, Shuting Dong, Xiangyuan Lan, Chun Yuan 0003
NeurIPS6
2024 High-Resolution Image Harmonization with Adaptive-Interval Color Transformation
abstract
Existing high-resolution image harmonization methods typically rely on global color adjustments or the upsampling of parameter maps. However, these methods ignore local variations, leading to inharmonious appearances. To address this problem, we propose an Adaptive-Interval Color Transformation method (AICT), which predicts pixel-wise color transformations and adaptively adjusts the sampling interval to model local non-linearities of the color transformation at high resolution. Specifically, a parameter network is first designed to generate multiple position-dependent 3-dimensional lookup tables (3D LUTs), which use the color and position of each pixel to perform pixel-wise color transformations. Then, to enhance local variations adaptively, we separate a color transform into a cascade of sub-transformations using two 3D LUTs to achieve the non-uniform sampling intervals of the color transform. Finally, a global consistent weight learning method is proposed to predict an image-level weight for each color transform, utilizing global information to enhance the overall harmony. Extensive experiments demonstrate that our AICT achieves state-of-the-art performance with a lightweight architecture. The code is available at https://github.com/aipixel/AICT.
Quanling Meng, Qinglin Liu, Zonglin Li 0004, Xiangyuan Lan, Shengping Zhang, Liqiang Nie
NeurIPS4
2024 Local context attention learning for fine-grained scene graph generation
abstract
Fine-grained scene graph generation aims to parse the objects and their fine-grained relationships within scenes. Despite the significant progress in recent years, their performance is still limited by two major issues: (1) ambiguous perception under a global view; (2) the lack of reliable, fine-grained annotations. We argue that understanding the local context is important in addressing the two issues. However, previous works often overlook it, which limits their effectiveness in fine-grained scene graph generation. To tackle this challenge, we introduce a Local-context Attention Learning method that concentrates on local context and can generate high-reliability, fine-grained annotations. It comprises two components: (1) The Fine-grained Location Attention Network (FLAN), a multi-branch network that encompasses global and local branches, can attend to local informative context and perceive granularity levels in different regions, thereby adaptively enhancing the learning of fine-grained locations. (2) The Fine-grained Location Label Transfer (FLLT) method identifies coarse-grained labels inconsistent with the local context and determines which labels should be transferred through the global confidence thresholding strategy, finally transferring them to reliable local context-consistent fine-grained ones. Experiments conducted on the Visual Genome, OpenImage, and GQA-200 datasets show that the proposed methods achieve significant improvements on the fine-grained scene graph generation task. By addressing the challenge mentioned above, our method also achieves state-of-the-art performances on the three datasets.
Xuhan Zhu, Ruiping Wang 0001, Xiangyuan Lan, Yaowei Wang 0001
Pattern Recognit.3
2024 Efficient Image Classification via Structured Low-Rank Matrix Factorization Regression
abstract
In real-world applications involving sparse coding and low-rank matrix recovery problems, linear regression methods usually struggle to effectively capture the structured correlations present in data matrices. This limitation arises from representation approaches that treat images as vectors and handle testing samples individually, overlooking these correlations. To address these challenges, we propose a novel approach that leverages the low-rank property to capture the global and intrinsic structure of residual and coefficient matrices, departing from the assumption of independent and identically distributed (I.I.D) data. Our method introduces nonconvex and nonsmooth low-rank matrix regression models guided by the extended matrix variate power exponential distribution (M.P.E.D). By incorporating factorization strategies into the regression coefficient matrix and utilizing the Schatten-$p$norm with three distinct values of$p$, we enhance computational efficiency. Our formulation enables efficient subproblem solving through the introduction of auxiliary variables and the use of singular value threshold operators. We achieve closed-form solutions using the proposed multi-variable alternating direction method of multipliers (ADMM). Theoretical analysis establishes the local convergence properties and computational complexity of our optimization algorithm. Furthermore, we conduct numerical experiments on various image datasets, including face, object, and digital, to demonstrate the superior performance and computational efficiency of our methods compared to several related regression approaches. The source codes for our method are available athttps://github.com/ZhangHengMin/TIFS_SLRMFR.
Hengmin Zhang, Jian Yang 0003, Jianjun Qian, Guangwei Gao, Xiangyuan Lan, Zhiyuan Zha, Bihan Wen
IEEE Trans. Inf. Forensics Secur.5
2024 Limb-Aware Virtual Try-On Network With Progressive Clothing Warping
abstract
Image-based virtual try-on aims to transfer an in-shop clothing image to a person image. Most existing methods adopt a single global deformation to perform clothing warping directly, which lacks fine-grained modeling of in-shop clothing and leads to distorted clothing appearance. In addition, existing methods usually fail to generate limb details well because they are limited by the used clothing-agnostic person representation without referring to the limb textures of the person image. To address these problems, we propose Limb-aware Virtual Try-on Network named PL-VTON, which performs fine-grained clothing warping progressively and generates high-quality try-on results with realistic limb details. Specifically, we present Progressive Clothing Warping (PCW) that explicitly models the location and size of in-shop clothing and utilizes a two-stage alignment strategy to progressively align the in-shop clothing with the human body. Moreover, a novel gravity-aware loss that considers the fit of the person wearing clothing is adopted to better handle the clothing edges. Then, we design Person Parsing Estimator (PPE) with a non-limb target parsing map to semantically divide the person into various regions, which provides structural constraints on the human body and therefore alleviates texture bleeding between clothing and body regions. Finally, we introduce Limb-aware Texture Fusion (LTF) that focuses on generating realistic details in limb regions, where a coarse try-on result is first generated by fusing the warped clothing image with the person image, then limb textures are further fused with the coarse result under limb-aware guidance to refine limb details. Extensive experiments demonstrate that our PL-VTON outperforms the state-of-the-art methods both qualitatively and quantitatively.
Shengping Zhang, Weigang Zhang, Xiangyuan Lan, Hongxun Yao, Qingming Huang
IEEE Trans. Multim.4
2023 Strip-MLP: Efficient Token Interaction for Vision MLP
abstract
Token interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token interaction on the spatial dimension is highly dependent on the spatial resolution of the feature maps, which limits the model’s expressive ability, especially in deep layers where the feature are down-sampled to a small spatial size. To address this issue, we present a novel method called Strip-MLP to enrich the token interaction power in three ways. Firstly, we introduce a new MLP paradigm called Strip MLP layer that allows the token to interact with other tokens in a cross-strip manner, enabling the tokens in a row (or column) to contribute to the information aggregations in adjacent but different strips of rows (or columns). Secondly, a Cascade Group Strip Mixing Module (CGSMM) is proposed to overcome the performance degradation caused by small spatial feature size. The module allows tokens to interact more effectively in the manners of within-patch and cross-patch, which is independent to the feature spatial size. Finally, based on the Strip MLP layer, we propose a novel Local Strip Mixing Module (LSMM) to boost the token interaction power in the local region. Extensive experiments demonstrate that Strip-MLP significantly improves the performance of MLP-based models on small datasets and obtains comparable or even better results on ImageNet. In particular, Strip-MLP models achieve higher average Top-1 accuracy than existing MLP-based models by +2.44% on Caltech-101 and +2.16% on CIFAR-100. The source codes will be available at https://github.com/Med-Process/Strip_MLP.
Guiping Cao, Shengda Luo, Wenjian Huang 0001, Xiangyuan Lan, Dongmei Jiang, Yaowei Wang 0001, Jianguo Zhang 0001
ICCV4
2023 Towards Adaptable Graph Representation Learning: An Adaptive Multi-Graph Contrastive Transformer
abstract
Significant progress has been made in graph representation learning in recent years. However, most of these methods model spatial relationships via predefined graphs or decouple spatial-temporal representations, which limits the generalization and effectiveness of the model. To address these issues, we introduce an adaptive multi-graph contrastive transformer (AMGCT) for general spatial-temporal graph representation learning. Specifically, we first propose adaptive multi-graph contrastive learning (AMGCL). Without any expert knowledge, AMGCL can gradually generate adaptive spatial graphs with different topologies to learn spatial representations from different views. Cross-graph contrastive learning further explores potential correlations between different views, making each view's features more discriminative. In addition, to avoid insufficient interaction caused by decoupling spatial-temporal information in existing methods, we design a coupled graph transformer (CGT) to consider spatial relationships at each stage of temporal modeling, explore complementary information between spatial and temporal domains, and obtain more compact spatial-temporal representations. Experimental results on two different spatial-temporal graph datasets and tasks demonstrate that the proposed method achieves excellent performance.
Yan Li 0121, Liang Zhang 0042, Xiangyuan Lan, Dongmei Jiang
ACM Multimedia3
2023 Robust Tracking via Uncertainty-Aware Semantic Consistency
abstract
Robust tracking has a variety of practical applications. Despite many years of progress, it is still a difficult problem due to enormous uncertainties in real-world scenes. To address this issue, we propose a robust anchor-free based tracking model with uncertainty estimation. Within the model, a new data-driven uncertainty estimation strategy is proposed to generate uncertainty-aware features with promising discriminative and descriptive power. Then, a simple yet effective pyramid-wise cross correlation operation is constructed to extract multi-scale semantic features that provide rich correlation information for uncertainty-aware estimation and thus enhances the tracking robustness. Finally, a semantic consistency checking branch is designed to further estimate uncertainty of output results from the classification and regression branches by adaptively generating semantically consistent labels. Experiments on six benchmarks (i.e., OTB100, VOT2018, VOT2020, TrackingNet, GOT-10K and LaSOT) show the competing performance of our tracker with 130 FPS.
Jie Ma 0006, Xiangyuan Lan, Bineng Zhong 0001, Guorong Li, Zhenjun Tang, Xianxian Li, Rongrong Ji
IEEE Trans. Circuits Syst. Video Technol.2
2023 Detecting and Tracking of Multiple Mice Using Part Proposal Networks
abstract
The study of mouse social behaviors has been increasingly undertaken in neuroscience research. However, automated quantification of mouse behaviors from the videos of interacting mice is still a challenging problem, where object tracking plays a key role in locating mice in their living spaces. Artificial markers are often applied for multiple mice tracking, which are intrusive and consequently interfere with the movements of mice in a dynamic environment. In this article, we propose a novel method to continuously track several mice and individual parts without requiring any specific tagging. First, we propose an efficient and robust deep-learning-based mouse part detection scheme to generate part candidates. Subsequently, we propose a novel Bayesian-inference integer linear programming (BILP) model that jointly assigns the part candidates to individual targets with necessary geometric constraints while establishing pair-wise association between the detected parts. There is no publicly available dataset in the research community that provides a quantitative test bed for part detection and tracking of multiple mice, and we here introduce a new challenging Multi-Mice PartsTrack dataset that is made of complex behaviors. Finally, we evaluate our proposed approach against several baselines on our new datasets, where the results show that our method outperforms the other state-of-the-art approaches in terms of accuracy. We also demonstrate the generalization ability of the proposed approach on tracking zebra and locust.
Zheheng Jiang, Long Chen 0019, Xiangrong Zhang, Xiangyuan Lan, Danny Crookes, Ming-Hsuan Yang 0001, Huiyu Zhou 0001
IEEE Trans. Neural Networks Learn. Syst.6
2022 Convolution by Multiplication: Accelerated Two- Stream Fourier Domain Convolutional Neural Network for Facial Expression Recognition
abstract
Facial expression plays an important role in human communication as a type of nonverbal language and has been widely used in various areas such as psychology, human-computer interaction and robotics. Nowadays, convolutional neural network is a promising approach for facial expression recognition. However, convolutional layers can be time-consuming and computationally expensive because a large number of parameters participate in the calculations and need to be updated during training. To improve the performance of deep neural network in facial expression recognition and accelerate training and calculation, we propose a novel framework which adopts efficient element-wise multiplication to replace traditional convolution. To disentangle reliable feature representation for more effective recognition and further enhance the recognition performance while maintaining the efficiency, we propose a representation scheme which can retain informative feature components while removing unreliable ones in Fourier domain based on the proposed multiplication framework. Extensive comparison and ablation studies are conducted on several benchmark datasets, which shows the efficiency and effectiveness of the proposed model.
Xingming Zhang 0001, Xiangyuan Lan, Haoxiang Wang 0002
IEEE Trans. Circuits Syst. Video Technol.3
2022 Continuous Prediction of Lower-Limb Kinematics From Multi-Modal Biomedical Signals
abstract
The fast-growing techniques of measuring and fusing multi-modal biomedical signals enable advanced motor intent decoding schemes of lower-limb exoskeletons, meeting the increasing demand for rehabilitative or assistive applications of take-home healthcare. Challenges of exoskeletons’ motor intent decoding schemes remain in making a continuous prediction to compensate for the hysteretic response caused by mechanical transmission. In this paper, we solve this problem by proposing an ahead-of-time continuous prediction of lower-limb kinematics, with the prediction of knee angles during level walking as a case study. Firstly, an end-to-end kinematics prediction network(KinPreNet),1consisting of a feature extractor and an angle predictor, is proposed and experimentally compared with features and methods traditionally used in ahead-of-time prediction of gait phases. Secondly, inspired by the electromechanical delay(EMD), we further explore our algorithm’s capability of compensating response delay of mechanical transmission by validating the performance of the different sections of prediction time. And we experimentally reveal the time boundary of compensating the hysteretic response. Thirdly, a comparison of employing EMG signals or not is performed to reveal the EMG and kinematic signals’ collaborated contributions to the continuous prediction. During the experiments, EMG signals of nine muscles and knee angles calculated from inertial measurement unit (IMU) signals are recorded from ten healthy subjects. Our algorithm can predict knee angles with the averaged RMSE of 3.98 deg which is better than the 15.95-deg averaged RMSE of utilizing the traditional methods of ahead-of-time prediction. The best prediction time is in the interval of 27ms and 108ms. To the best of our knowledge, this is the first study of continuously predicting lower-limb kinematics in an ahead-of-time manner based on the electromechanical delay (EMD).
Chunzhi Yi, Feng Jiang 0001, Shengping Zhang, Hao Guo 0015, Chifu Yang, Zhen Ding, Baichun Wei, Xiangyuan Lan, Huiyu Zhou 0001
IEEE Trans. Circuits Syst. Video Technol.8
2022 Cohesive Multi-Modality Feature Learning and Fusion for COVID-19 Patient Severity Prediction
abstract
The outbreak of coronavirus disease (COVID-19) has been a nightmare to citizens, hospitals, healthcare practitioners, and the economy in 2020. The overwhelming number of confirmed cases and suspected cases put forward an unprecedented challenge to the hospital's capacity of management and medical resource distribution. To reduce the possibility of cross-infection and attend a patient according to his severity level, expertly diagnosis and sophisticated medical examinations are often required but hard to fulfil during a pandemic. To facilitate the assessment of a patient's severity, this paper proposes a multi-modality feature learning and fusion model for end-to-end covid patient severity prediction using the blood test supported electronic medical record (EMR) and chest computerized tomography (CT) scan images. To evaluate a patient's severity by the co-occurrence of salient clinical features, the High-order Factorization Network (HoFN) is proposed to learn the impact of a set of clinical features without tedious feature engineering. On the other hand, an attention-based deep convolutional neural network (CNN) using pre-trained parameters are used to process the lung CT images. Finally, to achieve cohesion of cross-modality representation, we design a loss function to shift deep features of both-modality into the same feature space which improves the model's performance and robustness when one modality is absent. Experimental results demonstrate that the proposed multi-modality feature learning and fusion model achieves high performance in an authentic scenario.
Jinzhao Zhou, Xingming Zhang 0001, Ziwei Zhu 0005, Xiangyuan Lan, Lunkai Fu, Haoxiang Wang 0002, Hanchun Wen
IEEE Trans. Circuits Syst. Video Technol.4
2022 Learning Temporal Similarity of Remote Photoplethysmography for Fast 3D Mask Face Presentation Attack Detection
abstract
To detect 3D mask face presentation attack, remote Photoplethysmography (rPPG), a biomedical technique that measures the heartbeat signal remotely with a normal RGB camera, is adopted as a robust liveness cue. Although existing rPPG-based solutions exhibit strong performance in experiments, the required observation time is too long (10-12 seconds) to be user-friendly in real applications such as E-payment and smartphone unlock. To shorten the observation time (within 1-second), we propose a fast rPPG-based 3D mask presentation attack detection (PAD) method by analyzing the similarity of rPPG signals in the time domain. In particular, based on facial and background local rPPG signals, we design a set of temporal similarity features to investigate the robust properties of rPPG shape and phase. Following the same direction, we refine the traditional rPPG extractor into a learnable network to cooperate with our TSrPPG feature for better robustness. An effective but lightweight spatiotemporal convolution network is constructed with a self-supervised learning strategy, aiming at enhancing the consistency of genuine facial rPPG signals and reducing the correlation of rPPG signals on masked faces. Extensive experiments are conducted on 3DMAD, HKBU-MARs V1+ and V2+, and CSMAD, which totally involve 18772 short-term video slots with a large number of real-world variations, in terms of mask type, mask transmittance, lighting condition, recording device, resolution of facial region, and compression configuration. Our proposed method persists the good performance of rPPG-based solution with only 1-second observation and outperforms the state-of-the-art competitors on discriminability and generalizability. Evaluations on prints attack, display attack, and disguise attacks with transparent masks, make-up and tattoo further exhibit its potential on handling a wider variety of attacks. To our best knowledge, this is the first work that addresses the length of observation time issue of rPPG-based 3D mask PAD.
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.2
2021 Driver distraction detection using capsule network
Deepak Kumar Jain 0001, Rachna Jain, Xiangyuan Lan, Yash Upadhyay, Anuj Thareja
Neural Comput. Appl.3
2021 Multi-Channel Remote Photoplethysmography Correspondence Feature for 3D Mask Face Presentation Attack Detection
abstract
With the advancement of 3D printing technologies, 3D mask presentation attack becomes a critical challenge in face recognition. To tackle the 3D mask presentation attack detection (PAD), remote Photoplethysmography (rPPG) is employed as an intrinsic detection cue which is independent of the mask material and appearance quality. Although the effectiveness of existing rPPG-based methods has been verified, they may not be robust enough when rPPG signals are contaminated by noise. To identify the heartbeat information from the noisy raw rPPG signals, we propose a new 3D mask PAD feature, multi-channel rPPG correspondence feature (MCCFrPPG) with the global noise-aware template learning and verification framework. To further boost the discriminability, temporal variation of the rPPG signal is considered and extracted through the multi-channel time-frequency analysis scheme. This paper also extends HKBU-MARs V2 dataset with more customized high-quality masks and increases the number of videos by two times. Comprehensive experiments were performed on existing 3D mask datasets and the extended HKBU-MARs V2+, which totally covers 3 types of masks, 12 different light settings and 6 cameras. The results not only justify the effectiveness and robustness of the proposed MCCFrPPG on 3D mask attacks but also indicate its potential on handling the replay attack with camera motion and dim light.
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.2
2021 SecureFace: Face Template Protection
abstract
It has been shown that face images can be reconstructed from their representations (templates). We propose a randomized CNN to generate protected face biometric templates given the input face image and a user-specific key. The use of user-specific keys introduces randomness to the secure template and hence strengthens the template security. To further enhance the security of the templates, instead of storing the key, we store a secure sketch that can be decoded to generate the key with genuine queries submitted to the system. We have evaluated the proposed protected template generation method using three benchmarking datasets for the face (FRGC v2.0, CFP, and IJB-A). The experimental results justify that the protected template generated by the proposed method are non-invertible and cancellable, while preserving the verification performance.
Guangcan Mai, Kai Cao 0001, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.3
2021 Spatial-temporal Regularized Multi-modality Correlation Filters for Tracking with Re-detection
abstract
The development of multi-spectrum image sensing technology has brought great interest in exploiting the information of multiple modalities (e.g., RGB and infrared modalities) for solving computer vision problems. In this article, we investigate how to exploit information from RGB and infrared modalities to address two important issues in visual tracking: robustness and object re-detection. Although various algorithms that attempt to exploit multi-modality information in appearance modeling have been developed, they still face challenges that mainly come from the following aspects: (1) the lack of robustness to deal with large appearance changes and dynamic background, (2) failure in re-capturing the object when tracking loss happens, and (3) difficulty in determining the reliability of different modalities. To address these issues and perform effective integration of multiple modalities, we propose a new tracking-by-detection algorithm called Adaptive Spatial-temporal Regulated Multi-Modality Correlation Filter. Particularly, an adaptive spatial-temporal regularization is imposed into the correlation filter framework in which the spatial regularization can help to suppress effect from the cluttered background while the temporal regularization enables the adaptive incorporation of historical appearance cues to deal with appearance changes. In addition, a dynamic modality weight learning algorithm is integrated into the correlation filter training, which ensures that more reliable modalities gain more importance in target tracking. Experimental results demonstrate the effectiveness of the proposed method.
Xiangyuan Lan, Zifei Yang, Wei Zhang 0021, Pong C. Yuen
ACM Trans. Multim. Comput. Commun. Appl.1
2020 Regularized Fine-Grained Meta Face Anti-Spoofing
abstract
Face presentation attacks have become an increasingly critical concern when face recognition is widely applied. Many face anti-spoofing methods have been proposed, but most of them ignore the generalization ability to unseen attacks. To overcome the limitation, this work casts face anti-spoofing as a domain generalization (DG) problem, and attempts to address this problem by developing a new meta-learning framework called Regularized Fine-grained Meta-learning. To let our face anti-spoofing model generalize well to unseen attacks, the proposed framework trains our model to perform well in the simulated domain shift scenarios, which is achieved by finding generalized learning directions in the meta-learning process. Specifically, the proposed framework incorporates the domain knowledge of face anti-spoofing as the regularization so that meta-learning is conducted in the feature space regularized by the supervision of domain knowledge. This enables our model more likely to find generalized learning directions with the regularized meta-learning for face anti-spoofing task. Besides, to further enhance the generalization ability of our model, the proposed framework adopts a fine-grained learning strategy that simultaneously conducts meta-learning in a variety of domain shift scenarios in each iteration. Extensive experiments on four public datasets validate the effectiveness of the proposed method.
Rui Shao 0001, Xiangyuan Lan, Pong C. Yuen
AAAI2
2020 Facial Expression Recognition Using Spatial-Temporal Semantic Graph Network
abstract
Motions of facial components convey significant information of facial expressions. Although remarkable advancement has been made, the dynamic of facial topology has not been fully exploited. In this paper, a novel facial expression recognition (FER) algorithm called Spatial Temporal Semantic Graph Network (STSGN) is proposed to automatically learn spatial and temporal patterns through end-to-end feature learning from facial topology structure. The proposed algorithm not only has greater discriminative power to capture the dynamic patterns of facial expression and stronger generalization capability to handle different variations but also higher interpretability. Experimental evaluation on two popular datasets, CK+ and Oulu-CASIA, shows that our algorithm achieves more competitive results than other state-of-the-art methods.
Jinzhao Zhou, Xingming Zhang 0001, Yang Liu 0182, Xiangyuan Lan
ICIP4
2020 Temporal Similarity Analysis of Remote Photoplethysmography for Fast 3D Mask Face Presentation Attack Detection
abstract
To tackle the 3D mask face presentation attack, remote Photoplethysmography (rPPG), a biomedical technique that can detect heartbeat signal remotely, is employed as an intrinsic liveness cue. Although existing rPPG-based methods exhibit encouraging results, they require long observation time (10-12 seconds) to identify the heartbeat information, which limits their employment in real applications such as smartphone unlock and e-payment. To shorten the observation time (within 1-second) while keeping the performance, we propose a fast rPPG-based 3D mask presentation attack detection (PAD) method by analyzing the similarity of local facial rPPG signals in the time domain. In particular, a set of temporal similarity features of facial and background local rPPG signals are designed and fused to adapt the real world variations based on rPPG shape and phase properties. For better evaluation under practical variations, we build the HKBU-MARsV2+ dataset that includes 16 masks from 2 types and 6 lighting conditions. Finally, extensive experiments are conducted on 11092 shortterm video slots from 4 datasets with a large number of real- world variations, in terms of mask type, lighting condition, camera, resolution of face region, and compression setting. Results show that the proposed TSrPPG outperforms the state-of-the-art competitors dramatically on discriminabil- ity and generalizability. To our best knowledge, this is the first work that addresses the length of observation time issue of rPPG-based 3D mask PAD.
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
WACV2
2020 Fusion of iris and sclera using phase intensive rubbersheet mutual exclusion for periocular recognition
Deepak Kumar Jain 0001, Xiangyuan Lan, Manikandan Ramachandran
Image Vis. Comput.2
2020 Deep Refinement: capsule network with attention mechanism-based system for text classification
Deepak Kumar Jain 0001, Rachna Jain, Yash Upadhyay, Abhishek Kathuria, Xiangyuan Lan
Neural Comput. Appl.5
2020 Modality-correlation-aware sparse representation for RGB-infrared object tracking
Xiangyuan Lan, Mang Ye, Shengping Zhang, Huiyu Zhou 0001, Pong C. Yuen
Pattern Recognit. Lett.1
2020 Bi-Directional Center-Constrained Top-Ranking for Visible Thermal Person Re-Identification
abstract
Visible thermal person re-identification (VT-REID) is a task of matching person images captured by thermal and visible cameras, which is an extremely important issue in night-time surveillance applications. Existing cross-modality recognition works mainly focus on learning sharable feature representations to handle the cross-modality discrepancies. However, apart from the cross-modality discrepancy caused by different camera spectrums, VT-REID also suffers from large cross-modality and intra-modality variations caused by different camera environments and human poses, and so on. In this paper, we propose a dual-path network with a novel bi-directional dual-constrained top-ranking (BDTR) loss to learn discriminative feature representations. It is featured in two aspects: 1) end-to-end learning without extra metric learning step and 2) the dual-constraint simultaneously handles the cross-modality and intra-modality variations to ensure the feature discriminability. Meanwhile, a bi-directional center-constrained top-ranking (eBDTR) is proposed to incorporate the previous two constraints into a single formula, which preserves the properties to handle both cross-modality and intra-modality variations. The extensive experiments on two cross-modality re-ID datasets demonstrate the superiority of the proposed method compared to the state-of-the-arts.
Mang Ye, Xiangyuan Lan, Zheng Wang 0007, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.2
2020 Improving Night-Time Pedestrian Retrieval With Distribution Alignment and Contextual Distance
abstract
Night-time pedestrian retrieval is a cross-modality retrieval task of retrieving person images between day-time visible images and night-time thermal images. It is a very challenging problem due to modality difference, camera variations, and person variations, but it plays an important role in night-time video surveillance. The existing cross-modality retrieval usually focuses on learning modality sharable feature representations to bridge the modality gap. In this article, we propose to utilize auxiliary information to improve the retrieval performance, which consistently improves the performance with different baseline loss functions. Our auxiliary information contains two major parts: cross-modality feature distribution and contextual information. The former aligns the cross-modality feature distributions between two modalities to improve the performance, and the latter optimizes the cross-modality distance measurement with the contextual information. We also demonstrate that abundant annotated visible pedestrian images, which are easily accessible, help to improve the cross-modality pedestrian retrieval as well. The proposed method is featured in two aspects: the auxiliary information does not need additional human intervention or annotation; it learns discriminative feature representations in an end-to-end deep learning manner. Extensive experiments on two cross-modality pedestrian retrieval datasets demonstrate the superiority of the proposed method, achieving much better performance than the state-of-the-arts.
Mang Ye, Xiangyuan Lan, Hongyuan Zhu 0002
IEEE Trans. Ind. Informatics3
2020 Cross-Modality Person Re-Identification via Modality-Aware Collaborative Ensemble Learning
abstract
Visible thermal person re-identification (VT-ReID) is a challenging cross-modality pedestrian retrieval problem due to the large intra-class variations and modality discrepancy across different cameras. Existing VT-ReID methods mainly focus on learning cross-modality sharable feature representations by handling the modality-discrepancy in feature level. However, the modality difference in classifier level has received much less attention, resulting in limited discriminability. In this paper, we propose a novel modality-aware collaborative ensemble (MACE) learning method with middle-level sharable two-stream network (MSTN) for VT-ReID, which handles the modality-discrepancy in both feature level and classifier level. In feature level, MSTN achieves much better performance than existing methods by capturing sharable discriminative middlelevel features in convolutional layers. In classifier level, we introduce both modality-specific and modality-sharable identity classifiers for two modalities to handle the modality discrepancy. To utilize the complementary information among different classifiers, we propose an ensemble learning scheme to incorporate the modality sharable classifier and the modality specific classifiers. In addition, we introduce a collaborative learning strategy, which regularizes modality-specific identity predictions and the ensemble outputs. Extensive experiments on two cross-modality datasets demonstrate that the proposed method outperforms current state-of-the-art by a large margin, achieving rank- 1/mAP accuracy 51.64%/50.11% on the SYSU-MM01 dataset, and 72.37%/69.09% on the RegDB dataset.
Mang Ye, Xiangyuan Lan, Qingming Leng, Jianbing Shen
IEEE Trans. Image Process.2
2020 Fine-Grained Spatial Alignment Model for Person Re-Identification With Focal Triplet Loss
abstract
Recent advances of person re-identification have well advocated the usage of human body cues to boost performance. However, most existing methods still retain on exploiting a relatively coarse-grained local information. Such information may include redundant backgrounds that are sensitive to the apparently similar persons when facing challenging scenarios like complex poses, inaccurate detection, occlusion and misalignment. In this paper we propose a novel Fine-Grained Spatial Alignment Model (FGSAM) to mine fine-grained local information to handle the aforementioned challenge effectively. In particular, we first design a pose resolve net with channel parse blocks (CPB) to extract pose information in pixel-level. This network allows the proposed model to be robust to complex pose variations while suppressing the redundant backgrounds caused by inaccurate detection and occlusion. Given the extracted pose information, a locally reinforced alignment mode is further proposed to address the misalignment problem between different local parts by considering different local parts along with attribute information in a fine-grained way. Finally, a focal triplet loss is designed to effectively train the entire model, which imposes a constraint on the intra-class and an adaptively weight adjustment mechanism to handle the hard sample problem. Extensive evaluations and analysis on Market1501, DukeMTMC-reid and PETA datasets demonstrate the effectiveness of FGSAM in coping with the problems of misalignment, occlusion and complex poses.
Qinqin Zhou 0001, Bineng Zhong 0001, Xiangyuan Lan, Gan Sun, Yulun Zhang 0001, Baochang Zhang 0001, Rongrong Ji
IEEE Trans. Image Process.3
2019 Multi-Adversarial Discriminative Deep Domain Generalization for Face Presentation Attack Detection
abstract
Face presentation attacks have become an increasingly critical issue in the face recognition community. Many face anti-spoofing methods have been proposed, but they cannot generalize well on "unseen" attacks. This work focuses on improving the generalization ability of face anti-spoofing methods from the perspective of the domain generalization. We propose to learn a generalized feature space via a novel multi-adversarial discriminative deep domain generalization framework. In this framework, a multi-adversarial deep domain generalization is performed under a dual-force triplet-mining constraint. This ensures that the learned feature space is discriminative and shared by multiple source domains, and thus is more generalized to new face presentation attacks. An auxiliary face depth supervision is incorporated to further enhance the generalization ability. Extensive experiments on four public datasets validate the effectiveness of the proposed method.
Rui Shao 0001, Xiangyuan Lan, Jiawei Li 0003, Pong C. Yuen
CVPR2
2019 LRDNN: Local-refining based Deep Neural Network for Person Re-Identification with Attribute Discerning
abstract
Recently, pose or attribute information has been widely used to solve person re-identification (re-ID) problem. However, the inaccurate output from pose or attribute modules will impair the final person re-ID performance. Since re-ID, pose estimation and attribute recognition are all based on the person appearance information, we propose a Local-refining based Deep Neural Network (LRDNN) to aggregate pose estimation and attribute recognition to improve the re-ID performance. To this end, we add a pose branch to extract the local spatial information and optimize the whole network on both person identity and attribute objectives. To diminish the negative affect from unstable pose estimation, a novel structure called channel parse block (CPB) is introduced to learn weights on different feature channels in pose branch. Then two branches are combined with compact bilinear pooling. Experimental results on Market1501 and DukeMTMC-reid datasets illustrate the effectiveness of the proposed method.
Qinqin Zhou 0001, Bineng Zhong 0001, Xiangyuan Lan, Gan Sun, Yulun Zhang 0001, Mengran Gou
IJCAI3
2019 Modality-aware Collaborative Learning for Visible Thermal Person Re-Identification
abstract
Visible thermal person re-identification (VT-ReID) is a cross-modality pedestrian retrieval problem, which automatically searches persons between day-time visible images and night-time thermal images. Despite the extensive progress in single-modality ReID, the cross-modality pedestrian retrieval problem has limited attention due to its challenges in modality discrepancy and large intra-class variations across cameras. Existing cross-modality ReID methods usually solve this problem by learning cross-modality feature representations with modality-sharable classifier. However, this learning strategy may lose discriminative information in different modalities. In this paper, we propose a novel modality-aware collaborative (MAC) learning method on top of a two-stream network for VT-ReID, which handles the modality-discrepancy in both feature level and classifier level. In feature level, it handles the modality discrepancy by a two-stream network with different parameters. In classifier level, it contains two separate modality-specific identity classifiers for two modalities to capture the modality-specific information, and they have the same network architecture but different parameters. In addition, we introduce a collaborative learning scheme, which regularizes the modality-sharable and modality-specific identity classifiers by utilizing the relationship between different classifiers. Extensive experiments on two cross-modality person re-identification datasets demonstrate the superiority of the proposed method, achieving much better performance than the state-of-the-art.
Mang Ye, Xiangyuan Lan, Qingming Leng
ACM Multimedia2
2019 Adversarial auto-encoder for unsupervised deep domain adaptation
abstract
Unsupervised visual domain adaptation aims to train a classifier that works well on a target domain given labelled source samples and unlabelled target samples. The key issue in unsupervised visual domain adaptation is how to do the feature alignment between source and target domains. Inspired by the adversarial learning in generative adversarial networks, this study proposes a novel adversarial auto‐encoder for unsupervised deep domain adaptation. This method incorporates the auto‐encoder with the adversarial learning so that the domain similarity and reconstruction information from the decoder can be exploited to facilitate the adversarial domain adaptation in the encoder. Extensive experiments on various visual recognition tasks show that the proposed method performs favourably against or better than competitive state‐of‐the‐art methods.
Rui Shao 0001, Xiangyuan Lan
IET Image Process.2
2019 Joint Discriminative Learning of Deep Dynamic Textures for 3D Mask Face Anti-Spoofing
abstract
Three-dimensional mask spoofing attacks have been one of the main challenges in face recognition. Compared with a 3D mask, a real face displays different facial motion patterns that are reflected by different facial dynamic textures. However, a large portion of these facial motion differences is subtle. We find that the subtle facial motion can be fully captured by multiple deep dynamic textures from a convolutional layer of a convolutional neural network, but not all deep dynamic textures from different spatial regions and different channels of a convolutional layer are useful for differentiation of subtle motions between real faces and 3D masks. In this paper, we propose a novel feature learning model to learn discriminative deep dynamic textures for 3D mask face anti-spoofing. A novel joint discriminative learning strategy is further incorporated in the learning model to jointly learn the spatial- and channel-discriminability of the deep dynamic textures. The proposed joint discriminative learning strategy can be used to adaptively weight the discriminability of the learned feature from different spatial regions or channels, which ensures that more discriminative deep dynamic textures play more important roles in face/mask classification. Experiments on several publicly available data sets validate that the proposed method achieves promising results in intra- and cross-data set scenarios.
Rui Shao 0001, Xiangyuan Lan, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.2
2018 Robust Collaborative Discriminative Learning for RGB-Infrared Tracking
abstract
Tracking target of interests is an important step for motion perception in intelligent video surveillance systems. While most recently developed tracking algorithms are grounded in RGB image sequences, it should be noted that information from RGB modality is not always reliable (e.g. in a dark environment with poor lighting condition), which urges the need to integrate information from infrared modality for effective tracking because of the insensitivity to illumination condition of infrared thermal camera. However, several issues encountered during the tracking process limit the fusing performance of these heterogeneous modalities: 1) the cross-modality discrepancy of visual and motion characteristics, 2) the uncertainty of degree of reliability in different modalities, and 3) large target appearance variations and background distractions within each modality. To address these issues, this paper proposes a novel and optimal discriminative learning framework for multi-modality tracking. In particular, the proposed discriminative learning framework is able to: 1) jointly eliminate outlier samples caused by large variations and learn discriminability-consistent features from heterogeneous modalities, and 2) collaboratively perform modality reliability measurement and target-background separation. Extensive experiments on RGB-infrared image sequences demonstrate the effectiveness of the proposed method.
Xiangyuan Lan, Mang Ye, Shengping Zhang, Pong C. Yuen
AAAI1
2018 Hierarchical Discriminative Learning for Visible Thermal Person Re-Identification
abstract
Person re-identification is widely studied in visible spectrum, where all the person images are captured by visible cameras. However, visible cameras may not capture valid appearance information under poor illumination conditions, e.g, at night. In this case, thermal camera is superior since it is less dependent on the lighting by using infrared light to capture the human body. To this end, this paper investigates a cross-modal re-identification problem, namely visible-thermal person re-identification (VT-REID). Existing cross-modal matching methods mainly focus on modeling the cross-modality discrepancy, while VT-REID also suffers from cross-view variations caused by different camera views. Therefore, we propose a hierarchical cross-modality matching model by jointly optimizing the modality-specific and modality-shared metrics. The modality-specific metrics transform two heterogenous modalities into a consistent space that modality-shared metric can be subsequently learnt. Meanwhile, the modality-specific metric compacts features of the same person within each modality to handle the large intra-modality intra-person variations (e.g. viewpoints, pose). Additionally, an improved two-stream CNN network is presented to learn the multi-modality sharable feature representations. Identity loss and contrastive loss are integrated to enhance the discriminability and modality-invariance with partially shared layer parameters. Extensive experiments illustrate the effectiveness and robustness of the proposed method.
Mang Ye, Xiangyuan Lan, Jiawei Li 0003, Pong C. Yuen
AAAI2
2018 Remote Photoplethysmography Correspondence Feature for 3D Mask Face Presentation Attack Detection
Si-Qi Liu 0003, Xiangyuan Lan, Pong C. Yuen
ECCV (16)2
2018 Robust Anchor Embedding for Unsupervised Video Person re-IDentification in the Wild
Mang Ye, Xiangyuan Lan, Pong C. Yuen
ECCV (7)2
2018 Visible Thermal Person Re-Identification via Dual-Constrained Top-Ranking
abstract
Cross-modality person re-identification between the thermal and visible domains is extremely important for night-time surveillance applications. Existing works in this filed mainly focus on learning sharable feature representations to handle the cross-modality discrepancies. However, besides the cross-modality discrepancy caused by different camera spectrums, visible thermal person re-identification also suffers from large cross-modality and intra-modality variations caused by different camera views and human poses. In this paper, we propose a dual-path network with a novel bi-directional dual-constrained top-ranking loss to learn discriminative feature representations. It is advantageous in two aspects: 1) end-to-end feature learning directly from the data without extra metric learning steps, 2) it simultaneously handles the cross-modality and intra-modality variations to ensure the discriminability of the learnt representations. Meanwhile, identity loss is further incorporated to model the identity-specific information to handle large intra-class variations. Extensive experiments on two datasets demonstrate the superior performance compared to the state-of-the-arts.
Mang Ye, Zheng Wang 0007, Xiangyuan Lan, Pong C. Yuen
IJCAI3
2018 Feature Constrained by Pixel: Hierarchical Adversarial Deep Domain Adaptation
abstract
In multimedia analysis, one objective of unsupervised visual domain adaptation is to train a classifier that works well on a target domain given labeled source samples and unlabeled target samples. Feature alignment of two domains is the key issue which should be addressed to achieve this objective. Inspired by the recent study of Generative Adversarial Networks (GAN) in domain adaptation, this paper proposes a new model based on Generative Adversarial Network, named Hierarchical Adversarial Deep Network (HADN), which jointly optimizes the feature-level and pixel-level adversarial adaptation within a hierarchical network structure. Specifically, the hierarchical network structure ensures that the knowledge from pixel-level adversarial adaptation can be back propagated to facilitate the feature-level adaptation, which achieves a better feature alignment under the constraint of pixel-level adversarial adaptation. Extensive experiments on various visual recognition tasks show that the proposed method performs favorably against or better than competitive state-of-the-art methods.
Rui Shao 0001, Xiangyuan Lan, Pong C. Yuen
ACM Multimedia2
2018 Learning Common and Feature-Specific Patterns: A Novel Multiple-Sparse-Representation-Based Tracker
abstract
The use of multiple features has been shown to be an effective strategy for visual tracking because of their complementary contributions to appearance modeling. The key problem is how to learn a fused representation from multiple features for appearance modeling. Different features extracted from the same object should share some commonalities in their representations while each feature should also have some feature-specific representation patterns which reflect its complementarity in appearance modeling. Different from existing multi-feature sparse trackers which only consider the commonalities among the sparsity patterns of multiple features, this paper proposes a novel multiple sparse representation framework for visual tracking which jointly exploits the shared and feature-specific properties of different features by decomposing multiple sparsity patterns. Moreover, we introduce a novel online multiple metric learning to efficiently and adaptively incorporate the appearance proximity constraint, which ensures that the learned commonalities of multiple features are more representative. Experimental results on tracking benchmark videos and other challenging videos demonstrate the effectiveness of the proposed tracker.
Xiangyuan Lan, Shengping Zhang, Pong C. Yuen, Rama Chellappa
IEEE Trans. Image Process.1
2018 Point-to-Set Distance Metric Learning on Deep Representations for Visual Tracking
abstract
For autonomous driving application, a car shall be able to track objects in the scene in order to estimate where and how they will move such that the tracker embedded in the car can efficiently alert the car for effective collision-avoidance. Traditional discriminative object tracking methods usually train a binary classifier via a support vector machine (SVM) scheme to distinguish the target from its background. Despite demonstrated success, the performance of the SVM-based trackers is limited because the classification is carried out only depending on support vectors (SVs) but the target's dynamic appearance may look similar to the training samples that have not been selected as SVs, especially when the training samples are not linearly classifiable. In such cases, the tracker may drift to the background and fail to track the target eventually. To address this problem, in this paper, we propose to integrate the point-to-set/image-to-imageSet distance metric learning (DML) into visual tracking tasks and take full advantage of all the training samples when determining the best target candidate. The point-to-set DML is conducted on convolutional neural network features of the training data extracted from the starting frames. When a new frame comes, target candidates are first projected to the common subspace using the learned mapping functions, and then the candidate having the minimal distance to the target template sets is selected as the tracking result. Extensive experimental results show that even without model update the proposed method is able to achieve favorable performance on challenging image sequences compared with several state-of-the-art trackers.
Shengping Zhang, Yuankai Qi, Feng Jiang 0001, Xiangyuan Lan, Pong C. Yuen, Huiyu Zhou 0001
IEEE Trans. Intell. Transp. Syst.4
2017 Robust MIL-Based Feature Template Learning for Object Tracking
abstract
Because of appearance variations, training samples of the tracked targets collected by the online tracker are required for updating the tracking model. However, this often leads to tracking drift problem because of potentially corrupted samples: 1) contaminated/outlier samples resulting from large variations (e.g. occlusion, illumination), and 2) misaligned samples caused by tracking inaccuracy. Therefore, in order to reduce the tracking drift while maintaining the adaptability of a visual tracker, how to alleviate these two issues via an effective model learning (updating) strategy is a key problem to be solved. To address these issues, this paper proposes a novel and optimal model learning (updating) scheme which aims to simultaneously eliminate the negative effects from these two issues mentioned above in a unified robust feature template learning framework. Particularly, the proposed feature template learning framework is capable of: 1) adaptively learning uncontaminated feature templates by separating out contaminated samples, and 2) resolving label ambiguities caused by misaligned samples via a probabilistic multiple instance learning (MIL) model. Experiments on challenging video sequences show that the proposed tracker performs favourably against several state-of-the-art trackers.
Xiangyuan Lan, Pong C. Yuen, Rama Chellappa
AAAI1
2017 Deep convolutional dynamic texture learning with adaptive channel-discriminability for 3D mask face anti-spoofing
abstract
3D mask spoofing attack has been one of the main challenges in face recognition. A real face displays a different motion behaviour compared to a 3D mask spoof attempt, which is reflected by different facial dynamic textures. However, the different dynamic information usually exists in the subtle texture level, which cannot be fully differentiated by traditional hand-crafted texture-based methods. In this paper, we propose a novel method for 3D mask face anti-spoofing, namely deep convolutional dynamic texture learning, which learns robust dynamic texture information from fine-grained deep convolutional features. Moreover, channel-discriminability constraint is adaptively incorporated to weight the discriminability of feature channels in the learning process. Experiments on both public datasets validate that the proposed method achieves promising results under intra and cross dataset scenario.
Rui Shao 0001, Xiangyuan Lan, Pong C. Yuen
IJCB2
2017 Robust Visual Tracking via Basis Matching
abstract
Most existing tracking approaches are based on either the tracking by detection framework or the tracking by matching framework. The former needs to learn a discriminative classifier using positive and negative samples, which will cause tracking drift due to unreliable samples. The latter usually performs tracking by matching local interest points between a target candidate and the tracked target, which is not robust to target appearance changes over time. In this paper, we propose a novel tracking by matching framework for robust tracking based on basis matching rather than point matching. In particular, we learn the target model from target images using a set of Gabor basis functions, which have large responses on the corresponding spatial positions after a max pooling. During tracking, a target candidate is evaluated by computing the responses of the Gabor basis functions on their corresponding spatial positions. The experimental results on a set of challenging sequences validate that the performance of the proposed tracking method outperforms those of several state-of-the-art methods.
Shengping Zhang, Xiangyuan Lan, Yuankai Qi, Pong C. Yuen
IEEE Trans. Circuits Syst. Video Technol.2
2017 A Biologically Inspired Appearance Model for Robust Visual Tracking
abstract
In this paper, we propose a biologically inspired appearance model for robust visual tracking. Motivated in part by the success of the hierarchical organization of the primary visual cortex (area V1), we establish an architecture consisting of five layers: whitening, rectification, normalization, coding, and pooling. The first three layers stem from the models developed for object recognition. In this paper, our attention focuses on the coding and pooling layers. In particular, we use a discriminative sparse coding method in the coding layer along with spatial pyramid representation in the pooling layer, which makes it easier to distinguish the target to be tracked from its background in the presence of appearance variations. An extensive experimental study shows that the proposed method has higher tracking accuracy than several state-of-the-art trackers.
Shengping Zhang, Xiangyuan Lan, Hongxun Yao, Huiyu Zhou 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.2
2016 Robust Joint Discriminative Feature Learning for Visual Tracking
Xiangyuan Lan, Shengping Zhang, Pong C. Yuen
IJCAI1
2015 Joint Sparse Representation and Robust Feature-Level Fusion for Multi-Cue Visual Tracking
abstract
Visual tracking using multiple features has been proved as a robust approach because features could complement each other. Since different types of variations such as illumination, occlusion, and pose may occur in a video sequence, especially long sequence videos, how to properly select and fuse appropriate features has become one of the key problems in this approach. To address this issue, this paper proposes a new joint sparse representation model for robust feature-level fusion. The proposed method dynamically removes unreliable features to be fused for tracking by using the advantages of sparse representation. In order to capture the non-linear similarity of features, we extend the proposed method into a general kernelized framework, which is able to perform feature fusion on various kernel spaces. As a result, robust tracking performance is obtained. Both the qualitative and quantitative experimental results on publicly available videos show that the proposed method outperforms both sparse representation-based and fusion based-trackers.
Xiangyuan Lan, Andy Jinhua Ma, Pong C. Yuen, Rama Chellappa
IEEE Trans. Image Process.1
2014 Multi-cue Visual Tracking Using Robust Feature-Level Fusion Based on Joint Sparse Representation
abstract
The use of multiple features for tracking has been proved as an effective approach because limitation of each feature could be compensated. Since different types of variations such as illumination, occlusion and pose may happen in a video sequence, especially long sequence videos, how to dynamically select the appropriate features is one of the key problems in this approach. To address this issue in multi-cue visual tracking, this paper proposes a new joint sparse representation model for robust feature-level fusion. The proposed method dynamically removes unreliable features to be fused for tracking by using the advantages of sparse representation. As a result, robust tracking performance is obtained. Experimental results on publicly available videos show that the proposed method outperforms both existing sparse representation based and fusion-based trackers.
Xiangyuan Lan, Andy Jinhua Ma, Pong C. Yuen
CVPR1