Jiaxin Chen 0002

dblp:65/1392-2 · DBLP profile ↗
← Back
47ranked-venue papers
10as first author
30since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 41 · 10 first-author · 25 since 2021Artificial intelligence and machine learning · 31 · 5 first-author · 22 since 2021
YearPublicationVenuePosition
2026 MP-ISMoE: Mixed-Precision Interactive Side Mixture-of-Experts for Efficient Transfer Learning
abstract
Parameter-efficient transfer learning (PETL) has emerged as a pivotal paradigm for adapting pre-trained foundation models to downstream tasks, significantly reducing trainable parameters yet suffering from substantial memory overhead caused by gradient backpropagation during fine-tuning. While memory-efficient transfer learning (METL) circumvents this challenge by bypassing backbone gradient computation via lightweight small side networks, its stringent memory constraint severely limits learning capacity of side networks, thereby significantly compromising performance. To address these limitations, we propose a novel Mixed-Precision Interactive Side Mixture-of-Experts framework (MP-ISMoE). Specifically, we first propose an Gaussian Noise Perturbed Iterative Quantization (GNP-IQ) scheme to quantize weights into lower-bits while effectively decreasing quantization errors. By leveraging memory conserved from GNP-IQ, we subsequently employ Interactive Side Mixture-of-Experts (ISMoE) to scale up side networks without sacrificing overall memory efficiency. Different from conventional mixture-of-experts, ISMoE learns to select optimal experts by interacting with salient features from frozen backbones, thus suppressing knowledge forgetting and boosting performance. Extensive experiments across diverse vision-language and language-only tasks demonstrate that MP-ISMoE remarkably promotes accuracy compared to state-of-the-art METL approaches, while maintaining comparable parameter and memory efficiency.
Zimeng Wu, Shengcai Liao, Shujiang Wu, Jiaxin Chen 0002
AAAI5
2026 Parameter-Efficient Tuning for Fine-Grained Recognition via Channel-Wise Importance Equalization and Diversity Navigation
abstract
Parameter-efficient tuning (PET) has achieved promising performance on various downstream vision tasks. Despite their effectiveness for general classification, existing PET approaches neglect the over-concentration of channel-wise saliency and the feature redundancy of pre-trained models during fine-tuning, thus leaving much room for improvement when applied to the downstream fine-grained recognition tasks. To address these issues, we propose a novel parameter-efficient tuning approach tailored for fine-grained recognition (FG-PET). Specifically, FG-PET first employs a Channel-wise Importance Equalization (CIE) module. It suppresses the concentrated salient channels while strengthening the remaining majority ones during fine-tuning, notably mitigating the over-concentrated saliency, thus evoking more channels within pre-trained models to deliver abundant local visual clues. Furthermore, FG-PET develops an Efficient Navigator for Diversity (EFIND) by introducing a center-based loss and orthogonal constraints on features generated from distinct attention heads. It alleviates the redundancy between different attention maps, thus enforcing the models to explore diverse subtle visual differences in various discriminative local regions, which are critical for fine-grained recognition. Extensive experimental results on five public fine-grained benchmarks based on distinct ViT models demonstrate that the proposed method remarkably boosts the performance of existing PET approaches, and generalizes well to general classification tasks. The source code is available at FG-PET.
Hanwen Zhong, Jiaxin Chen 0002, Di Huang 0001, Yunhong Wang 0001
IEEE Trans. Image Process.2
2025 TCAQ-DM: Timestep-Channel Adaptive Quantization for Diffusion Models
abstract
Diffusion models have achieved remarkable success in the image and video generation tasks. Nevertheless, they often require a large amount of memory and time overhead during inference, due to the complex network architecture and considerable number of timesteps for iterative diffusion. Recently, the post-training quantization (PTQ) technique has proved a promising way to reduce the inference cost by quantizing the float-point operations to low-bit ones. However, most of them fail to tackle with the large variations in the distribution of activations across distinct channels and timesteps, as well as the inconsistent of input between quantization and inference on diffusion models, thus leaving much room for improvement. To address the above issues, we propose a novel method dubbed Timestep-Channel Adaptive Quantization for Diffusion Models (TCAQ-DM). Specifically, we develop a timestep-channel joint reparameterization (TCR) module to balance the activation range along both the timesteps and channels, facilitating the successive reconstruction procedure. Subsequently, we employ a dynamically adaptive quantization (DAQ) module that mitigate the quantization error by selecting an optimal quantizer for each post-Softmax layers according to their specific types of distributions. Moreover, we present a progressively aligned reconstruction (PAR) strategy to mitigate the bias caused by the input mismatch. Extensive experiments on various benchmarks and distinct diffusion models demonstrate that the proposed method substantially outperforms the state-of-the-art approaches in most cases, especially yielding comparable FID metrics to the full precision model on CIFAR-10 in the W6A6 setting, while enabling generating available images in the W4A4 settings.
Haocheng Huang, Jiaxin Chen 0002, Jinyang Guo 0002, Ruiyi Zhan, Yunhong Wang 0001
AAAI2
2025 3D²-Actor: Learning Pose-Conditioned 3D-Aware Denoiser for Realistic Gaussian Avatar Modeling
abstract
Advancements in neural implicit representations and differentiable rendering have markedly improved the ability to learn animatable 3D avatars from sparse multi-view RGB videos. However, current methods that map observation space to canonical space often face challenges in capturing pose-dependent details and generalizing to novel poses. While diffusion models have demonstrated remarkable zero-shot capabilities in 2D image generation, their potential for creating animatable 3D avatars from 2D inputs remains underexplored. In this work, we introduce 3D²-Actor, a novel approach featuring a pose-conditioned 3D-aware human modeling pipeline that integrates iterative 2D denoising and 3D rectifying steps. The 2D denoiser, guided by pose cues, generates detailed multi-view images that provide the rich feature set necessary for high-fidelity 3D reconstruction and pose rendering. Complementing this, our Gaussian-based 3D rectifier renders images with enhanced 3D consistency through a two-stage projection strategy and a novel local coordinate representation. Additionally, we propose an innovative sampling strategy to ensure smooth temporal continuity across frames in video synthesis. Our method effectively addresses the limitations of traditional numerical solutions in handling ill-posed mappings, producing realistic and animatable 3D human avatars. Experimental results demonstrate that 3D²-Actor excels in high-fidelity avatar modeling and robustly generalizes to novel poses.
Zichen Tang, Hongyu Yang 0001, Hanchen Zhang, Jiaxin Chen 0002, Di Huang 0001
AAAI4
2025 Unified Knowledge Maintenance Pruning and Progressive Recovery with Weight Recalling for Large Vision-Language Models
abstract
Large Vision-Language Model (LVLM), leveraging Large Language Model (LLM) as the cognitive core, has recently become one of the most representative multimodal model paradigms. However, with the expansion of unimodal branches, \emph{i.e.} visual encoder and LLM, the storage and computational burdens intensify, posing challenges for deployment. Structured pruning has proved promising in compressing large models by trimming a large portion of insignificant network structures. Nevertheless, most of them are predominantly designed for LLMs, either relying on unitary importance metrics that fail to deal with modality-wise imbalances or adopting generic pruning and recovery paradigms that overlook the unique calibration status and capability requirements of large models, leading to substantial performance degradation. To address these issues, we propose a novel structured pruning approach for LVLMs, dubbed Unified Knowledge Maintenance Pruning and Progressive Recovery with Weight Recalling (UKMP). Specifically, we design a Unified Knowledge Maintenance Importance (UKMI) metric, which simultaneously considers balancing the block-wise and modality-wise importance by adaptive normalization, optimizing the importance estimation by refining gradient-based criteria, and maintaining the knowledge capacity of LVLMs by using the angle distribution information entropy. Moreover, we develop a LoRA-based Progressive Distillation (LPD) method that recalls the pruned weights and performs progressive distillation for comprehensive recovery. Extensive experimental results across various vision-language tasks demonstrate the effectiveness of our approach, comparing to the state-of-the-art structured pruning methods.
Zimeng Wu, Jiaxin Chen 0002, Yunhong Wang 0001
AAAI2
2025 FIMA-Q: Post-Training Quantization for Vision Transformers by Fisher Information Matrix Approximation
abstract
Post-training quantization (PTQ) has stood out as a cost-effective and promising model compression paradigm in recent years, as it avoids computationally intensive model retraining. Nevertheless, current PTQ methods for Vision Transformers (ViTs) still suffer from significant accuracy degradation, especially under low-bit quantization. To address these shortcomings, we analyze the prevailing Hessian-guided quantization loss, and uncover certain limitations of conventional Hessian approximations. By following the block-wise reconstruction framework, we propose a novel PTQ method for ViTs, dubbed FIMA-Q. Specifically, we firstly establish the connection between KL divergence and FIM, which enables fast computation of the quantization loss during reconstruction. We further propose an efficient FIM approximation method, namely DPLR-FIM, by employing the diagonal plus low-rank principle, and formulate the ultimate quantization loss. Our extensive experiments, conducted across various vision tasks with representative ViT-based architectures on public datasets, demonstrate that our method substantially promotes the accuracy compared to the state-of-the-art approaches, especially in the case of low-bit quantization. The source code is available at https://github.com/ShiheWang/FIMA-Q.
Zhuguanyu Wu, Shihe Wang, Jiaxin Chen 0002, Yunhong Wang 0001
CVPR4
2025 APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision Transformers
abstract
Vision Transformers (ViTs) have become one of the most commonly used backbones for vision tasks. Despite their remarkable performance, they often suffer significant accuracy drops when quantized for practical deployment, particularly by post-training quantization (PTQ) under ultra-low bits. Recently, reconstruction-based PTQ methods have shown promising performance in quantizing Convolutional Neural Networks (CNNs). However, they fail when applied to ViTs, primarily due to the inaccurate estimation of output importance and the substantial accuracy degradation in quantizing post-GELU activations. To address these issues, we propose APHQ-ViT, a novel PTQ approach based on importance estimation with Average Perturbation Hessian (APH). Specifically, we first thoroughly analyze the current approximation approaches with Hessian loss, and propose an improved average perturbation Hessian loss. To deal with the quantization of the post-GELU activations, we design an MLP Reconstruction (MR) method by replacing the GELU function in MLP with ReLU and reconstructing it by the APH loss on a small unlabeled calibration set. Extensive experiments demonstrate that APHQ-ViT using linear quantizers outperforms existing PTQ methods by substantial margins in 3-bit and 4-bit across different vision tasks. The source code is available at https://github.com/GoatWu/APHQ-ViT.
Zhuguanyu Wu, Jiaxin Chen 0002, Jinyang Guo 0002, Di Huang 0001, Yunhong Wang 0001
CVPR3
2025 CMAE-3D: Contrastive Masked AutoEncoders for Self-Supervised 3D Object Detection
Yanan Zhang 0005, Jiaxin Chen 0002, Di Huang 0001
Int. J. Comput. Vis.2
2025 Correction: CMAE-3D: Contrastive Masked AutoEncoders for Self-Supervised 3D Object Detection
Yanan Zhang 0005, Jiaxin Chen 0002, Di Huang 0001
Int. J. Comput. Vis.2
2025 Cross-Modal Contrastive Masked AutoEncoder for Compressed Video Pre-Training
abstract
In this paper, we propose a novel Transformer based approach, namely Cross-modal Contrastive Masked AutoEncoder (C2MAE), to Self-Supervised Learning (SSL) on compressed videos. A unified Transformer encoder is employed to discover relationships of visual tokens from RGBs, motion vectors and residuals. A hybrid SSL framework is proposed, which combines the complementary advantages of Masked Image Modeling (MIM) and Contrastive Learning (CL) pretext tasks, for powerful representation learning. The MIM branch extends VideoMAE by a new Fine-Grained Motion-aware Masking (FGMM) strategy and a modified Multi-modal Reconstruction (MR) task, where FGMM computes motion saliency maps as motion priors to guide the masks so that it well fits for the data properties in the compressed domain and the MR task highlights the reconstruction of raw videos by joint representations from corresponding compressed videos in addition to that in each single modality. The CL branch introduces the Contrastive Cross-modal Learning (CCL) module, and the features from a compressed video clip and the ones from its raw video counterpart are compared instead of widely used augmented data. Due to these designs, C2MAE significantly enhances interactions across modalities to compensate the sparsity of I-frames and the coarse and noisy nature of P-frames, thus delivering much stronger pre-trained models. Extensive experiments are conducted on the UCF-101, HMDB-51 and Kinetics-400 benchmarks with state-of-the-art results reported, demonstrating its effectiveness.
Jiaxin Chen 0002, Guohao Li 0010, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
IEEE Trans. Image Process.2
2025 Sharing Task-Relevant Information in Visual Prompt Tuning by Cross-Layer Dynamic Connection
abstract
Recent progress has shown great potential of visual prompt tuning (VPT) when adapting pre-trained vision transformers to various downstream tasks. However, most existing solutions independently optimize prompts at each layer, thereby neglecting the usage of task-relevant information encoded in prompt tokens across layers. Additionally, existing prompt structures are prone to interference from task-irrelevant noise in input images, which can adversely affect the sharing of task-relevant information. In this paper, we propose a novel VPT approach, SVPT. It innovatively incorporates a cross-layer dynamic connection (CDC) for input prompt tokens from adjacent layers, enabling effective sharing of task-relevant information. Furthermore, we design a dynamic aggregation (DA) module that facilitates selective sharing of information between layers. The combination of CDC and DA enhances the flexibility of the attention process within the VPT framework. Building upon these foundations, SVPT introduces an attentive enhancement (AE) mechanism that automatically identifies salient image tokens and refines them with prompt tokens in an additive manner. Extensive experiments on 24 image classification and semantic segmentation benchmarks clearly demonstrate the advantages of the proposed SVPT, compared to the state-of-the-art counterparts.
Jiaxin Chen 0002, Di Huang 0001
IEEE Trans. Image Process.2
2024 AdaLog: Post-training Quantization for Vision Transformers with Adaptive Logarithm Quantizer
Zhuguanyu Wu, Jiaxin Chen 0002, Hanwen Zhong, Di Huang 0001, Yunhong Wang 0001
ECCV (27)2
2024 Transforming Vision Transformer: Towards Efficient Multi-Task Asynchronous Learner
abstract
Multi-Task Learning (MTL) for Vision Transformer aims at enhancing the model capability by tackling multiple tasks simultaneously. Most recent works have predominantly focused on designing Mixture-of-Experts (MoE) structures and integrating Low-Rank Adaptation (LoRA) to efficiently perform multi-task learning. However, their rigid combination hampers both the optimization of MoE and the effectiveness of reparameterization of LoRA, leading to sub-optimal performance and low inference speed. In this work, we propose a novel approach dubbed Efficient Multi-Task Learning (EMTAL) by transforming a pre-trained Vision Transformer into an efficient multi-task learner during training, and reparameterizing the learned structure for efficient inference. Specifically, we firstly develop the MoEfied LoRA structure, which decomposes the pre-trained Transformer into a low-rank MoE structure and employ LoRA to fine-tune the parameters. Subsequently, we take into account the intrinsic asynchronous nature of multi-task learning and devise a learning Quality Retaining (QR) optimization mechanism, by leveraging the historical high-quality class logits to prevent a well-trained task from performance degradation. Finally, we design a router fading strategy to integrate the learned parameters into the original Transformer, archiving efficient inference. Extensive experiments on public benchmarks demonstrate the superiority of our method, compared to the state-of-the-art multi-task learning approaches.
Hanwen Zhong, Jiaxin Chen 0002, Di Huang 0001, Yunhong Wang 0001
NeurIPS2
2024 OTAMatch: Optimal Transport Assignment With PseudoNCE for Semi-Supervised Learning
abstract
In semi-supervised learning (SSL), many approaches follow the effective self-training paradigm with consistency regularization, utilizing threshold heuristics to alleviate label noise. However, such threshold heuristics lead to the underutilization of crucial discriminative information from the excluded data. In this paper, we present OTAMatch, a novel SSL framework that reformulates pseudo-labeling as an optimal transport (OT) assignment problem and simultaneously exploits data with high confidence to mitigate the confirmation bias. Firstly, OTAMatch models the pseudo-label allocation task as a convex minimization problem, facilitating end-to-end optimization with all pseudo-labels and employing the Sinkhorn-Knopp algorithm for efficient approximation. Meanwhile, we incorporate epsilon-greedy posterior regularization and curriculum bias correction strategies to constrain the distribution of OT assignments, improving the robustness with noisy pseudo-labels. Secondly, we propose PseudoNCE, which explicitly exploits pseudo-label consistency with threshold heuristics to maximize mutual information within self-training, significantly boosting the balance of convergence speed and performance. Consequently, our proposed approach achieves competitive performance on various SSL benchmarks. Specifically, OTAMatch substantially outperforms the previous state-of-the-art SSL algorithms in realistic and challenging scenarios, exemplified by a notable 9.45% error rate reduction over SoftMatch on ImageNet with 100K-label split, underlining its robustness and effectiveness.
Junjie Liu 0003, Debang Li, Qiuyu Huang, Jiaxin Chen 0002, Di Huang 0001
IEEE Trans. Image Process.5
2023 Adaptive Sparse Convolutional Networks with Global Context Enhancement for Faster Object Detection on Drone Images
abstract
Object detection on drone images with low-latency is an important but challenging task on the resource-constrained unmanned aerial vehicle (UAV) platform. This paper investigates optimizing the detection head based on the sparse convolution, which proves effective in balancing the accuracy and efficiency. Nevertheless, it suffers from inadequate integration of contextual information of tiny objects as well as clumsy control of the mask ratio in the presence of foreground with varying scales. To address the issues above, we propose a novel global context-enhanced adaptive sparse convolutional network (CEASC). It first develops a context-enhanced group normalization (CE-GN) layer, by replacing the statistics based on sparsely sampled features with the global contextual ones, and then designs an adaptive multi-layer masking strategy to generate optimal mask ratios at distinct scales for compact foreground coverage, promoting both the accuracy and efficiency. Extensive experimental results on two major benchmarks, i.e. VisDrone and UAVDT, demonstrate that CEASC remarkably reduces the GFLOPs and accelerates the inference procedure when plugging into the typical state-of-the-art detection frameworks (e.g. RetinaNet and GFL V1) with competitive performance. Code is available at https://github.com/Cuogeihong/CEASC.
Bowei Du, Yecheng Huang, Jiaxin Chen 0002, Di Huang 0001
CVPR3
2023 OcTr: Octree-Based Transformer for 3D Object Detection
abstract
A key challenge for LiDAR-based 3D object detection is to capture sufficient features from large scale 3D scenes especially for distant or/and occluded objects. Albeit recent efforts made by Transformers with the long sequence modeling capability, they fail to properly balance the accuracy and efficiency, suffering from inadequate receptive fields or coarse-grained holistic correlations. In this paper, we propose an Octree-based Transformer, named OcTr, to address this issue. It first constructs a dynamic octree on the hierarchical feature pyramid through conducting self-attention on the top level and then recursively propagates to the level below restricted by the octants, which captures rich global context in a coarse-to-fine manner while maintaining the computational complexity under control. Furthermore, for enhanced foreground perception, we propose a hybrid positional embedding, composed of the semantic-aware positional embedding and attention mask, to fully exploit semantic and geometry clues. Extensive experiments are conducted on the Waymo Open Dataset and KITTI Dataset, and OcTr reaches newly state-of-the-art results.
Yanan Zhang 0005, Jiaxin Chen 0002, Di Huang 0001
CVPR3
2023 DR-Tune: Improving Fine-tuning of Pretrained Visual Models by Distribution Regularization with Semantic Calibration
abstract
The visual models pretrained on large-scale benchmarks encode general knowledge and prove effective in building more powerful representations for downstream tasks. Most existing approaches follow the fine-tuning paradigm, either by initializing or regularizing the downstream model based on the pretrained one. The former fails to retain the knowledge in the successive fine-tuning phase, thereby prone to be over-fitting, and the latter imposes strong constraints to the weights or feature maps of the downstream model without considering semantic drift, often incurring insufficient optimization. To deal with these issues, we propose a novel fine-tuning framework, namely distribution regularization with semantic calibration (DR-Tune). It employs distribution regularization by enforcing the downstream task head to decrease its classification error on the pretrained feature distribution, which prevents it from over-fitting while enabling sufficient training of downstream encoders. Furthermore, to alleviate the interference by semantic drift, we develop the semantic calibration (SC) module to align the global shape and class centers of the pretrained and downstream feature distributions. Extensive experiments on widely used image classification datasets show that DR-Tune consistently improves the performance when combing with various backbones under different pretraining strategies. Code is available at: https://github.com/weeknan/DR-Tune.
Jiaxin Chen 0002, Di Huang 0001
ICCV2
2023 MIEP: Channel Pruning with Multi-granular Importance Estimation for Object Detection
abstract
This paper investigates compressing a pre-trained deep object detector to a lightweight one by channel pruning, which has proved effective and flexible in promoting efficiency. However, the majority of existing works trim channels based on a monotonous criterion for general purposes, i.e., the importance to the task-specific loss. They are prone to overly prune intermediate layers and simultaneously leave large intra-layer redundancy, severely deteriorating the detection accuracy. To address the issues above, we propose a novel channel pruning approach with multi-granular importance estimation (MIEP), consisting of the Feature-level Object-sensitive Importance (FOI) and the Intra-layer Redundancy-aware Importance (IRI). The former puts large weights on channels that are critical for object representation through the guidance of object features from the pre-trained model, and mitigates over-pruning when combined with the task-specific loss. The latter groups highly correlated channels based on clustering, which are subsequently pruned with priority to decrease redundancy. Extensive experiments on the COCO and VOC benchmarks demonstrate that MIEP remarkably outperforms the state-of-the-art channel pruning approaches, achieves a better balance between accuracy and efficiency compared to lightweight object detectors, and generalizes well to various detection frameworks (e.g., Faster-RCNN and FSAF) and tasks (e.g., classification).
Liangwei Jiang, Jiaxin Chen 0002, Di Huang 0001, Yunhong Wang 0001
ACM Multimedia2
2023 Compressed Video Prompt Tuning
abstract
Compressed videos offer a compelling alternative to raw videos, showing the possibility to significantly reduce the on-line computational and storage cost. However, current approaches to compressed video processing generally follow the resource-consuming pre-training and fine-tuning paradigm, which does not fully take advantage of such properties, making them not favorable enough for widespread applications. Inspired by recent successes of prompt tuning techniques in computer vision, this paper presents the first attempt to build a prompt based representation learning framework, which enables effective and efficient adaptation of pre-trained raw video models to compressed video understanding tasks. To this end, we propose a novel prompt tuning approach, namely Compressed Video Prompt Tuning (CVPT), emphatically dealing with the challenging issue caused by the inconsistency between pre-training and downstream data modalities. Specifically, CVPT replaces the learnable prompts with compressed modalities (\emph{e.g.} Motion Vectors and Residuals) by re-parameterizing them into conditional prompts followed by layer-wise refinement. The conditional prompts exhibit improved adaptability and generalizability to instances compared to conventional individual learnable ones, and the Residual prompts enhance the noisy motion cues in the Motion Vector prompts for further fusion with the visual cues from I-frames. Additionally, we design Selective Cross-modal Complementary Prompt (SCCP) blocks. After inserting them into the backbone, SCCP blocks leverage semantic relations across diverse levels and modalities to improve cross-modal interactions between prompts and input flows. Extensive evaluations on HMDB-51, UCF-101 and Something-Something v2 demonstrate that CVPT remarkably outperforms the state-of-the-art counterparts, delivering a much better balance between accuracy and efficiency.
Jiaxin Chen 0002, Xiuguo Bao, Di Huang 0001
NeurIPS2
2023 SAMP: Sub-task Aware Model Pruning with Layer-Wise Channel Balancing for Person Search
Zimeng Wu, Jiaxin Chen 0002, Yunhong Wang 0001
PRCV (10)2
2023 Learning Multi-Attention Context Graph for Group-Based Re-Identification
abstract
Learning to re-identify or retrieve a group of people across non-overlapped camera systems has important applications in video surveillance. However, most existing methods focus on (single) person re-identification (re-id), ignoring the fact that people often walk in groups in real scenarios. In this work, we take a step further and consider employing context information for identifying groups of people, i.e., group re-id. On the one hand, group re-id is more challenging than single person re-id, since it requires both a robust modeling of local individual person appearance (with different illumination conditions, pose/viewpoint variations, and occlusions), as well as full awareness of global group structures (with group layout and group member variations). On the other hand, we believe that person re-id can be greatly enhanced by incorporating additional visual context from neighboring group members, a task which we formulate as group-aware (single) person re-id. In this paper, we propose a novel unified framework based on graph neural networks to simultaneously address the above two group-based re-id tasks, i.e., group re-id and group-aware person re-id. Specifically, we construct a context graph with group members as its nodes to exploit dependencies among different people. A multi-level attention mechanism is developed to formulate both intra-group and inter-group context, with an additional self-attention module for robust graph-level representations by attentively aggregating node-level features. The proposed model can be directly generalized to tackle group-aware person re-id using node-level representations. Meanwhile, to facilitate the deployment of deep learning models on these tasks, we build a new group re-id dataset which contains more than 3.8K images with 1.5K annotated groups, an order of magnitude larger than existing group re-id datasets. Extensive experiments on the novel dataset as well as three existing datasets clearly demonstrate the effectiveness of the proposed framework for both group-based re-id tasks.
Yichao Yan, Jie Qin 0004, Bingbing Ni, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Wei-Shi Zheng 0001, Xiaokang Yang 0001, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 UFPMP-Det: Toward Accurate and Efficient Object Detection on Drone Imagery
abstract
This paper proposes a novel approach to object detection on drone imagery, namely Multi-Proxy Detection Network with Unified Foreground Packing (UFPMP-Det). To deal with the numerous instances of very small scales, different from the common solution that divides the high-resolution input image into quite a number of chips with low foreground ratios to perform detection on them each, the Unified Foreground Packing (UFP) module is designed, where the sub-regions given by a coarse detector are initially merged through clustering to suppress background and the resulting ones are subsequently packed into a mosaic for a single inference, thus significantly reducing overall time cost. Furthermore, to address the more serious confusion between inter-class similarities and intra-class variations of instances, which deteriorates detection performance but is rarely discussed, the Multi-Proxy Detection Network (MP-Det) is presented to model object distributions in a fine-grained manner by employing multiple proxy learning, and the proxies are enforced to be diverse by minimizing a Bag-of-Instance-Words (BoIW) guided optimal transport loss. By such means, UFPMP-Det largely promotes both the detection accuracy and efficiency. Extensive experiments are carried out on the widely used VisDrone and UAVDT datasets, and UFPMP-Det reports new state-of-the-art scores at a much higher speed, highlighting its advantages. The code is available at https://github.com/PuAnysh/UFPMP-Det.
Yecheng Huang, Jiaxin Chen 0002, Di Huang 0001
AAAI2
2022 Reading Chinese in Natural Scenes with a Bag-of-Radicals Prior
Qingjie Liu 0001, Jiaxin Chen 0002, Yunhong Wang 0001
BMVC3
2022 Entropy-based Active Learning for Object Detection with Progressive Diversity Constraint
abstract
Active learning is a promising alternative to alleviate the issue of high annotation cost in the computer vision tasks by consciously selecting more informative samples to label. Active learning for object detection is more challenging and existing efforts on it are relatively rare. In this paper, we propose a novel hybrid approach to address this problem, where the instance-level uncertainty and diversity are jointly considered in a bottom-up manner. To balance the computational complexity, the proposed approach is designed as a two-stage procedure. At the first stage, an Entropy-based Non-Maximum Suppression (ENMS) is presented to estimate the uncertainty of every image, which performs NMS according to the entropy in the feature space to remove predictions with redundant information gains. At the second stage, a diverse prototype (DivProto) strategy is explored to ensure the diversity across images by progressively converting it into the intra-class and inter-class diversities of the entropy-based class-specific prototypes. Extensive experiments are conducted on MS COCO and Pascal VOC, and the proposed approach achieves state of the art results and significantly outperforms the other counter-parts, highlighting its superiority.
Jiaxin Chen 0002, Di Huang 0001
CVPR2
2022 Target-Relevant Knowledge Preservation for Multi-Source Domain Adaptive Object Detection
abstract
Domain adaptive object detection (DAOD) is a promising way to alleviate performance drop of detectors in new scenes. Albeit great effort made in single source domain adaptation, a more generalized task with multiple source domains remains not being well explored, due to knowledge degradation during their combination. To address this issue, we propose a novel approach, namely target-relevant knowledge preservation (TRKP), to unsupervised multi-source DAOD. Specifically, TRKP adopts the teacher-student framework, where the multi-head teacher network is built to extract knowledge from labeled source domains and guide the student network to learn detectors in unlabeled target domain. The teacher network is further equipped with an adversarial multi-source disentanglement (AMSD) module to preserve source domain-specific knowledge and simultaneously perform cross-domain alignment. Besides, a holistic target-relevant mining (HTRM) scheme is developed to re-weight the source images according to the source-target relevance. By this means, the teacher network is enforced to capture target-relevant knowledge, thus benefiting decreasing domain shift when mentoring object detection in the target domain. Extensive experiments are conducted on various widely used benchmarks with new state-of-the-art scores reported, highlighting the effectiveness.
Jiaxin Chen 0002, Mengzhe He, Yiru Wang 0003, Bo Li 0114, Bingqi Ma, Weihao Gan, Wei Wu 0021, Yali Wang 0001, Di Huang 0001
CVPR2
2022 CAT-Det: Contrastively Augmented Transformer for Multimodal 3D Object Detection
abstract
In autonomous driving, LiDAR point-clouds and RGB images are two major data modalities with complementary cues for 3D object detection. However, it is quite difficult to sufficiently use them, due to large inter-modal discrepancies. To address this issue, we propose a novel framework, namely Contrastively Augmented Transformer for multi-modal 3D object Detection (CAT-Det). Specifically, CAT-Det adopts a two-stream structure consisting of a Pointformer (PT) branch, an Imageformer (IT) branch along with a Cross-Modal Transformer (CMT) module. PT, IT and CMT jointly encode intra-modal and inter-modal long-range contexts for representing an object, thus fully exploring multi-modal information for detection. Furthermore, we propose an effective One-way Multimodal Data Augmentation (OMDA) approach via hierarchical contrastive learning at both the point and object levels, significantly improving the accuracy only by augmenting point-clouds, which is free from complex generation of paired samples of the two modalities. Extensive experiments on the KITTI benchmark show that CAT-Det achieves a new state-of-the-art, highlighting its effectiveness.
Yanan Zhang 0005, Jiaxin Chen 0002, Di Huang 0001
CVPR2
2022 Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
abstract
Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues (e.g., motion vectors and residuals). However, this task severely suffers from the coarse and noisy dynamics and the insufficient fusion of the heterogeneous RGB and motion modalities. To address the two issues above, this paper proposes a novel framework, namely Attentive Cross-modal Interaction Network with Motion Enhancement (MEACI-Net). It follows the two-stream architecture, i.e. one for the RGB modality and the other for the motion modality. Particularly, the motion stream employs a multi-scale block embedded with a denoising module to enhance representation learning. The interaction between the two streams is then strengthened by introducing the Selective Motion Complement (SMC) and Cross-Modality Augment (CMA) modules, where SMC complements the RGB modality with spatio-temporally attentive local motion features and CMA further combines the two modalities with selective feature augmentation. Extensive experiments on the UCF-101, HMDB-51 and Kinetics-400 benchmarks demonstrate the effectiveness and efficiency of MEACI-Net.
Jiaxin Chen 0002, Dongming Zhang 0004, Xiuguo Bao, Di Huang 0001
IJCAI2
2022 SHREC'22 track: Sketch-based 3D shape retrieval in the wild
Jie Qin 0004, Shuaihang Yuan, Jiaxin Chen 0002, Boulbaba Ben Amor, Yi Fang 0006, Nhat Hoang-Xuan, Chi-Bien Chu, Khoi-Nguyen Nguyen-Ngoc, Thien-Tri Cao, Nhat-Khang Ngô, Tuan-Luc Huynh, Hai-Dang Nguyen, Minh-Triet Tran, Haoyang Luo, Jianning Wang, Zheng Zhang 0006, Zihao Xin, Yang Wang 0023, Haiqin Chen, Qunying Zhou
Comput. Graph.3
2022 Video Person Re-Identification Using Attribute-Enhanced Features
abstract
In this work we propose to boost video-based person re-identification (Re-ID) by using attribute-enhanced feature presentation. To this end, we not only try to use the ID-relevant attributes more effectively, but also for the first time in literature harness the ID-irrelevant attributes to help model training. The former mainly include gender, age, clothing characteristics, etc., which contain rich and supplementary information about the pedestrian; the latter include viewpoint, action, etc., which are seldom used for identification previously. In particular, we use the attributes to enhance the significant areas of the image with a novel Attribute Salient Region Enhance (ASRE) module that can attend more accurately to the body of the pedestrian, so as to better separate the target from the background. Furthermore, we find that many ID-irrelevant but subject-relevant factors, like the view angle and movement of the target pedestrian, have great impact on the two-dimensional appearance of a pedestrian. We then propose to exploit both the ID-relevant and the ID-irrelevant attributes via a novel triplet loss called the Viewpoint and Action-Invariant (VAI) triplet loss. Based on the above, we design an Attribute Salience Assisted Network (ASA-Net) to perform attribute recognition along with identity recognition, and use the attributes for feature enhancement and hard sample mining. Extensive experiments on MARS and DukeMTMC-VideoReID datasets show that our method outperforms the state-of-the-arts. Also, the visualizations of learning results further prove the effectiveness of the proposed method.
Tianrui Chai, Annan Li, Jiaxin Chen 0002, Xinyu Mei, Yunhong Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2021 PR-GCN: A Deep Graph Convolutional Network with Point Refinement for 6D Pose Estimation
abstract
RGB-D based 6D pose estimation has recently achieved remarkable progress, but still suffers from two major limitations: (1) ineffective representation of depth data and (2) insufficient integration of different modalities. This paper proposes a novel deep learning approach, namely Graph Convolutional Network with Point Refinement (PR-GCN), to simultaneously address the issues above in a unified way. It first introduces the Point Refinement Network (PRN) to polish 3D point clouds, recovering missing parts with noise removed. Subsequently, the Multi-Modal Fusion Graph Convolutional Network (MMF-GCN) is presented to strengthen RGB-D combination, which captures geometry-aware inter-modality correlation through local information propagation in the graph convolutional network. Extensive experiments are conducted on three widely used benchmarks, and state-of-the-art performance is reached. Besides, it is also shown that the proposed PRN and MMF-GCN modules are well generalized to other frameworks.
Guangyuan Zhou, Huiqun Wang, Jiaxin Chen 0002, Di Huang 0001
ICCV3
2020 Auto-Encoding Twin-Bottleneck Hashing
abstract
Conventional unsupervised hashing methods usually take advantage of similarity graphs, which are either pre-computed in the high-dimensional space or obtained from random anchor points. On the one hand, existing methods uncouple the procedures of hash function learning and graph construction. On the other hand, graphs empirically built upon original data could introduce biased prior knowledge of data relevance, leading to sub-optimal retrieval performance. In this paper, we tackle the above problems by proposing an efficient and adaptive code-driven graph, which is updated by decoding in the context of an auto-encoder. Specifically, we introduce into our framework twin bottlenecks (i.e., latent variables) that exchange crucial information collaboratively. One bottleneck (i.e., binary codes) conveys the high-level intrinsic data structure captured by the code-driven graph to the other (i.e., continuous variables for low-level detail information), which in turn propagates the updated network feedback for the encoder to learn more discriminative binary codes. The auto-encoding learning objective literally rewards the code-driven graph to learn an optimal encoder. Moreover, the proposed model can be simply optimized by gradient descent without violating the binary constraints. Experiments on benchmarked datasets clearly show the superiority of our framework over the state-of-the-art hashing methods. Our source code can be found at https://github.com/ymcidence/TBH.
Yuming Shen, Jie Qin 0004, Jiaxin Chen 0002, Mengyang Yu, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Ling Shao 0001
CVPR3
2020 Learning Multi-Granular Hypergraphs for Video-Based Person Re-Identification
abstract
Video-based person re-identification (re-ID) is an important research topic in computer vision. The key to tackling the challenging task is to exploit both spatial and temporal clues in video sequences. In this work, we propose a novel graph-based framework, namely Multi-Granular Hypergraph (MGH), to pursue better representational capabilities by modeling spatiotemporal dependencies in terms of multiple granularities. Specifically, hypergraphs with different spatial granularities are constructed using various levels of part-based features across the video sequence. In each hypergraph, different temporal granularities are captured by hyperedges that connect a set of graph nodes (i.e., part-based features) across different temporal ranges. Two critical issues (misalignment and occlusion) are explicitly addressed by the proposed hypergraph propagation and feature aggregation schemes. Finally, we further enhance the overall video representation by learning more diversified graph-level representations of multiple granularities based on mutual information minimization. Extensive experiments on three widely-adopted benchmarks clearly demonstrate the effectiveness of the proposed framework. Notably, 90.0% top-1 accuracy on MARS is achieved using MGH, outperforming the state-of-the-arts.
Yichao Yan, Jie Qin 0004, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Ying Tai, Ling Shao 0001
CVPR3
2020 Learning Attentive and Hierarchical Representations for 3D Shape Recognition
Jiaxin Chen 0002, Jie Qin 0004, Yuming Shen, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
ECCV (15)1
2020 Deep Local Binary Coding for Person Re-Identification by Delving into the Details
abstract
Person re-identification (ReID) has recently received extensive research interests due to its diverse applications in multimedia analysis and computer vision. However, the majority of existing works focus on improving matching accuracy, while ignoring matching efficiency. In this work, we present a novel binary representation learning framework for efficient person ReID, namely Deep Local Binary Coding (DLBC). Different from existing deep binary ReID approaches, DLBC attempts to learn discriminative binary codes by explicitly interacting with local visual details. Specifically, DLBC first extracts a set of local features from spatially salient regions of pedestrian images. Subsequently, DLBC formulates a new binary-local semantic mutual information (BSMI) maximization term, based on which a self-lifting (SL) block is built to further exploit the semantic importance of local features. The BSMI term together with the SL block simultaneously enhances the dependency of binary codes on selected local features as well as their robustness to cross-view visual inconsistency. In addition, an efficient optimizing method is developed to train the proposed deep models with orthogonal and binary constraints. Extensive experiments reveal that DLBC significantly minimizes the accuracy gap between binary ReID methods and the state-of-the-art real-valued ones, whilst remarkably reducing query time and memory cost.
Jiaxin Chen 0002, Jie Qin 0004, Yichao Yan, Lei Huang 0015, Li Liu 0004, Fan Zhu 0001, Ling Shao 0001
ACM Multimedia1
2020 Zero-VAE-GAN: Generating Unseen Features for Generalized and Transductive Zero-Shot Learning
abstract
Zero-shot learning (ZSL) is a challenging task due to the lack of unseen class data during training. Existing works attempt to establish a mapping between the visual and class spaces through a common intermediate semantic space. The main limitation of existing methods is the strong bias towards seen class, known as the domain shift problem, which leads to unsatisfactory performance in both conventional and generalized ZSL tasks. To tackle this challenge, we propose to convert ZSL to the conventional supervised learning by generating features for unseen classes. To this end, a joint generative model that couples variational autoencoder (VAE) and generative adversarial network (GAN), called Zero-VAE-GAN, is proposed to generate high-quality unseen features. To enhance the class-level discriminability, an adversarial categorization network is incorporated into the joint framework. Besides, we propose two self-training strategies to augment unlabeled unseen features for the transductive extension of our model, addressing the domain shift problem to a large extent. Experimental results on five standard benchmarks and a large-scale dataset demonstrate the superiority of our generative model over the state-of-the-art methods for conventional, especially generalized ZSL tasks. Moreover, the further improvement of the transductive setting demonstrates the effectiveness of the proposed self-training strategies.
Xingsong Hou, Jie Qin 0004, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Zhao Zhang 0001, Ling Shao 0001
IEEE Trans. Image Process.4
2019 Deep Sketch-Shape Hashing With Segmented 3D Stochastic Viewing
abstract
Sketch-based 3D shape retrieval has been extensively studied in recent works, most of which focus on improving the retrieval accuracy, whilst neglecting the efficiency. In this paper, we propose a novel framework for efficient sketch-based 3D shape retrieval, i.e., Deep Sketch-Shape Hashing (DSSH), which tackles the challenging problem from two perspectives. Firstly, we propose an intuitive 3D shape representation method to deal with unaligned shapes with arbitrary poses. Specifically, the proposed Segmented Stochastic-viewing Shape Network models discriminative 3D representations by a set of 2D images rendered from multiple views, which are stochastically selected from non-overlapping spatial segments of a 3D sphere. Secondly, Batch-Hard Binary Coding (BHBC) is developed to learn semantics-preserving compact binary codes by mining the hardest samples. The overall framework is jointly learned by developing an alternating iteration algorithm. Extensive experimental results on three benchmarks show that DSSH improves both the retrieval efficiency and accuracy remarkably, compared to the state-of-the-art methods.
Jiaxin Chen 0002, Jie Qin 0004, Li Liu 0004, Fan Zhu 0001, Fumin Shen, Jin Xie 0001, Ling Shao 0001
CVPR1
2018 Deep Cross-Modality Adaptation via Semantics Preserving Adversarial Learning for Sketch-Based 3D Shape Retrieval
Jiaxin Chen 0002, Yi Fang 0006
ECCV (13)1
2018 GII Representation-Based Cross-View Gait Recognition by Discriminative Projection With List-Wise Constraints
abstract
Remote person identification by gait is one of the most important topics in the field of computer vision and pattern recognition. However, gait recognition suffers severely from the appearance variance caused by the view change. It is very common that gait recognition has a high performance when the view is fixed but the performance will have a sharp decrease when the view variance becomes significant. Existing approaches have tried all kinds of strategies like tensor analysis or view transform models to slow down the trend of performance decrease but still have potential for further improvement. In this paper, a discriminative projection with list-wise constraints (DPLC) is proposed to deal with view variance in cross-view gait recognition, which has been further refined by introducing a rectification term to automatically capture the principal discriminative information. The DPLC with rectification (DPLCR) embeds list-wise relative similarity measurement among intraclass and inner-class individuals, which can learn a more discriminative and robust projection. Based on the original DPLCR, we have introduced the kernel trick to exploit nonlinear cross-view correlations and extended DPLCR to deal with the problem of multiview gait recognition. Moreover, a simple yet efficient gait representation, namely gait individuality image (GII), based on gait energy image is proposed, which could better capture the discriminative information for cross view gait recognition. Experiments have been conducted in the CASIA-B database and the experimental results demonstrate the outstanding performance of both the DPLCR framework and the new GII representation. It is shown that the DPLCR-based cross-view gait recognition has outperformed the-state-of-the-art approaches in almost all cases under large view variance. The combination of the GII representation and the DPLCR has further enhanced the performance to be a new benchmark for cross-view gait recognition.
Zhaoxiang Zhang 0001, Jiaxin Chen 0002, Qiang Wu 0001, Ling Shao 0001
IEEE Trans. Cybern.2
2017 Fast Person Re-identification via Cross-Camera Semantic Binary Transformation
abstract
Numerous methods have been proposed for person re-identification, most of which however neglect the matching efficiency. Recently, several hashing based approaches have been developed to make re-identification more scalable for large-scale gallery sets. Despite their efficiency, these works ignore cross-camera variations, which severely deteriorate the final matching accuracy. To address the above issues, we propose a novel hashing based method for fast person re-identification, namely Cross-camera Semantic Binary Transformation (CSBT). CSBT aims to transform original high-dimensional feature vectors into compact identity-preserving binary codes. To this end, CSBT first employs a subspace projection to mitigate cross-camera variations, by maximizing intra-person similarities and inter-person discrepancies. Subsequently, a binary coding scheme is proposed via seamlessly incorporating both the semantic pairwise relationships and local affinity information. Finally, a joint learning framework is proposed for simultaneous subspace projection learning and binary coding based on discrete alternating optimization. Experimental results on four benchmarks clearly demonstrate the superiority of CSBT over the state-of-the-art methods.
Jiaxin Chen 0002, Yunhong Wang 0001, Jie Qin 0004, Li Liu 0004, Ling Shao 0001
CVPR1
2017 Zero-Shot Action Recognition with Error-Correcting Output Codes
abstract
Recently, zero-shot action recognition (ZSAR) has emerged with the explosive growth of action categories. In this paper, we explore ZSAR from a novel perspective by adopting the Error-Correcting Output Codes (dubbed ZSECOC). Our ZSECOC equips the conventional ECOC with the additional capability of ZSAR, by addressing the domain shift problem. In particular, we learn discriminative ZSECOC for seen categories from both category-level semantics and intrinsic data structures. This procedure deals with domain shift implicitly by transferring the well-established correlations among seen categories to unseen ones. Moreover, a simple semantic transfer strategy is developed for explicitly transforming the learned embeddings of seen categories to better fit the underlying structure of unseen categories. As a consequence, our ZSECOC inherits the promising characteristics from ECOC as well as overcomes domain shift, making it more discriminative for ZSAR. We systematically evaluate ZSECOC on three realistic action benchmarks, i.e. Olympic Sports, HMDB51 and UCF101. The experimental results clearly show the superiority of ZSECOC over the state-of-the-art methods.
Jie Qin 0004, Li Liu 0004, Ling Shao 0001, Fumin Shen, Bingbing Ni, Jiaxin Chen 0002, Yunhong Wang 0001
CVPR6
2016 Person re-identification by distance metric learning to discrete hashing
abstract
Most of the existing works on person re-identification have focused on improving matching rate at top ranks. Few efforts are devoted to address the problem of efficient storage and fast search for person re-identification. In this paper, we investigate the prevailing hashing method, originally designed for large scale image retrieval, for fast person re-identification with efficient storage. We propose a novel hashing approach, namely Distance Metric Learning to Discrete Hashing (DMLDH), which jointly learns a discriminative projection via metric learning to alleviate cross-view variations, and a hashing function for discriminative binary coding by minimizing inner-class Hamming distances and maximizing inter-class Hamming distances. To deal with the formulated non-convex optimization problem, we develop an alternative iteration algorithm by solving several subproblems with analytical solutions. Experimental results on benchmarks demonstrate that the proposed method outperforms the state-of-the-art hashing approaches.
Jiaxin Chen 0002, Yunhong Wang 0001
ICIP1
2016 A fast adaptive spatio-temporal 3D feature for video-based person re-identification
abstract
Video-based person re-identification has become a hot topic in the field of research on computer vision and intelligent surveillance, which is more robust to the variations in a person's appearance than single-shot based methods and involves space-time information. However, the most existing spatiotemporal features have been proposed for action recognition that they mainly focus on the exact spatial changes over time. Unlike action recognition, pedestrians captured in person re-identification problem show similar and cyclic walking activities. The essential spatio-temporal information for person re-identification is the statistical information over time. In this paper, we propose a novel spatio-temporal feature, namely Fast Adaptive Spatio-Temporal 3D feature (FAST3D), for video-based person re-identification. The feature is able to extract the statistical motion information based on densely computed multi-direction gradients and an adaptive fusion process. We evaluate our method on two challenging datasets and the experimental results show the effectiveness and efficiency of the proposed feature.
Zheng Liu 0014, Jiaxin Chen 0002, Yunhong Wang 0001
ICIP2
2016 Person Re-identification by Exploiting Spatio-Temporal Cues and Multi-view Metric Learning
abstract
In this letter, we introduce a new spatio-temporal feature, namely optical flow energy image (OFEI), for video-based person re-identification. OFEI aims to exploit spatio-temporally stable regions across frames, which can capture discriminative cues such as human body parts and carry-on stuffs. Furthermore, we propose a novel matching method, denoted by multi-view relevance metric learning with list-wise constraints (mvRMLLC), to integrate the spatio-temporal (i.e., OFEI) and appearance features. Unlike previous works, mvRMLLC assumes that multiple features are generated from different views with distinct data distributions, while their similarities should be globally consistent. Multiple similarity metrics are then learned and fused by maximizing their global consistency and simultaneously allowing local discrepancies. Extensive experiments on two benchmarks demonstrate that OFEI outperforms the state-of-the-art spatio-temporal features, and mvRMLLC could further enhance the overall performance significantly.
Jiaxin Chen 0002, Yunhong Wang 0001, Yuan Yan Tang
IEEE Signal Process. Lett.1
2016 Beyond Semantic Attributes: Discrete Latent Attributes Learning for Zero-Shot Recognition
abstract
In this letter, we propose a novel approach for learning semantics-driven attributes, which are discriminative for zero-shot visual recognition. Latent attributes are derived in a principled manner, aiming at maintaining class-level semantic relatedness and attribute-wise balancedness. Unlike existing methods that binarize learned real-valued attributes via a quantization stage, we directly learn the optimal binary attributes by effectively addressing a discrete optimization problem. Particularly, we propose a class-wise discrete descent algorithm, based on which latent attributes of each class are learned iteratively. Moreover, we propose to simultaneously predict multiple attributes from low-level features via multioutput neural networks (MONN), which can model intrinsic correlation among attributes and make prediction more tractable. Extensive experiments on two standard datasets clearly demonstrate the superiority of our method over the state-of-the-arts.
Jie Qin 0004, Yunhong Wang 0001, Li Liu 0004, Jiaxin Chen 0002, Ling Shao 0001
IEEE Signal Process. Lett.4
2016 Corrections to "Relevance Metric Learning for Person Re-Identification by Exploiting Listwise Similarities"
Jiaxin Chen 0002, Zhaoxiang Zhang 0001, Yunhong Wang 0001
IEEE Trans. Image Process.1
2015 Relevance Metric Learning for Person Re-Identification by Exploiting Listwise Similarities
abstract
Person re-identification aims to match people across non-overlapping camera views, which is an important but challenging task in video surveillance. In order to obtain a robust metric for matching, metric learning has been introduced recently. Most existing works focus on seeking a Mahalanobis distance by employing sparse pairwise constraints, which utilize image pairs with the same person identity as positive samples, and select a small portion of those with different identities as negative samples. However, this training strategy has abandoned a large amount of discriminative information, and ignored the relative similarities. In this paper, we propose a novel relevance metric learning method with listwise constraints (RMLLCs) by adopting listwise similarities, which consist of the similarity list of each image with respect to all remaining images. By virtue of listwise similarities, RMLLC could capture all pairwise similarities, and consequently learn a more discriminative metric by enforcing the metric to conserve predefined similarity lists in a low-dimensional projection subspace. Despite the performance enhancement, RMLLC using predefined similarity lists fails to capture the relative relevance information, which is often unavailable in practice. To address this problem, we further introduce a rectification term to automatically exploit the relative similarities, and develop an efficient alternating iterative algorithm to jointly learn the optimal metric and the rectification term. Extensive experiments on four publicly available benchmarking data sets are carried out and demonstrate that the proposed method is significantly superior to the state-of-the-art approaches. The results also show that the introduction of the rectification term could further boost the performance of RMLLC.
Jiaxin Chen 0002, Zhaoxiang Zhang 0001, Yunhong Wang 0001
IEEE Trans. Image Process.1
2014 Relevance Metric Learning for Person Re-identification by Exploiting Global Similarities
abstract
Person re-identification aims to match people across non-overlapping camera views, which is an important and challenging task. In order to obtain a robust metric for measuring (dis)similarities of (un)matched image pairs, metric learning has been introduced recently. Most existing works focus on seeking a Mahalanobis distance by employing sparse pair wise (dis)similarity constraints. However, the pair wise constraints have ignored a large portion of useful similarity information, and could not provide global similarity information. This paper proposes a novel metric learning method that could effectively exploit the global similarities. Specifically, we predefine lists of similarity scores, and measure (dis)similarities by the relevance of feature vectors. Subsequently, we learn a relevance metric by using the predefined list wise constraints, where the learnt metric is enforced to conserve predefined list wise similarities. Our main contributions lie on three folds: (1) we propose a metric learning method, which could effectively encode the global similarity information by using list wise constraints, (2) we formulate the relevance metric learning into a convex optimization problem, which could be solved efficiently, (3) we further kernelize the proposed method to support nonlinear mappings. The proposed method is experimentally validated on benchmark datasets, and outperforms state-of-the-art metric learning methods.
Jiaxin Chen 0002, Zhaoxiang Zhang 0001, Yunhong Wang 0001
ICPR1