Andy Jinhua Ma

dblp:119/1514 · also Andy J. Ma · DBLP profile ↗
← Back
79ranked-venue papers
10as first author
48since 2021 · last 2026
0000-0002-0165-8416ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 51 · 7 first-author · 29 since 2021Artificial intelligence and machine learning · 37 · 8 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 8 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 CSQDA: A Parameter-Efficient and Memory-Efficient Tuning Method for Medical Image Classification
Yiqian Li, Andy Jinhua Ma
MMM (2)2
2026 SSMC: spatial-spectral mask consistency learning for semi-supervised medical image segmentation
Yidan Qin, Chenyu Cai, Andy Jinhua Ma
Mach. Vis. Appl.5
2026 DiffusionEngine: Diffusion model is scalable data engine for object detection
Manlin Zhang, Jie Wu 0032, Yuxi Ren, Ming Li 0010, Andy Jinhua Ma
Pattern Recognit.6
2026 Federated Single-Positive Multi-Label Learning
abstract
Single-positive multi-label learning (SPMLL) aims to train a multi-label classifier from data with single-positive label, to predict all applicable labels during testing. However, existing SPMLL methods are tailored for centralized datasets, which fail to be directly deployed to distributed setting like federated learning. In this paper, we start the first attempt to study federated single-positive multi-label learning (FedSPMLL), aiming to collaboratively train a SPMLL model from distributed data. To achieve this, we need to address challenges caused by label incompleteness: limited generalization ability of local model and overweighting contribution of client with local dataset suffering from severe label incompleteness. To this end, we propose a novelFedLOGmethod, guidingFedSPMLL with predicateLOGic-modeled label correlation. Enabling the informative knowledge extraction from limited data, we propose to model label correlation within local dataset using predicate logic. To alleviate false negative label issue, we propose to transfer confident label correlation knowledge to local model by self-distillation. To downweight the contribution of unreliable client owning dataset with severe label incompleteness, we propose a new measurement of label incompleteness to adjust client contribution for a fair aggregation. We establish a comprehensive FedSPMLL benchmark. And extensive experiments demonstrate the superiority of our FedLOG method.
Mang Ye, Andy Jinhua Ma, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.3
2025 Video Individual Counting for Moving Drones
abstract
Video Individual Counting (VIC) has received increasing attention for its importance in intelligent video surveillance. Existing works are limited in two aspects, i.e., dataset and method. Previous datasets are captured with fixed or rarely moving cameras with relatively sparse individuals, restricting evaluation for a highly varying view and time in crowded scenes. Existing methods rely on localization followed by association or classification, which struggle under dense and dynamic conditions due to inaccurate localization of small targets. To address these issues, we introduce the MovingDroneCrowd Dataset, featuring videos captured by fast-moving drones in crowded scenes under diverse illuminations, shooting heights and angles. We further propose a Shared Density map-guided Network (SDNet) using a Depth-wise Cross-Frame Attention (DCFA) module to directly estimate shared density maps between consecutive frames, from which the inflow and outflow density maps are derived by subtracting the shared density maps from the global density maps. The inflow density maps across frames are summed up to obtain the number of unique pedestrians in a video. Experiments on our datasets and publicly available ones show the superiority of our method over the state of the arts in highly dynamic and complex crowded scenes. Our dataset and codes have been released publicly.
Yaowu Fan, Jia Wan 0001, Tao Han 0002, Antoni B. Chan, Andy Jinhua Ma
ICCV5
2025 Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video Synthesis
abstract
Distribution Matching Distillation (DMD) is a promising score distillation technique that compresses pre-trained teacher diffusion models into efficient one-step or multi-step student generators. Nevertheless, its reliance on the reverse Kullback-Leibler (KL) divergence minimization potentially induces mode collapse (or mode-seeking) in certain applications. To circumvent this inherent drawback, we propose Adversarial Distribution Matching (ADM), a novel framework that leverages diffusion-based discriminators to align the latent predictions between real and fake score estimators for score distillation in an adversarial manner. In the context of extremely challenging one-step distillation, we further improve the pre-trained generator by adversarial distillation with hybrid discriminators in both latent and pixel spaces. Different from the mean squared error used in DMD2 pre-training, our method incorporates the distributional loss on ODE pairs collected from the teacher model, and thus providing a better initialization for score distillation fine-tuning in the next stage. By combining the adversarial distillation pre-training with ADM fine-tuning into a unified pipeline termed DMDX, our proposed method achieves superior one-step performance on SDXL compared to DMD2 while consuming less GPU time. Additional experiments that apply multi-step ADM distillation on SD3-Medium, SD3.5-Large, and CogVideoX set a new benchmark towards efficient image and video synthesis.
Yanzuo Lu, Yuxi Ren, Xin Xia 0005, Shanchuan Lin, Xuefeng Xiao 0001, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai
ICCV7
2025 FIE: Filtering, Inference and Enhancement for Multi-modal Object Re-identification
Qingcheng Yang, Yanzuo Lu, Andy Jinhua Ma
PRCV (16)3
2025 Background-aware Prior Mask Refinement and Feature Augmentation for Few-Shot Segmentation
abstract
Few-shot semantic segmentation enables the rapid deployment of segmentation models to new categories with limited annotated data. In recent works, it is a promising approach to generating prior masks based on pixel-wise similarity as the guidance for segmentation. Since the prior masks are generated by a frozen pre-trained backbone, they may be biased to wrongly activate the non-target classes in the background regions. To address this issue, we propose a Background-aware Feature Augmentation (BFA) method for few-shot semantic segmentation. Our method first estimates foreground and background prior masks, and then fuses them by a learnable convolution. With the help of the background information, more accurate prior masks can be generated by deactivating the background regions. By utilizing the refined prior masks, foreground and background features are augmented to encode foreground and background information for segmentation. The proposed BFA is a plug-and-play method which can be easily integrated into existing works based on the prototype matching framework. Experiments on PASCAL-5i and COCO-20i datasets demonstrate the superiority of our method compared to the state of the arts.
Hongrong He, Andy Jinhua Ma
SMC3
2025 Contrastive Cross-modal Prototype Prediction and Fusion for Video Anomaly Detection*
abstract
Video Anomaly Detection (VAD) identifies unexpected events by learning normal behavior from surveillance footage, assuming only normal training data is available. Previous methods, which focus on frame reconstruction or prediction tasks, are constrained by insufficient semantic sensitivity due to a reliance on pixel-level errors and inadequate adaptability to diverse normal patterns. Although contrastive learning methods attempt to address these issues by learning normal subcategories through manually constructed positive-negative pairs, they still encounter semantic ambiguity arising from inappropriate contrastive strategies. To overcome these limitations, we propose a Contrastive Cross-modal Prototype Prediction and Fusion (C2P2F) framework. Specifically, our method comprises three stages: (1) First, the Contrastive Prototype Prediction (CPP) module separately learns normal patterns of appearance and motion on RGB frames and optical flow inputs. Without constructing positive-negative pairs, we perform a cross-view prototype prediction task to discern inherent normal patterns within the data. (2) Then, the Cross-Modal Prototype Fusion (CMPF) conducts alternative training with RGB and optical flow inputs to establish comprehensive normal representations from two complementary modalities, enforcing consistency between cross-modal prototypes and learning semantically rich normal patterns. (3) Finally, the Prototype Number Adjustment (PNA) module is employed to mitigate initialization bias. Integrating these components, our approach adaptively models diverse normalcy and enhances anomaly discrimination via joint optimization of prototype stability and cross-modal consistency, with experiments on three benchmarks showing its superiority.
Junqiao Wang, Jiawen Peng, Andy Jinhua Ma
SMC4
2025 Adversarial Style Mixup and Improved Temporal Alignment for Cross-Domain Few-Shot Action Recognition
Kaiyan Cao, Jiawen Peng, Xinyuan Hou, Andy Jinhua Ma
Comput. Vis. Image Underst.5
2025 Language-guided Alignment and Distillation for Source-free Domain Adaptation
Jiawen Peng, Andy Jinhua Ma
Neurocomputing4
2025 Invariant prompting with classifier rectification for continual learning
Chunsing Lo, Andy Jinhua Ma
Image Vis. Comput.3
2025 Vision-Language Adaptive Clustering and Meta-Adaptation for Unsupervised Few-Shot Action Recognition
abstract
Unsupervised few-shot action recognition is a practical but challenging task, which adapts knowledge learned from unlabeled videos to novel action classes with only limited labeled data. Without annotated data of base action classes for meta-learning, it cannot achieve satisfactory performance due to the low-quality pseudo-classes and episodes. Though vision-language pre-training models such as CLIP can be employed to improve the quality of pseudo-classes and episodes, the performance improvements may still be limited by using only the visual encoder in the absence of textual modality information. In this paper, we propose fully exploiting the multimodal knowledge of a pre-trained vision-language model CLIP in a novel framework for unsupervised video meta-learning. Textual modality is automatically generated for each unlabeled video by a video-to-text transformer. Multimodal adaptive clustering for episodic sampling (MACES) based on a video-text ensemble distance metric is proposed to accurately estimate pseudo-classes, which constructs high-quality few-shot tasks (episodes) for episodic training. Vision-language meta-adaptation (VLMA) is designed for adapting the pre-trained model to novel tasks by category-aware vision-language contrastive learning and confidence-based reliable bidirectional knowledge distillation. The final prediction is obtained by multimodal adaptive inference. Extensive experiments on five benchmarks demonstrate the superiority of our method for unsupervised few-shot action recognition.
Jiawen Peng, Yanzuo Lu, Jian-Huang Lai, Andy Jinhua Ma
IEEE Trans. Circuits Syst. Video Technol.5
2025 Learning Crowd Scale and Distribution for Weakly Supervised Crowd Counting and Localization
abstract
The count supervision used in weakly-supervised crowd counting is derived from the number of point annotations, which means that the labeling cost is not effectively reduced. Moreover, due to the lack of spatial information about the pedestrians during training, previous works struggle to accurately learn the positions of individuals. To address these challenges, we propose a crowd counting and localization method based on scene-specific synthetic data for surveillance scenarios, which can accurately predict the number and location of person without any manually labeled point-wise or count-wise annotations. Our method dynamically adjust scene-specific synthetic data to minimize domain differences from surveillance scenes by learning the crowd scale and distribution. Specifically, based on realistic synthetic data, the models learn precise location and scale information, which can then regenerate new synthetic data with a more reasonable pedestrian distribution and scale and generate high-quality pseudo point-wise annotations. Subsequently, the counter is trained using our proposed robust soft-weighted loss function, under the joint supervision of auto-generated point-wise annotations on synthetic data and pseudo point-wise annotations on real data in an end-to-end manner. Our proposed loss function, based on the designed weighted optimal transport, effectively mitigates noise in pseudo point-wise labels and is not only insensitive to hyperparemeters but also exhibits superior generalization ability on real data. We conduct comprehensive experiments across multiple scene-specific datasets, demonstrating our method’s superiority in counting and localization performance over count-supervised, fully-supervised, and state-of-the-art domain adaption algorithms. Code is available athttps://github.com/fyw1999/LCSD.
Yaowu Fan, Jia Wan 0001, Andy Jinhua Ma
IEEE Trans. Circuits Syst. Video Technol.3
2024 MLNet: Mutual Learning Network with Neighborhood Invariance for Universal Domain Adaptation
abstract
Universal domain adaptation (UniDA) is a practical but challenging problem, in which information about the relation between the source and the target domains is not given for knowledge transfer. Existing UniDA methods may suffer from the problems of overlooking intra-domain variations in the target domain and difficulty in separating between the similar known and unknown class. To address these issues, we propose a novel Mutual Learning Network (MLNet) with neighborhood invariance for UniDA. In our method, confidence-guided invariant feature learning with self-adaptive neighbor selection is designed to reduce the intra-domain variations for more generalizable feature representation. By using the cross-domain mixup scheme for better unknown-class identification, the proposed method compensates for the misidentified known-class errors by mutual learning between the closed-set and open-set classifiers. Extensive experiments on three publicly available benchmarks demonstrate that our method achieves the best results compared to the state-of-the-arts in most cases and significantly outperforms the baseline across all the four settings in UniDA. Code is available at https://github.com/YanzuoLu/MLNet.
Yanzuo Lu, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai
AAAI3
2024 Memory-Guided Contrastive and Triplet Separation for Weakly-Supervised Disease Detection
abstract
Weakly-supervised disease detection has the great potential to alleviate the time-consuming and labor-intensive burden of manual annotations in instance level. While existing methods extract normality prototypes encoding normal patterns to improve the detection performance, they may fail to detect subtle anomalies without considering the diverse abnormal patterns. In this paper, we propose to recognize both normal and abnormal patterns for feature representation learning in disease detection based on a dual memory network. To learn discriminative memory banks and classifiers, dual memory loss is incorporated with a feature magnitude separation loss for model training. With the proposed contrastive feature separation loss and triplet feature separation loss, feature discriminability is further improved to obtain a more separable decision boundary. Additionally, we collect a large-scale CT dataset to evaluate lung tumor detection. Extensive experiments on the publicly available PANDA-MIL dataset and the collected LUNG-MIL dataset demonstrate the superiority of our proposed method compared to the state-of-the-art approaches for weakly-supervised disease detection.
Jinwen She, Andy Jinhua Ma
BIBM3
2024 Coarse-to-Fine Latent Diffusion for Pose-Guided Person Image Synthesis
abstract
Diffusion model is a promising approach to image generation and has been employed for Pose-Guided Person Image Synthesis (PGPIS) with competitive performance. While existing methods simply align the person appearance to the target pose, they are prone to overfitting due to the lack of a high-level semantic understanding on the source person image. In this paper, we propose a novel Coarse-to-Fine Latent Diffusion (CFLD) method for PGPIS. In the absence of image-caption pairs and textual prompts, we de-velop a novel training paradigm purely based on images to control the generation process of a pre-trained text-to-image diffusion model. A perception-refined decoder is designed to progressively refine a set of learnable queries and extract semantic understanding of person images as a coarse-grained prompt. This allows for the decoupling of fine-grained appearance and pose information controls at different stages, and thus circumventing the potential over-fitting problem. To generate more realistic texture details, a hybrid- granularity attention module is proposed to encode multi-scale fine-grained appearance features as bias terms to augment the coarse-grained prompt. Both quantitative and qualitative experimental results on the DeepFashion benchmark demonstrate the superiority of our method over the state of the arts for PGPIS. Code is available at https://github.com/YanzuoLu/CFLD.
Yanzuo Lu, Manlin Zhang, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai
CVPR3
2024 Confidence-guided Source Label Refinement and Class-aware Pseudo-Label Thresholding for Domain-Adaptive Semantic Segmentation
abstract
Domain-Adaptive Semantic Segmentation (DASS) aims to transfer a segmentation model trained on a labeled source domain to an unlabeled target domain. Most existing methods overlook the errors in the source labels and directly utilize the erroneous source labels for training. This may result in wrong source domain knowledge being wrongly transferred to the target domain, which leads to similar classes cannot be separated well in the target domain. To this end, we propose a novel Confidence-guided Online REfinement (CORE) method, which introduces a refinement network trained on the target data to rectify erroneous source labels online. This can effectively avoid transferring erroneous knowledge from the source domain to the target domain. On the other hand, using the same threshold for all classes in the target domain may result in a severe class imbalance in self-training. Therefore, we propose Class-aware Adaptive Thresholding (CAT) to calculate different thresholds for different target classes, which adaptively filter target pixels with high confidence for better adaptation. Our proposed CORE-CAT method can be easily integrated with existing methods and achieves superior performance across various domain-adaptive semantic segmentation benchmarks.
Hongwei Chu, Jiawen Peng, Andy Jinhua Ma
IJCNN4
2024 Pair Shuffle Consistency for Semi-supervised Medical Image Segmentation
Chenyu Cai, Andy Jinhua Ma
MICCAI (8)4
2023 Gradient Adjusted and Weight Rectified Mean Teacher for Source-Free Object Detection
Jiawen Peng, Yanxu Hu, Andy Jinhua Ma
ICANN (7)5
2023 Transformer Based Prototype Learning for Weakly-Supervised Histopathology Tissue Semantic Segmentation
Jinwen She, Yanxu Hu, Andy Jinhua Ma
ICANN (4)3
2023 Dual Episodic Sampling and Momentum Consistency Regularization for Unsupervised Few-shot Learning
abstract
Unsupervised Few-shot Learning (UFSL) is a practical approach to adapting knowledge learned from unlabeled data of base classes to novel classes with limited labeled data. Nevertheless, most existing UFSL methods may not learn generalizable features in latter training epochs due to the simplicity of meta-learning tasks constructed by data augmentation. To address this issue, we propose two novel components, namely Dual Episodic Sampling (DES) and Momentum Consistency Regularization (MCR) for UFSL. In the DES, two types of sampling strategies are used to construct harder training tasks with multiple augmentations to generate each pseudo-class of increased diversity. The MCR constrains the consistency of the backbone encoder with its momentum counterpart to learn better generalized features for novel classes. Experimental results on four datasets verify the superiority of our method for unsupervised few-shot image classification.
Yanxu Hu, Andy Jinhua Ma
ICME4
2023 E2: Entropy Discrimination and Energy Optimization for Source-free Universal Domain Adaptation
abstract
Universal domain adaptation (UniDA) transfers knowledge under both distribution and category shifts. Most UniDA methods accessible to source-domain data during model adaptation may result in privacy policy violation and source-data transfer inefficiency. To address this issue, we propose a novel source-free UniDA method coupling confidence-guided entropy discrimination and likelihood-induced energy optimization. The entropy-based separation of target-known and unknown classes is too conservative for known-class prediction. Thus, we derive the confidence-guided entropy by scaling the normalized prediction score with the known-class confidence, that more known-class samples are correctly predicted. Due to difficult estimation of the marginal distribution without source-domain data, we constrain the target-domain marginal distribution by maximizing (minimizing) the known (unknown)-class likelihood, which equals free energy optimization. Theoretically, the overall optimization amounts to decreasing and increasing internal energy of known and unknown classes in physics, respectively. Extensive experiments demonstrate the superiority of the proposed method.
Andy Jinhua Ma, Pong C. Yuen
ICME2
2023 Discriminative Gradient Adjustment with Coupled Knowledge Distillation for Class Incremental Learning
abstract
Class Incremental Learning (CIL) is a promising approach to addressing the catastrophic forgetting problem when learning for new categories. Though recent works based on dynamic architectures achieve convincing performance, data imbalance caused by limited size of memory and compression of the increasingly growing network are challenges to be solved. In this paper, we propose the novel Discriminative Gradient Adjustment (DGA) and Coupled Knowledge Distillation strategy (CKD) for these two challengs. The DGA mitigates the data imbalance problem by designing the loss function with a static global balance factor and a ground-truth-based dynamic factor. The CKD fully utilizes intermediate layers of the dual-branch models by feature-level distillation with moving-average weight updating for network compression. Extensive experiments on CIFAR100 and ImageNet100 datasets demonstrate the superiority of our method for CIL.
Yanxu Hu, Jiawen Peng, Andy Jinhua Ma
ICME4
2023 Collaborative Learning of Diverse Experts for Source-free Universal Domain Adaptation
abstract
Source-free universal domain adaptation (SFUniDA) is a challenging yet practical problem that adapts the source model to the target domain in the presence of distribution and category shifts without accessing source domain data. Most existing methods are developed based on a single-expert target model for both known- and unknown-class data training, such that the known- and unknown-class data in the target domain may not be separated well from each other. To address this issue, we propose a novel Cobllaborative Learning of Diverse Experts (CoDE) method for SFUniDA. In our method, unknown-class compatible source model training is designed to reserve space for the potential target unknown-class data. Two diverse experts are learned to better recognize the target known- and unknown-class data respectively by the specialized entropy discrimination. We improve the transferability of both experts by collaboratively correcting the possible misclassification errors with consistency and diversity learning. The final prediction with high confidence is obtained by gating the diverse experts based on soft neighbor density. Extensive experiments on four publicly available benchmarks demonstrate the superiority of our method compared to the state of the art.
Yanzuo Lu, Yanxu Hu, Andy Jinhua Ma
ACM Multimedia4
2023 Patch Shuffle and Pixel Contrast: Dual Consistency Learning for Semi-supervised Lung Tumor Segmentation
Chenyu Cai, Manlin Zhang, Yanxu Hu, Andy Jinhua Ma
PRCV (5)6
2023 DilateFormer: Multi-Scale Dilated Transformer for Visual Recognition
abstract
As ade factosolution, the vanilla Vision Transformers (ViTs) are encouraged to model long-range dependencies between arbitrary image patches while the global attended receptive field leads to quadratic computational cost. Another branch of Vision Transformers exploits local attention inspired by CNNs, which only models the interactions between patches in small neighborhoods. Although such a solution reduces the computational cost, it naturally suffers from small attended receptive fields, which may limit the performance. In this work, we explore effective Vision Transformers to pursue a preferable trade-off between the computational complexity and size of the attended receptive field. By analyzing the patch interaction of global attention in ViTs, we observe two key properties in the shallow layers, namely locality and sparsity, indicating the redundancy of global dependency modeling in shallow layers of ViTs. Accordingly, we propose Multi-Scale Dilated Attention (MSDA) to modellocalandsparsepatch interaction within the sliding window. With a pyramid architecture, we construct a Multi-Scale Dilated Transformer (DilateFormer) by stacking MSDA blocks at low-level stages and global multi-head self-attention blocks at high-level stages. Our experiment results show that our DilateFormer achieves state-of-the-art performance on various vision tasks. On ImageNet-1 K classification task, DilateFormer achieves comparable performance with 70% fewer FLOPs compared with existing state-of-the-art models. Our DilateFormer-Base achieves 85.6% top-1 accuracy on ImageNet-1 K classification task, 53.5% box mAP/46.1% mask mAP on COCO object detection/instance segmentation task and 51.1% MS mIoU on ADE20 K semantic segmentation task.
Jiayu Jiao, Yu-Ming Tang, Kun-Yu Lin, Yipeng Gao, Andy Jinhua Ma, Yaowei Wang 0001, Wei-Shi Zheng 0001
IEEE Trans. Multim.5
2022 Suppressing Static Visual Cues via Normalizing Flows for Self-Supervised Video Representation Learning
abstract
Despite the great progress in video understanding made by deep convolutional neural networks, feature representation learned by existing methods may be biased to static visual cues. To address this issue, we propose a novel method to suppress static visual cues (SSVC) based on probabilistic analysis for self-supervised video representation learning. In our method, video frames are first encoded to obtain latent variables under standard normal distribution via normalizing flows. By modelling static factors in a video as a random variable, the conditional distribution of each latent variable becomes shifted and scaled normal. Then, the less-varying latent variables along time are selected as static cues and suppressed to generate motion-preserved videos. Finally, positive pairs are constructed by motion-preserved videos for contrastive learning to alleviate the problem of representation bias to static cues. The less-biased video representation can be better generalized to various downstream tasks. Extensive experiments on publicly available benchmarks demonstrate that the proposed method outperforms the state of the art when only single RGB modality is used for pre-training.
Manlin Zhang, Andy Jinhua Ma
AAAI3
2022 Anatomical prior-inspired label refinement for weakly supervised liver tumor segmentation with volume-level labels
Fei Lyu 0004, Andy Jinhua Ma, Pong C. Yuen
BMVC2
2022 Adversarial Feature Augmentation for Cross-domain Few-Shot Classification
Yanxu Hu, Andy Jinhua Ma
ECCV (20)2
2022 Region-Interactive Proposal Network and Class-Interactive Feature Learning for Few-Shot Object Detection
abstract
Few-shot object detection is a promising approach to solving the problem of detecting novel objects with only limited annotated data for training. Most existing methods are developed based on the progress in few-shot classification, which pay little attention to improving the localization module and modelling class interrelation. To address these issues, this paper proposes two novel modules, namely Region-interactive Proposal Network (Ri-PN) and Class-interactive Feature Learning (Ci-FL), for better localization and classification performance, respectively. In the Ri-PN, regions of novel classes are interacted with base classes via graph convolution instead of background due to the stronger relevance between base and novel classes together with the guidance of supervised regions loss. On the other hand, the Ci-FL refines class-specific features in prototypical learning by attentive graph convolutional network. Experimental results on PASCAL VOC and MS COCO datasets verify the superiority of our method for few-shot object detection.
Yanxu Hu, Faming Wu, Andy Jinhua Ma
ICME4
2022 Learning to Mitigate Extreme Distribution Bias for Few-Shot Object Detection
abstract
Few-shot object detection is an important but challenging task where only a few instances of novel categories are available. The widely used approach is to pretrain a detector on base classes with abundant samples and then fine-tune it for novel classes. Due to the extreme data imbalance between base and novel classes, the detection performance of novel classes degrades with the distribution bias. To overcome this limitation, we propose a distribution calibration strategy and a class discrimination regularization method for better few-shot detection. Based on theoretical analysis on decision margins of base and novel classes, the decision area of novel classes is enlarged to balance the prediction probability. On the other hand, to increase the separability of inter-class distributions, the similarity between class-specific representations is minimized. Extensive experiments on PASCAL VOC and MS COCO datasets verify the effectiveness and generalization ability of our method to improve few-shot object detection.
Faming Wu, Yanxu Hu, Andy Jinhua Ma
ICME4
2022 Source-free Temporal Attentive Domain Adaptation for Video Action Recognition
abstract
With the rapidly increasing video data, many video analysis techniques have been developed and achieved success in recent years. To mitigate the distribution bias of video data across domains, unsupervised video domain adaptation (UVDA) has been proposed and become an active research topic. Nevertheless, existing UVDA methods need to access source domain data during training, which may result in problems of privacy policy violation and transfer inefficiency. To address this issue, we propose a novel source-free temporal attentive domain adaptation (SFTADA) method for video action recognition under the more challenging UVDA setting, such that source domain data is not required for learning the target domain. In our method, an innovative Temporal Attentive aGgregation (TAG) module is designed to combine frame-level features with varying importance weights for video-level representation generation. Without source domain data and label information in the target domain and during testing, an MLP-based attention network is trained to approximate the attentive aggregation function based on class centroids. By minimizing frame-level and video-level loss functions, both the temporal and spatial domain shifts in cross-domain video data can be reduced. Extensive experiments on four benchmark datasets demonstrate the effectiveness of our proposed method in solving the challenging source-free UVDA task.
Peipeng Chen, Andy Jinhua Ma
ICMR2
2022 Improving Pre-trained Masked Autoencoder via Locality Enhancement for Person Re-identification
Yanzuo Lu, Manlin Zhang, Andy Jinhua Ma, Xiaohua Xie, Jian-Huang Lai
PRCV (2)4
2022 Multi-level Attentive Adversarial Learning with Temporal Dilation for Unsupervised Video Domain Adaptation
abstract
Most existing works on unsupervised video domain adaptation attempt to mitigate the distribution gap across domains in frame and video levels. Such two-level distribution alignment approach may suffer from the problems of insufficient alignment for complex video data and misalignment along the temporal dimension. To address these issues, we develop a novel framework of Multi-level Attentive Adversarial Learning with Temporal Dilation (MA2L- TD). Given frame-level features as input, multi-level temporal features are generated and multiple domain discriminators are individually trained by adversarial learning for them. For better distribution alignment, level-wise attention weights are calculated by the degree of domain confusion in each level. To mitigate the negative effect of misalignment, features are aggregated with the attention mechanism determined by individual domain discriminators. Moreover, temporal dilation is designed for sequential non-repeatability to balance the computational efficiency and the possible number of levels. Extensive experimental results show that our proposed method outperforms the state of the art on four benchmark datasets.1
Peipeng Chen, Andy Jinhua Ma
WACV3
2022 Hierarchical feature disentangling network for universal domain adaptation
Peipeng Chen, Yue Gao 0009, Youngsun Pan, Andy Jinhua Ma
Pattern Recognit.6
2022 Revealing Task-Relevant Model Memorization for Source-Protected Unsupervised Domain Adaptation
abstract
Source-data-free unsupervised domain adaptation (SF-UDA) is an approach to improve model performance in the target domain without accessing the source data. Some SF-UDA methods have been proposed and achieved promising results using the information from source-model parameters. However, current research on information security confirms the ability of a well-trained model to memorize its training data. Therefore, SF-UDA methods that access model parameters remain at risk of privacy disclosure. This paper introduces a new topic of source-protected UDA (SP-UDA) that adapts the source model to the target domain while protecting the source-domain data and model privacy. In SP-UDA, only a black-box source model and a set of unlabeled target data are available for domain adaptation. We consider SP-UDA from a new perspective of model memorization revelation. A Source-Protected Generative Model (SPGM) is developed to reveal task-relevant memorization from the source model. SPGM directly distills the inverse process of the source model without access to source-model parameters to meet the privacy protection objective in SP-UDA. The SPGM is learned under the supervision of a newly designed metric named privacy-protected transfer (PPT). The PPT metric measures the transferability and desensitization of the generated data to encourage the SPGM to extract task-relevant information rather than the unintended memorization. A set of desensitized pseudo data is then generated as substitutes for the real source data in UDA. The performance of the proposed method has been validated in four cross-dataset recognition applications with encouraging results.
Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
IEEE Trans. Inf. Forensics Secur.2
2022 Weakly Supervised Liver Tumor Segmentation Using Couinaud Segment Annotation
abstract
Automatic liver tumor segmentation is of great importance for assisting doctors in liver cancer diagnosis and treatment planning. Recently, deep learning approaches trained with pixel-level annotations have contributed many breakthroughs in image segmentation. However, acquiring such accurate dense annotations is time-consuming and labor-intensive, which limits the performance of deep neural networks for medical image segmentation. We note that Couinaud segment is widely used by radiologists when recording liver cancer-related findings in the reports, since it is well-suited for describing the localization of tumors. In this paper, we propose a novel approach to train convolutional networks for liver tumor segmentation using Couinaud segment annotations. Couinaud segment annotations are image-level labels with values ranging from 1 to 8, indicating a specific region of the liver. Our proposed model, namely CouinaudNet, can estimate pseudo tumor masks from the Couinaud segment annotations as pixel-wise supervision for training a fully supervised tumor segmentation model, and it is composed of two components: 1) an inpainting network with Couinaud segment masks which can effectively remove tumors for pathological images by filling the tumor regions with plausible healthy-looking intensities; 2) a difference spotting network for segmenting the tumors, which is trained with healthy-pathological pairs generated by an effective tumor synthesis strategy. The proposed method is extensively evaluated on two liver tumor segmentation datasets. The experimental results demonstrate that our method can achieve competitive performance compared to the fully supervised counterpart and the state-of-the-art methods while requiring significantly less annotation effort.
Fei Lyu 0004, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
IEEE Trans. Medical Imaging2
2022 Learning From Synthetic CT Images via Test-Time Training for Liver Tumor Segmentation
abstract
Automatic liver tumor segmentation could offer assistance to radiologists in liver tumor diagnosis, and its performance has been significantly improved by recent deep learning based methods. These methods rely on large-scale well-annotated training datasets, but collecting such datasets is time-consuming and labor-intensive, which could hinder their performance in practical situations. Learning from synthetic data is an encouraging solution to address this problem. In our task, synthetic tumors can be injected to healthy images to form training pairs. However, directly applying the model trained using the synthetic tumor images on real test images performs poorly due to the domain shift problem. In this paper, we propose a novel approach, namely Synthetic-to-Real Test-Time Training (SR-TTT), to reduce the domain gap between synthetic training images and real test images. Specifically, we add a self-supervised auxiliary task, i.e., two-step reconstruction, which takes the output of the main segmentation task as its input to build an explicit connection between these two tasks. Moreover, we design a scheduled mixture strategy to avoid error accumulation and bias explosion in the training process. During test time, we adapt the segmentation model to each test image with self-supervision from the auxiliary task so as to improve the inference performance. The proposed method is extensively evaluated on two public datasets for liver tumor segmentation. The experimental results demonstrate that our proposed SR-TTT can effectively mitigate the synthetic-to-real domain shift problem in the liver tumor segmentation task, and is superior to existing state-of-the-art approaches.
Fei Lyu 0004, Mang Ye, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
IEEE Trans. Medical Imaging3
2022 Multi-Level Temporal Dilated Dense Prediction for Action Recognition
abstract
3D convolutional neural networks have achieved great success for action recognition. However, large variations of temporal dynamics have not been properly processed and low-level features have not been fully exploited in most existing works. To solve these two problems, we present a general and flexible framework, namely multi-level temporal dilated dense prediction network, which can incorporate with most of existing methods as backbone to improve the temporal modeling capacity. In the proposed method, a novel temporal dilated dense prediction block is designed to fully utilize temporal features with various temporal dilated rates for dense prediction while maintaining relatively low computational cost. To fuse information from low to high levels, our method combines the predictions from multiple such blocks inserted at different stages of the backbone network. In-depth analysis is given to show that short- to long-term temporal dependencies can be captured and multi-level spatio-temporal features are effectively fused for video action recognition by the proposed method. Experimental results demonstrate that our method achieves impressive performance improvement on four publicly available action recognition benchmarks including Charades, Kinetics, Something-Something-V1 and HMDB51.
Manlin Zhang, Andy Jinhua Ma
IEEE Trans. Multim.5
2021 Removing the Background by Adding the Background: Towards Background Robust Self-Supervised Video Representation Learning
abstract
Self-supervised learning has shown great potentials in improving the video representation ability of deep neural networks by getting supervision from the data itself. However, some of the current methods tend to cheat from the background, i.e., the prediction is highly dependent on the video background instead of the motion, making the model vulnerable to background changes. To mitigate the model reliance towards the background, we propose to remove the background impact by adding the background. That is, given a video, we randomly select a static frame and add it to every other frames to construct a distracting video sample. Then we force the model to pull the feature of the distracting video and the feature of the original video closer, so that the model is explicitly restricted to resist the background influence, focusing more on the motion changes. We term our method as Background Erasing (BE). It is worth noting that the implementation of our method is so simple and neat and can be added to most of the SOTA methods without much efforts. Specifically, BE brings 16.4% and 19.1% improvements with MoCo on the severely biased datasets UCF101 and HMDB51, and 14.5% improvement on the less biased dataset Diving48.
Ke Li 0015, Andy Jinhua Ma, Hao Cheng 0012, Feiyue Huang, Rongrong Ji, Xing Sun 0001
CVPR5
2021 Improving Weakly Supervised Object Localization by Uncertainty Estimation of pseudo supervision
abstract
Pseudo bounding box supervision is a promising approach for weakly supervised object localization (WSOL) with only image-level labels. However, the generated pseudo bounding boxes may be inaccurate or even completely non-overlapped with the objects of interest. In this paper, we propose to estimate the uncertainty of pseudo bounding boxes such that the negative impact caused by inaccurate estimation of pseudo supervision could be alleviated for better WSOL. The refined bounding boxes and corresponding variance uncertainties are learned by training a neural network regressor to penalize the erroneous estimations. To the best of our knowledge, this is the first work to incorporate uncertainty information of pseudo bounding boxes for WSOL. Experimental results show that our method not only outperforms previous state-of-the-art methods in CUB-200-2011 and ILSVRC datasets but also gives more precise bounding box prediction when the IoU threshold is higher.
Andy Jinhua Ma, Nanxi Guo
ICME2
2021 Incorporating Decision-level Reconstruction Quality in Adversarial Autoencoders for Anomaly Detection
abstract
Anomaly detection, also known as one-class classification, is a challenging problem due to the absence of abnormal data for training. One of the promising approaches for image anomaly detection is to employ adversarial autoencoders for anomaly measurement by computing the pixel-level reconstruction errors. However, such pixel-level anomaly measure is very sensitive to individual pixels with huge reconstruction difference, resulting in large errors even on normal images. In this paper, we propose a novel anomaly detection method by combining decision-level reconstruction quality with pixel-level reconstruction error. In the proposed method, an adversarial autoencoder is trained to not only measure pixel-level reconstruction error, but also select typical and atypical normal samples from the train set. These selected samples are used to generate positive samples with good reconstruction quality and pseudo negative samples with poor quality respectively for training a classifier which measures the decision-level reconstruction quality of input samples. Experiments on public benchmark datasets show that our method achieves better results by incorporating decision-level reconstruction quality with pixel-level reconstruction error for anomaly detection, and outperforms a wide range of existing methods.
Luyuan Li, Andy Jinhua Ma
IJCNN2
2021 A Segmentation-Assisted Model for Universal Lesion Detection with Partial Labels
Fei Lyu 0004, Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
MICCAI (5)3
2021 Learning Spatio-temporal Representation by Channel Aliasing Video Perception
abstract
In this paper, we propose a novel pretext task namely Channel Aliasing Video Perception (CAVP) for self-supervised video representation learning. The main idea of our approach is to generate channel aliasing videos, which carry different motion cues simultaneously by assembling distinct channels from different videos. With the generated channel aliasing videos, we propose to recognize the number of different motion flows within a channel aliasing video for perception of discriminative motion cues. As a plug-and-play method, the proposed pretext task can be integrated into a co-training framework with other self-supervised learning methods to further improve the performance. Experimental results on publicly available action recognition benchmarks verify the effectiveness of our method for spatio-temporal representation learning.
Manlin Zhang, Andy Jinhua Ma
ACM Multimedia4
2021 AnchorConv: Anchor Convolution for Point Clouds Analysis
Youngsun Pan, Andy Jinhua Ma
PRCV (2)2
2021 Importance-aware personalized learning for early risk prediction using static and dynamic health data
abstract
OBJECTIVE: Accurate risk prediction is important for evaluating early medical treatment effects and improving health care quality. Existing methods are usually designed for dynamic medical data, which require long-term observations. Meanwhile, important personalized static information is ignored due to the underlying uncertainty and unquantifiable ambiguity. It is urgent to develop an early risk prediction method that can adaptively integrate both static and dynamic health data. MATERIALS AND METHODS: Data were from 6367 patients with Peptic Ulcer Bleeding between 2007 and 2016. This article develops a novel End-to-end Importance-Aware Personalized Deep Learning Approach (eiPDLA) to achieve accurate early clinical risk prediction. Specifically, eiPDLA introduces a long short-term memory with temporal attention to learn sequential dependencies from time-stamped records and simultaneously incorporating a residual network with correlation attention to capture their influencing relationship with static medical data. Furthermore, a new multi-residual multi-scale network with the importance-aware mechanism is designed to adaptively fuse the learned multisource features, automatically assigning larger weights to important features while weakening the influence of less important features. RESULTS: Extensive experimental results on a real-world dataset illustrate that our method significantly outperforms the state-of-the-arts for early risk prediction under various settings (eg, achieving an AUC score of 0.944 at 1 year ahead of risk prediction). Case studies indicate that the achieved prediction results are highly interpretable. CONCLUSION: These results reflect the importance of combining static and dynamic health data, mining their influencing relationship, and incorporating the importance-aware mechanism to automatically identify important features. The achieved accurate early risk prediction results save precious time for doctors to timely design effective treatments and improve clinical outcomes.
Qingxiong Tan, Mang Ye, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
J. Am. Medical Informatics Assoc.3
2021 Explainable Uncertainty-Aware Convolutional Recurrent Neural Network for Irregular Medical Time Series
abstract
Influenced by the dynamic changes in the severity of illness, patients usually take examinations in hospitals irregularly, producing a large volume of irregular medical time-series data. Performing diagnosis prediction from the irregular medical time series is challenging because the intervals between consecutive records significantly vary along time. Existing methods often handle this problem by generating regular time series from the irregular medical records without considering the uncertainty in the generated data, induced by the varying intervals. Thus, a novel Uncertainty-Aware Convolutional Recurrent Neural Network (UA-CRNN) is proposed in this article, which introduces the uncertainty information in the generated data to boost the risk prediction. To tackle the complex medical time series with subseries of different frequencies, the uncertainty information is further incorporated into the subseries level rather than the whole sequence to seamlessly adjust different time intervals. Specifically, a hierarchical uncertainty-aware decomposition layer (UADL) is designed to adaptively decompose time series into different subseries and assign them proper weights in accordance with their reliabilities. Meanwhile, an Explainable UA-CRNN (eUA-CRNN) is proposed to exploit filters with different passbands to ensure the unity of components in each subseries and the diversity of components in different subseries. Furthermore, eUA-CRNN incorporates with an uncertainty-aware attention module to learn attention weights from the uncertainty information, providing the explainable prediction results. The extensive experimental results on three real-world medical data sets illustrate the superiority of the proposed method compared with the state-of-the-art methods.
Qingxiong Tan, Mang Ye, Andy Jinhua Ma, Baoyao Yang, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
IEEE Trans. Neural Networks Learn. Syst.3
2020 DATA-GRU: Dual-Attention Time-Aware Gated Recurrent Unit for Irregular Multivariate Time Series
abstract
Due to the discrepancy of diseases and symptoms, patients usually visit hospitals irregularly and different physiological variables are examined at each visit, producing large amounts of irregular multivariate time series (IMTS) data with missing values and varying intervals. Existing methods process IMTS into regular data so that standard machine learning models can be employed. However, time intervals are usually determined by the status of patients, while missing values are caused by changes in symptoms. Therefore, we propose a novel end-to-end Dual-Attention Time-Aware Gated Recurrent Unit (DATA-GRU) for IMTS to predict the mortality risk of patients. In particular, DATA-GRU is able to: 1) preserve the informative varying intervals by introducing a time-aware structure to directly adjust the influence of the previous status in coordination with the elapsed time, and 2) tackle missing values by proposing a novel dual-attention structure to jointly consider data-quality and medical-knowledge. A novel unreliability-aware attention mechanism is designed to handle the diversity in the reliability of different data, while a new symptom-aware attention mechanism is proposed to extract medical reasons from original clinical records. Extensive experimental results on two real-world datasets demonstrate that DATA-GRU can significantly outperform state-of-the-art methods and provide meaningful clinical interpretation.
Qingxiong Tan, Mang Ye, Baoyao Yang, Si-Qi Liu 0003, Andy Jinhua Ma, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Pong C. Yuen
AAAI5
2020 Receptive Field Pyramid Network for Object Detection
abstract
Current state-of-the-art methods usually utilize feature pyramid to provide various receptive fields for detecting objects at different scales. However, the feature maps from low- to high-level layers have large semantic gaps and are with different spatial resolutions, so that their representational capacity differs and noise is introduced when fusing them. To overcome this limitation and carry out better object detection, we design a novel network named Receptive Field Pyramid Network (RFPN). The proposed method is derived based on a receptive field pyramid through dilated convolutions, such that all of the extracted feature maps are with strong semantics and the same resolution. Moreover, we propose a pyramid attention mechanism by iteratively leveraging information from previous receptive fields to give higher responses for objects of interest. Experimental results on publicly available datasets show that the proposed method achieves better results than existing methods for comparison.
Faming Wu, Andy Jinhua Ma, Yangshan Pan, Xiaowei Yan
ICASSP2
2020 Infrared-Visible Person Re-Identification Via Cross-Modality Batch Normalized Identity Embedding And Mutual Learning
abstract
Cross-Modality Infrared-Visible Person Re-identification (IV-REID) is an important application in intelligent video surveillance. Compared to traditional single-modality person re-identification (re-ID) task, IV-REID aims at matching pedestrian images across different spectrum camera views. In this work, we propose a simple but effective framework to reduce the modality discrepancy. First, a batch normalized cross-modality Identity (ID) embedding method is designed to ease the vanishing gradient problem in IVREID. Second, we introduce a novel mutual learning strategy for different single-modality ID Embedding method to further learn discriminative representations under different modalities. Albeit simple, extensive experiments show that our method outperforms the state-of-the-art on RegDB and SYSU-MMOI datasets. Source code is publicly available at: https://giChub.com/linyq17/IV-REID.
Andy Jinhua Ma
ICIP2
2020 Multi-Scale Feature Pyramids for Weakly Supervised Thoracic Disease Localization
abstract
Automatic localization of thoracic diseases has a wide range of applications which can assist radiologists for more efficient and better diagnosis. However, it is still a challenging task to locate the diseases accurately since strong location annotation may not be available and different thoracic diseases may vary in size greatly. In this paper, we propose a novel multi-scale feature pyramids model for weakly supervised disease localization on chest X-ray images. Our model leverages the multi-scale feature maps to learn location representation of lesions by fusing heatmaps generated from all these feature maps. Instead of linear combination of heatmaps, we reconFigure multi-scale feature maps with an Feature Pyramids Network (fpn) structure first. The FPN we conducted is a nonlinear combination of feature hierarchy, adding highly-nonlinear patterns in feature maps, enriching the representation space of heatmaps. Experiments on the ChestXray14 dataset show that the proposed method can significantly improve the 10-calization performance of small-sized diseases (nodules and masses) with competitive localization performance of large-sized diseases (cardiomegaly and pneumonia). Averagely, our method is superior to the state-of-the-art weakly supervised localization algorithms on the dataset, ChestXray14.
Andy Jinhua Ma, Youngsun Pan
ICIP2
2020 Multi-Scale Adversarial Cross-Domain Detection with Robust Discriminative Learning
abstract
Domain shift practically exists in almost all computer vision tasks including object detection, caused by which the performance drops evidently. Most existing methods for domain adaptation are specially designed for classification. For object detection, existing methods separate domain shift into image-level shift and instance-level shift and align image-level feature and instance-level feature respectively. However, we find that there are two problems which remain unsolved yet. First, the scale of objects is not the same even in an image. Second, negative transfer can affect model performance if not handled properly. We improve the performance of cross-domain detection from three perspectives: 1) using multiple dilated convolution kernels with different dilation rate to reduce the image-level domain discrepancy; 2) removing images or instances with low transferability to weaken the influence of negative transfer; 3) diversifying distributions by keeping instances' feature away from each other, and then pull them closer to the center of each category, so that make source samples distribution more dispersed and more robust for cross-domain detection. We test our model with Cityscapes [5], Foggy Cityscape [30] and SIM 10K [18] datasets, experimental results show that our method outperforms the state-of-the-art for object detection under the setting of unsupervised domain adaptation (UDA).
Youngsun Pan, Andy Jinhua Ma
WACV2
2020 Temporal matrix completion with locally linear latent factors for medical applications
Andy Jinhua Ma, Jacky C. P. Chan, Frodo Kin-Sun Chan, Pong C. Yuen, Terry Cheuk-Fung Yip, Yee-Kit Tse, Vincent Wai-Sun Wong, Grace Lai-Hung Wong
Artif. Intell. Medicine1
2020 Adversarial open set domain adaptation via progressive selection of transferable target samples
Andy Jinhua Ma, Yue Gao 0009, Youngsun Pan
Neurocomputing2
2020 Invariant subspace learning for time series data based on dynamic time warping distance
Huiqi Deng, Weifu Chen, Andy Jinhua Ma, Pong C. Yuen, Guo-Can Feng
Pattern Recognit.4
2019 UA-CRNN: Uncertainty-Aware Convolutional Recurrent Neural Network for Mortality Risk Prediction
abstract
Accurate prediction of mortality risk is important for evaluating early treatments, detecting high-risk patients and improving healthcare outcomes. Predicting mortality risk from the irregular clinical time series data is challenging due to the varying time intervals in the consecutive records. Existing methods usually solve this issue by generating regular time series data from the original irregular data without considering the uncertainty in the generated data, caused by varying time intervals. In this paper, we propose a novel Uncertainty-Aware Convolutional Recurrent Neural Network (UA-CRNN), which incorporates the uncertainty information in the generated data to improve the mortality risk prediction performance. To handle the complex clinical time series data with sub-series of different frequencies, we propose to incorporate the uncertainty information into the sub-series level rather than the whole time series data. Specifically, we design a novel hierarchical uncertainty-aware decomposition layer (UADL) to adaptively decompose time series into different sub-series and assign them proper weights according to their reliabilities. Experimental results on two real-world clinical datasets demonstrate that the proposed UA-CRNN method significantly outperforms state-of-the-art methods in both short-term and long-term mortality risk predictions.
Qingxiong Tan, Andy Jinhua Ma, Mang Ye, Baoyao Yang, Huiqi Deng, Vincent Wai-Sun Wong, Yee-Kit Tse, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Jessica Yuet-Ling Ching, Francis Ka-Leung Chan, Pong C. Yuen
CIKM2
2019 Spatial-Temporal Bottom-Up Top-Down Attention Model for Action Recognition
Andy Jinhua Ma
ICIG (1)2
2019 Variation Generalized Feature Learning via Intra-view Variation Adaptation
abstract
This paper addresses the variation generalized feature learning problem in unsupervised video-based person re-identification (re-ID). With advanced tracking and detection algorithms, large-scale intra-view positive samples can be easily collected by assuming that the image frames within the tracking sequence belong to the same person. Existing methods either directly use the intra-view positives to model cross-view variations or simply minimize the intra-view variations to capture the invariant component with some discriminative information loss. In this paper, we propose a Variation Generalized Feature Learning (VGFL) method to learn adaptable feature representation with intra-view positives. The proposed method can learn a discriminative re-ID model without any manually annotated cross-view positive sample pairs. It could address the unseen testing variations with a novel variation generalized feature learning algorithm. In addition, an Adaptability-Discriminability (AD) fusion method is introduced to learn adaptable video-level features. Extensive experiments on different datasets demonstrate the effectiveness of the proposed method.
Jiawei Li 0003, Mang Ye, Andy Jinhua Ma, Pong C. Yuen
IJCAI3
2019 Body Parts Synthesis for Cross-Quality Pose Estimation
abstract
Although encouraging results have been obtained in human pose estimation in recent years, the performance may degrade dramatically when the image quality differs between training and testing data sets. This paper addresses problems in cross-image-quality human pose estimation. To achieve this, we follow an unsupervised domain adaptation approach, in which labels in the target domain are unavailable. Unlike existing unsupervised domain adaptation methods that find label information from unlabeled data, the target pose information (label) is instead generated by synthesizing body parts with similar image-quality of the target domain. A translative dictionary is learned to associate the source and target domains, and a cross-quality adaptation model is developed to refine the source pose estimator using the synthesized target body parts. We perform cross-quality experiments on three data sets with different image quality by using two state-of-the-art pose estimators, and compare the proposed method with five unsupervised domain adaptation methods. Our experimental results show that the proposed method outperforms not only the source pose estimators, but also other unsupervised domain adaptation methods.
Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
IEEE Trans. Circuits Syst. Video Technol.2
2019 Dynamic Graph Co-Matching for Unsupervised Video-Based Person Re-Identification
abstract
Cross-camera label estimation from a set of unlabelled training data is an extremely important component in unsupervised person re-identification (re-ID) systems. With the estimated labels, existing advanced supervised learning methods can be leveraged to learn discriminative re-ID models. In this paper, we utilize the graph matching technique for accurate label estimation due to its advantages in optimal global matching and intra-camera relationship mining. However, the graph structure constructed with non-learnt similarity measurement cannot handle the large cross-camera variations, which leads to noisy and inaccurate label outputs. This paper designs a Dynamic Graph Matching (DGM) framework, which improves the label estimation process by iteratively refining the graph structure with better similarity measurement learnt from intermediate estimated labels. In addition, we design a positive re-weighting strategy to refine the intermediate labels, which enhances the robustness against inaccurate matching output and noisy initial training data. To fully utilize the abundant video information and reduce false matchings, a co-matching strategy is further incorporated into the framework. Comprehensive experiments conducted on three video benchmarks demonstrate that DGM outperforms state-of-the-art unsupervised re-ID methods and yields competitive performance to fully supervised upper bounds.
Mang Ye, Jiawei Li 0003, Andy Jinhua Ma, Liang Zheng 0001, Pong C. Yuen
IEEE Trans. Image Process.3
2018 Domain-Shared Group-Sparse Dictionary Learning for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation has been proved to be a promising approach to solve the problem of dataset bias. To employ source labels in the target domain, it is required to align the joint distributions of source and target data. To do this, the key research problem is to align conditional distributions across domains without target labels. In this paper, we propose a new criterion of domain-shared group-sparsity that is an equivalent condition for conditional distribution alignment. To solve the problem in joint distribution alignment, a domain-shared group-sparse dictionary learning method is developed towards joint alignment of conditional and marginal distributions. A classifier for target domain is trained using the domain-shared group-sparse coefficients and the target-specific information from the target data. Experimental results on cross-domain face and object recognition show that the proposed method outperforms eight state-of-the-art unsupervised domain adaptation algorithms.
Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
AAAI2
2018 A Hybrid Residual Network and Long Short-Term Memory Method for Peptic Ulcer Bleeding Mortality Prediction
Qingxiong Tan, Andy Jinhua Ma, Huiqi Deng, Vincent Wai-Sun Wong, Yee-Kit Tse, Terry Cheuk-Fung Yip, Grace Lai-Hung Wong, Jessica Yuet-Ling Ching, Francis Ka-Leung Chan, Pong C. Yuen
AMIA2
2018 Robust Shapelets Learning: Transform-Invariant Prototypes
Huiqi Deng, Weifu Chen, Andy Jinhua Ma, Pong C. Yuen, Guo-Can Feng
PRCV (3)3
2018 Semi-supervised Region Metric Learning for Person Re-identification
Jiawei Li 0003, Andy Jinhua Ma, Pong C. Yuen
Int. J. Comput. Vis.2
2018 Learning domain-shared group-sparse representation for unsupervised domain adaptation
Baoyao Yang, Andy Jinhua Ma, Pong C. Yuen
Pattern Recognit.2
2017 Dynamic Label Graph Matching for Unsupervised Video Re-identification
abstract
Label estimation is an important component in an unsupervised person re-identification (re-ID) system. This paper focuses on cross-camera label estimation, which can be subsequently used in feature learning to learn robust re-ID models. Specifically, we propose to construct a graph for samples in each camera, and then graph matching scheme is introduced for cross-camera labeling association. While labels directly output from existing graph matching methods may be noisy and inaccurate due to significant cross-camera variations, this paper propose a dynamic graph matching (DGM) method. DGM iteratively updates the image graph and the label estimation process by learning a better feature space with intermediate estimated labels. DGM is advantageous in two aspects: 1) the accuracy of estimated labels is improved significantly with the iterations; 2) DGM is robust to noisy initial training data. Extensive experiments conducted on three benchmarks including the large-scale MARS dataset show that DGM yields competitive performance to fully supervised baselines, and outperforms competing unsupervised learning methods.1
Mang Ye, Andy Jinhua Ma, Liang Zheng 0001, Jiawei Li 0003, Pong C. Yuen
ICCV2
2016 Process Monitoring in the Intensive Care Unit: Assessing Patient Mobility Through Activity Analysis with a Non-Invasive Mobility Sensor
Austin Reiter, Andy Jinhua Ma, Nishi Rawat, Christine Shrock, Suchi Saria
MICCAI (1)2
2015 Joint Sparse Representation and Robust Feature-Level Fusion for Multi-Cue Visual Tracking
abstract
Visual tracking using multiple features has been proved as a robust approach because features could complement each other. Since different types of variations such as illumination, occlusion, and pose may occur in a video sequence, especially long sequence videos, how to properly select and fuse appropriate features has become one of the key problems in this approach. To address this issue, this paper proposes a new joint sparse representation model for robust feature-level fusion. The proposed method dynamically removes unreliable features to be fused for tracking by using the advantages of sparse representation. In order to capture the non-linear similarity of features, we extend the proposed method into a general kernelized framework, which is able to perform feature fusion on various kernel spaces. As a result, robust tracking performance is obtained. Both the qualitative and quantitative experimental results on publicly available videos show that the proposed method outperforms both sparse representation-based and fusion based-trackers.
Xiangyuan Lan, Andy Jinhua Ma, Pong C. Yuen, Rama Chellappa
IEEE Trans. Image Process.2
2015 Cross-Domain Person Reidentification Using Domain Adaptation Ranking SVMs
abstract
This paper addresses a new person reidentification problem without label information of persons under nonoverlapping target cameras. Given the matched (positive) and unmatched (negative) image pairs from source domain cameras, as well as unmatched (negative) and unlabeled image pairs from target domain cameras, we propose an adaptive ranking support vector machines (AdaRSVMs) method for reidentification under target domain cameras without person labels. To overcome the problems introduced due to the absence of matched (positive) image pairs in the target domain, we relax the discriminative constraint to a necessary condition only relying on the positive mean in the target domain. To estimate the target positive mean, we make use of all the available data from source and target domains as well as constraints in person reidentification. Inspired by adaptive learning methods, a new discriminative model with high confidence in target positive mean and low confidence in target negative image pairs is developed by refining the distance model learnt from the source domain. Experimental results show that the proposed AdaRSVM outperforms existing supervised or unsupervised, learning or non-learning reidentification methods without using label information in target cameras. Moreover, our method achieves better reidentification performance than existing domain adaptation methods derived under equal conditional probability assumption.
Andy Jinhua Ma, Jiawei Li 0003, Pong C. Yuen, Ping Li 0001
IEEE Trans. Image Process.1
2014 Semi-Supervised Ranking for Re-identification with Few Labeled Image Pairs
Andy Jinhua Ma, Ping Li 0001
ACCV (4)1
2014 Query Based Adaptive Re-ranking for Person Re-identification
Andy Jinhua Ma, Ping Li 0001
ACCV (5)1
2014 Multi-cue Visual Tracking Using Robust Feature-Level Fusion Based on Joint Sparse Representation
abstract
The use of multiple features for tracking has been proved as an effective approach because limitation of each feature could be compensated. Since different types of variations such as illumination, occlusion and pose may happen in a video sequence, especially long sequence videos, how to dynamically select the appropriate features is one of the key problems in this approach. To address this issue in multi-cue visual tracking, this paper proposes a new joint sparse representation model for robust feature-level fusion. The proposed method dynamically removes unreliable features to be fused for tracking by using the advantages of sparse representation. As a result, robust tracking performance is obtained. Experimental results on publicly available videos show that the proposed method outperforms both existing sparse representation based and fusion-based trackers.
Xiangyuan Lan, Andy Jinhua Ma, Pong C. Yuen
CVPR2
2014 Reduced Analytic Dependency Modeling: Robust Fusion for Visual Recognition
Andy Jinhua Ma, Pong C. Yuen
Int. J. Comput. Vis.1
2013 Domain Transfer Support Vector Ranking for Person Re-identification without Target Camera Label Information
abstract
This paper addresses a new person re-identification problem without the label information of persons under non-overlapping target cameras. Given the matched (positive) and unmatched (negative) image pairs from source domain cameras, as well as unmatched (negative) image pairs which can be easily generated from target domain cameras, we propose a Domain Transfer Ranked Support Vector Machines (DTRSVM) method for re-identification under target domain cameras. To overcome the problems introduced due to the absence of matched (positive) image pairs in target domain, we relax the discriminative constraint to a necessary condition only relying on the positive mean in target domain. By estimating the target positive mean using source and target domain data, a new discriminative model with high confidence in target positive mean and low confidence in target negative image pairs is developed. Since the necessary condition may not truly preserve the discriminability, multi-task support vector ranking is proposed to incorporate the training data from source domain with label information. Experimental results show that the proposed DTRSVM outperforms existing methods without using label information in target cameras. And the top 30 rank accuracy can be improved by the proposed method upto 9.40% on publicly available person re-identification datasets.
Andy Jinhua Ma, Pong C. Yuen, Jiawei Li 0003
ICCV1
2013 Linear Dependency Modeling for Classifier Fusion and Feature Combination
abstract
This paper addresses the independent assumption issue in fusion process. In the last decade, dependency modeling techniques were developed under a specific distribution of classifiers or by estimating the joint distribution of the posteriors. This paper proposes a new framework to model the dependency between features without any assumption on feature/classifier distribution, and overcomes the difficulty in estimating the high-dimensional joint density. In this paper, we prove that feature dependency can be modeled by a linear combination of the posterior probabilities under some mild assumptions. Based on the linear combination property, two methods, namely, Linear Classifier Dependency Modeling (LCDM) and Linear Feature Dependency Modeling (LFDM), are derived and developed for dependency modeling in classifier level and feature level, respectively. The optimal models for LCDM and LFDM are learned by maximizing the margin between the genuine and imposter posterior probabilities. Both synthetic data and real datasets are used for experiments. Experimental results show that LCDM and LFDM with dependency modeling outperform existing classifier level and feature level combination methods under nonnormal distributions and on four real databases, respectively. Comparing the classifier level and feature level fusion methods, LFDM gives the best performance.
Andy Jinhua Ma, Pong C. Yuen, Jian-Huang Lai
IEEE Trans. Pattern Anal. Mach. Intell.1
2013 Supervised Spatio-Temporal Neighborhood Topology Learning for Action Recognition
abstract
Supervised manifold learning has been successfully applied to action recognition, in which class label information could improve the recognition performance. However, the learned manifold may not be able to well preserve both the local structure and global constraint of temporal labels in action sequences. To overcome this problem, this paper proposes a new supervised manifold learning algorithm called supervised spatio-temporal neighborhood topology learning (SSTNTL) for action recognition. By analyzing the topological characteristics in the context of action recognition, we propose to construct the neighborhood topology using both supervised spatial and temporal pose correspondence information. Employing the property in locality preserving projection (LPP), SSTNTL solves the generalized eigenvalue problem to obtain the best projections that not only separates data points from different classes, but also preserves local structures and temporal pose correspondence of sequences from the same class. Experimental results demonstrate that SSTNTL outperforms the manifold embedding methods with other topologies or local discriminant information. Moreover, compared with state-of-the-art action recognition algorithms, SSTNTL gives convincing performance for both human and gesture action recognition.
Andy Jinhua Ma, Pong C. Yuen, Wilman W. W. Zou, Jian-Huang Lai
IEEE Trans. Circuits Syst. Video Technol.1
2012 Reduced Analytical Dependency Modeling for Classifier Fusion
Andy Jinhua Ma, Pong C. Yuen
ECCV (3)1
2011 Linear dependency modeling for feature fusion
abstract
This paper addresses the independent assumption issue in fusion process. In the last decade, dependency modeling techniques were developed under a specific distribution of classifiers. This paper proposes a new framework to model the dependency between features without any assumption on feature/classifier distribution. In this paper, we prove that feature dependency can be modeled by a linear combination of the posterior probabilities under some mild assumptions. Based on the linear combination property, two methods, namely Linear Classifier Dependency Modeling (LCDM) and Linear Feature Dependency Modeling (LFDM), are derived and developed for dependency modeling in classifier level and feature level, respectively. The optimal models for LCDM and LFDM are learned by maximizing the margin between the genuine and imposter posterior probabilities. Both synthetic data and real datasets are used for experiments. Experimental results show that LFDM outperforms all existing combination methods.
Andy Jinhua Ma, Pong C. Yuen
ICCV1