VLDB 2026 Research / reviewers in the wild / expert
Xiang Xiang 0001
dblp:83/2686-1
· DBLP profile ↗
38ranked-venue papers
12as first author
29since 2021 · last 2026
0000-0003-0606-6193ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 20 since 2021Artificial intelligence and machine learning · 21 · 6 first-author · 20 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Random Amalgamation of Adapters for Flatter Loss Landscapes: Towards Class-Incremental Learning with Better StabilityabstractClass-incremental learning (CIL) enables models to continuously learn from streaming data while mitigating catastrophic forgetting of prior knowledge. Our research reveals that the CIL performance of pre-trained models (PTMs) varies significantly across different datasets, a phenomenon underexplored in existing studies. Through visualization, we observe that flatter loss landscapes correlate with superior CIL performance. This insight motivates us to enhance PTMs' CIL capability by promoting loss landscapes' flatness. Initially, we propose independently optimizing multiple adapter branches to equip PTMs with diverse learnable parameters, thereby improving stability during parameter updates. However, given computational and memory constraints, the number of adapters a PTM can accommodate is limited. To address this, we introduce a training strategy with randomized adapter amalgamation (RAA), compelling the model to maintain low loss across a broader and more continuous parameter space, significantly enhancing flatness. Furthermore, we refine existing sharpness-aware minimization techniques to further optimize the loss landscapes. Our extensive experiments and visualization results validate the efficacy of the method, resulting in the state-of-the-art (SOTA) performance. Xiang Xiang 0001, Jiaqi Gui |
AAAI | 2 |
| 2026 | AerialFusion: Co-Motion-Driven Unified Registration and Fusion on Multi-modal Data Streams from Aerial ViewabstractAerial multi-modal visual streams registration and fusion can generate more comprehensive scene information representations for UAVs' cross-modal perception. However, current challenges lie primarily in the essential difficulty of joint spatiotemporal representation learning from dynamic background and moving targets, and a critical shortage exists in large-scale, well-annotated multi-modal visual streams benchmark for UAV platforms. In this paper, we propose AerialFusion, a co-motion-driven unified UAVs visual streams registration and fusion that fully mines modality-invariant common features based on motion-aware, enabling spatiotemporally coherent registration and fusion. Specifically, 1) a Skewed Motion Distribution Field Co-Motion-Driven Image Registration, 2) a Co-Motion Generative Fusion, 3) a Streams-based Unified Learning. Furthermore, we introduce EUM3D, a registration and fusion benchmark for UAVs cross-modal perception. This benchmark contains 60 synchronized visible-infrared visual streams, or 122k spatially and temporally aligned pairs, most of which were taken at low-light scenes. And EUM3D provides pixel-level alignment guarantees via perspective-transform ground-truth. Extensive experiments reveal that AerialFusion surpasses current focus on image and static background fusion methods in aerial sequence scenarios, addressing spatiotemporal mismatches while suppressing cross-modal interference. Junhui Qiu, Xiang Xiang 0001, Hongyun Wang, Jiaqi Gui |
AAAI | 2 |
| 2026 | Frequency-Aware Domain GeneralizationabstractDeep Neural Networks (DNNs) exhibit surprising zero-shot generalization and emergent phenomena across various tasks. However, the underlying mechanisms behind these behaviors remain unclear. By analyzing the perception of image frequencies by DNNs, we establish the association between the generalization behavior and frequency-aware regions. DNNs with stronger generalization exhibit wider frequency-aware regions. Therefore, we improve the generalization performance by broadening the frequency awareness. Specifically, we enable DNNs to learn the relations between high-frequency components and semantic labels through frequency decomposition and mixup. Based on hierarchical feature alignment, we allow larger submodels to guide the frequency awareness of smaller submodels. Beyond training, we ensemble submodels to extract features from different frequency bands to enrich DNNs' frequency awareness during inference. We validate the effectiveness of our proposed method in image classification and object detection tasks in single and multi-source domain generalization scenarios. We also demonstrate the plug-and-play scalability of our method across existing approaches and different DNNs. Xiang Xiang 0001, Trac D. Tran |
IEEE Trans. Image Process. | 1 |
| 2025 | Overcoming Shortcut Problem in VLM for Robust Out-of-Distribution DetectionabstractVision-language models (VLMs), such as CLIP, have shown remarkable capabilities in downstream tasks. However, the coupling of semantic information between the foreground and the background in images leads to significant shortcut issues that adversely affect out-of-distribution (OOD) detection abilities. When confronted with a background OOD sample, VLMs are prone to misidentifying it as in-distribution (ID) data. In this paper, we analyze the OOD problem from the perspective of shortcuts in VLMs and propose OSPCoOp which includes background decoupling and mask-guided region regularization. We first decouple images into ID-relevant and ID-irrelevant regions and utilize the latter to generate a large number of augmented OOD background samples as pseudo-OOD supervision. We then use the masks from background decoupling to adjust the model’s attention, minimizing its focus on ID-irrelevant regions. To assess the model’s robustness against background interference, we introduce a new OOD evaluation dataset, ImageNet-Bg, which solely consists of background images with all ID-relevant regions removed. Our method demonstrates exceptional performance in few-shot scenarios, achieving strong results even in one-shot setting, and outperforms existing methods. The code and proposed ImageNet-Bg are available at https://github.com/HAIV-Lab/OSPCoOp_Imagenet-bg. Xiang Xiang 0001, Yifan Liang |
CVPR | 2 |
| 2025 | PTTA: Purifying Malicious Samples for Test-Time Model AdaptationabstractTest-Time Adaptation (TTA) enables deep neural networks to adapt to arbitrary distributions during inference. Existing TTA algorithms generally tend to select benign samples that help achieve robust online prediction and stable self-training. Although malicious samples that would undermine the model's optimization should be filtered out, it also leads to a waste of test data. To alleviate this issue, we focus on how to make full use of the malicious test samples for TTA by transforming them into benign ones, and propose a plug-and-play method, PTTA. The core of our solution lies in the purification strategy, which retrieves benign samples having opposite effects on the objective function to perform Mixup with malicious samples, based on a saliency indicator for encoding benign and malicious data. This strategy results in effective utilization of the information in malicious samples and an improvement of the models' online test accuracy. In this way, we can directly apply the purification loss to existing TTA algorithms without the need to carefully adjust the sample selection threshold. Extensive experiments on four types of TTA tasks as well as classification, segmentation, and adversarial defense demonstrate the effectiveness of our method. Code is available at https://github.com/HAIV-Lab/PTTA. Xiang Xiang 0001 |
ICML | 3 |
| 2025 | Decoupled Entropy MinimizationabstractEntropy Minimization (EM) is beneficial to reducing class overlap, bridging domain gap, and restricting uncertainty for various tasks in machine learning, yet its potential is limited. To study the internal mechanism of EM, we reformulate and decouple the classical EM into two parts with opposite effects: cluster aggregation driving factor (CADF) rewards dominant classes and prompts a peaked output distribution, while gradient mitigation calibrator (GMC) penalizes high-confidence classes based on predicted probabilities. Furthermore, we reveal the limitations of classical EM caused by its coupled formulation: 1) reward collapse impedes the contribution of high-certainty samples in the learning process, and 2) easy-class bias induces misalignment between output distribution and label distribution. To address these issues, we propose Adaptive Decoupled Entropy Minimization (AdaDEM), which normalizes the reward brought from CADF and employs a marginal entropy calibrator (MEC) to replace GMC. AdaDEM outperforms DEM*, an upper-bound variant of classical EM, and achieves superior performance across various imperfectly supervised learning tasks in noisy and dynamic environments. Xiang Xiang 0001 |
NeurIPS | 3 |
| 2025 | Few-Shot Font Generation via Attribute-Guided Diffusion with Style Contrastive Learning
Xiang Xiang 0001, Xiaofei Liao |
PRCV (6) | 2 |
| 2025 | Learning Visual-Semantic Hierarchical Attribute Space for Interpretable Open-Set RecognitionabstractIn the field of open-set recognition, conventional models often focus on addressing challenges within a single hierarchical category, and these methods frequently lack inter-pretability. In this paper, we propose a novel solution that utilizes attributes and hierarchical relationships to achieve interpretable open-set recognition. Our method is centered around the visual-semantic attribute space. By leveraging hierarchy division, we can decompose the attributes into more granular components, thereby yielding additional performance improvements. When confronted with an unfamiliar object, our method not only classifies it as an unknown category but also provides insights into the broader category and its associated attributes. This capability enhances interpretability by offering valuable information regarding the potential category and characteristics of the object. Experimental results demonstrate great performance improvements compared to existing methods. Xiang Xiang 0001 |
WACV | 2 |
| 2025 | Facial Action Unit Detection by Adaptively Constraining Self-Attention and Causally Deconfounding Sample
Zhiwen Shao, Hancheng Zhu, Yong Zhou 0003, Xiang Xiang 0001, Bing Liu 0016, Rui Yao 0006, Lizhuang Ma |
Int. J. Comput. Vis. | 4 |
| 2025 | GroupRF: Panoptic Scene Graph Generation with group relation tokensabstractPanoptic Scene Graph Generation (PSG) aims to predict a variety of relations between pairs of objects within an image, and indicate the objects by panoptic segmentation masks instead of bounding boxes . Existing PSG methods attempt to straightforwardly fuse the object tokens for relation prediction, thus failing to fully utilize the interaction between the pairwise objects. To address this problem, we propose a novel framework named Group R elation F ormer (GroupRF) to capture the fine-grained inter-dependency among all instances. Our method introduce a set of learnable tokens termed group rln tokens, which exploit fine-grained contextual interaction between object tokens with multiple attentive relations. In the process of relation prediction, we adopt multiple triplets to take advantage of the fine-grained interaction included in group rln tokens. We conduct comprehensive experiments on OpenPSG dataset, which show that our method outperforms the previous state-of-the-art method. Furthermore, we also show the effectiveness of our framework by ablation studies. Our code is available at https://github.com/WHY-student/GroupRF . Hongyun Wang, Jiachen Li 0002, Xiang Xiang 0001, Qing Xie 0002, Yanchun Ma, Yongjian Liu |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Aligning Logits Generatively for Principled Black-Box Knowledge Distillation in the WildabstractBlack-Box Knowledge Distillation (B2KD) is a conservative task in cloud-to-edge model compression, emphasizing the protection of data privacy and model copyrights on both the cloud and edge. With invisible data and models hosted on the server, B2KD aims to utilize only the API queries of the teacher model's inference results in the cloud to effectively distill a lightweight student model deployed on edge devices. B2KD faces challenges such as limited Internet exchange and edge-cloud disparity in data distribution. To address these issues, we theoretically provide a new optimization direction from logits to cell boundary, different from direct logits alignment, and formalize a workflow comprising deprivatization, distillation, and adaptation at test time. Guided by this, we propose a method, Mapping-Emulation KD (MEKD), to enhance the robust prediction and anti-interference capabilities of the student model on edge devices for any unknown data distribution in real-world scenarios. Our method does not differentiate between treating soft or hard responses and consists of: 1) deprivatization: emulating the inverse mapping of the teacher function with a generator, 2) distillation: aligning low-dimensional logits of the teacher and student models by reducing the distance of high-dimensional image points, and 3) adaptation: correcting the student's online prediction bias through a graph propagation-based only-forward test-time adaptation algorithm. Our method demonstrates inspiring performance for edge model distillation and adaptation across different teacher-student pairs. We validate the effectiveness of our method on multiple image recognition benchmarks and various Deep Neural Network models, achieving state-of-the-art performance and showcasing its practical value in remote sensing image recognition applications. Xiang Xiang 0001, Dongrui Wu, Zhigang Zeng, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Enhanced Dual-Pattern Matching With Vision-Language Representation for Out-of-Distribution DetectionabstractOut-of-distribution (OOD) detection presents a significant challenge in deploying pattern recognition and machine learning models, as they frequently fail to generalize to data from unseen distributions. Recent advancements in vision-language models (VLMs), particularly CLIP, have demonstrated promising results in OOD detection through their rich multimodal representations. However, current CLIP-based OOD detection methods predominantly rely on single-modality in-distribution (ID) data (e.g., textual cues), overlooking the valuable information contained in ID visual cues. In this work, we demonstrate that incorporating ID visual information is crucial for unlocking CLIP's full potential in OOD detection. We propose a novel approach, Dual-Pattern Matching (DPM), which effectively adapts CLIP for OOD detection by jointly exploiting both textual and visual ID patterns. Specifically, DPM refines visual and textual features through the proposed Domain-Specific Feature Aggregation (DSFA) and Prompt Enhancement (PE) modules. Subsequently, DPM stores class-wise textual features as textual patterns and aggregates ID visual features as visual patterns. During inference, DPM calculates similarity scores relative to both patterns to identify OOD data. Furthermore, we enhance DPM with lightweight adaptation mechanisms to further boost OOD detection performance. Comprehensive experiments demonstrate that DPM surpasses state-of-the-art methods on multiple benchmarks, highlighting the effectiveness of leveraging multimodal information for OOD detection. The proposed dual-pattern approach provides a simple yet robust framework for leveraging vision-language representations in OOD detection tasks. Xiang Xiang 0001, Zhigang Zeng, Xilin Chen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | High-level LoRA and hierarchical fusion for enhanced micro-expression recognition
Zhiwen Shao, Yong Zhou 0003, Xiang Xiang 0001, Jian Li 0054, Bing Liu 0016, Dit-Yan Yeung |
Vis. Comput. | 4 |
| 2024 | Masking Cascaded Self-attentions for Few-Shot Font-Generation Transformer
Xiang Xiang 0001 |
ACCV (5) | 2 |
| 2024 | Aligning Logits Generatively for Principled Black-Box Knowledge DistillationabstractBlack-Box Knowledge Distillation (B2KD) is a formu-lated problem for cloud-to-edge model compression with in-visible data and models hosted on the server. B2KD faces challenges such as limited Internet exchange and edge-cloud disparity of data distributions. In this paper, we for-malize a two-step workflow consisting of deprivatization and distillation, and theoretically provide a new optimization direction from logits to cell boundary different from direct logits alignment. With its guidance, we propose a new method Mapping-Emulation KD (MEKD) that distills a black-box cumbersome model into a lightweight one. Our method does not differentiate between treating soft or hard responses, and consists of: 1) deprivatization: emulating the inverse mapping of the teacher function with a genera-tor, and 2) distillation: aligning low-dimensional logits of the teacher and student models by reducing the distance of high-dimensional image points. For different teacher-student pairs, our method yields inspiring distillation per-formance on various benchmarks, and outperforms the pre-vious state-of-the-art approaches. Xiang Xiang 0001, Yuchuan Wu |
CVPR | 2 |
| 2024 | Semantically-Shifted Incremental Adapter-Tuning is A Continual ViTransformerabstractClass-incremental learning (CIL) aims to enable models to continuously learn new classes while overcoming catas-trophic forgetting. The introduction of pre-trained models has brought new tuning paradigms to CIL. In this paper, we revisit different parameter-efficient tuning (PET) meth-ods within the context of continual learning. We observe that adapter tuning demonstrates superiority over prompt-based methods, even without parameter expansion in each learning session. Motivated by this, we propose incremen-tally tuning the shared adapter without imposing parameter update constraints, enhancing the learning capacity of the backbone. Additionally, we employ feature sampling from stored prototypes to retrain a unified classifier, further im-proving its performance. We estimate the semantic shift of old prototypes without access to past samples and update stored prototypes session by session. Our proposed method eliminates model expansion and avoids retaining any im-age samples. It surpasses previous pre-trained model-based CIL methods and demonstrates remarkable continual learning capabilities. Experimental results on five CIL bench-marks validate the effectiveness of our approach, achieving state-of-the-art (SOTA) performance. Yuwen Tan, Qinhao Zhou, Xiang Xiang 0001, Yuchuan Wu |
CVPR | 3 |
| 2024 | Vision-Language Dual-Pattern Matching for Out-of-Distribution Detection
Xiang Xiang 0001 |
ECCV (85) | 3 |
| 2024 | Stay Open: Calibrating Weights Continuously for Detecting Out-of-Distribution Objects on the Fly
Xiang Xiang 0001, Qinhao Zhou |
PRCV (12) | 1 |
| 2024 | Expanding Hyperspherical Space for Few-Shot Class-Incremental LearningabstractIn today’s ever-changing world, the ability of machine learning models to continually learn new data without forgetting previous knowledge is of utmost importance. However, in the scenario of few-shot class-incremental learning (FSCIL), where models have limited access to new instances, this task becomes even more challenging. Current methods use prototypes as a replacement for classifiers, where the cosine similarity of instances to these prototypes is used for prediction. However, we have identified that the embedding space created by using the relu activation function is incomplete and crowded for future classes. To address this issue, we propose the Expanding Hyperspherical Space (EHS) method for FSCIL. In EHS, we utilize an odd-symmetric activation function to ensure the completeness and symmetry of embedding space. Additionally, we specify a region for base classes and reserve space for unseen future classes, which increases the distance between class distributions. Pseudo instances are also used to enable the model to anticipate possible upcoming samples. During inference, we provide rectification to the confidence to prevent bias towards base classes. We conducted experiments on benchmark datasets such as CIFAR100 and miniImageNet, which demonstrate that our proposed method achieves state-of-the-art performance. Xiang Xiang 0001 |
WACV | 2 |
| 2024 | Cross-Domain Few-Shot Incremental Learning for Point-Cloud RecognitionabstractSensing 3D objects is critical when 2D object recognition is not accessible. A robot pre-trained on a large point-cloud dataset will encounter unseen classes of 3D objects after deploying it. Therefore, the robot should be able to learn continuously in real-world scenarios. Few-shot class-incremental learning (FSCIL) requires the model to learn from few-shot new examples continually and not forget past classes. However, there is an implicit but strong assumption in the FSCIL that the distribution of the base and incremental classes is the same. In this paper, we focus on cross-domain FSCIL for point-cloud recognition. We decompose the catastrophic forgetting into base class forgetting and incremental class forgetting and alleviate them separately. We utilize the base model to discriminate base samples and new samples by treating base samples as in-distribution samples, and new objects as out-of-distribution samples. We retain the base model to avoid catastrophic forgetting of base classes and train an extra domain-specific module for all new samples to adapt to new classes. At inference, we first discriminate whether the sample belongs to the base class or the new class. Once classified at the model level, test samples are then passed to the corresponding model for class-level classification. To better mitigate the forgetting of new classes, we adopt the soft label and hard label replay together. Extensive experiments on synthetic-to-real incremental 3D datasets show that our proposed method can balance the performance between the base and new objects and outperforms the previous state-of-the-art methods. Yuwen Tan, Xiang Xiang 0001 |
WACV | 2 |
| 2024 | Curricular-balanced long-tailed learning
Xiang Xiang 0001, Xilin Chen 0001 |
Neurocomputing | 1 |
| 2023 | Decoupling MaxLogit for Out-of-Distribution DetectionabstractIn machine learning, it is often observed that standard training outputs anomalously high confidence for both in-distribution (ID) and out-of-distribution (OOD) data. Thus, the ability to detect OOD samples is critical to the model deployment. An essential step for OOD detection is post-hoc scoring. MaxLogit is one of the simplest scoring functions which uses the maximum logits as OOD score. To provide a new viewpoint to study the logit-based scoring function, we reformulate the logit into cosine similarity and logit norm and propose to use MaxCosine and MaxNorm. We empirically find that MaxCosine is a core factor in the effectiveness of MaxLogit. And the performance of MaxLogit is encumbered by MaxNorm. To tackle the problem, we propose the Decoupling MaxLogit (DML) for flexibility to balance MaxCosine and MaxNorm. To further embody the core of our method, we extend DML to DML+ based on the new insights that fewer hard samples and compact feature space are the key components to make logit-based methods effective. We demonstrate the effectiveness of our logit-based OOD detection methods on CIFAR-10, CIFAR-100 and ImageNet and establish state-of-the-art performance. Xiang Xiang 0001 |
CVPR | 2 |
| 2023 | Saliency Regularization for Self-Training with Partial AnnotationsabstractPartially annotated images are easy to obtain in multi-label classification. However, unknown labels in partially annotated images exacerbate the positive-negative imbalance inherent in multi-label classification, which affects supervised learning of known labels. Most current methods require sufficient image annotations, and do not focus on the imbalance of the labels in the supervised training phase. In this paper, we propose saliency regularization (SR) for a novel self-training framework. In particular, we model saliency on the class-specific maps, and strengthen the saliency of object regions corresponding to the present labels. Besides, we introduce consistency regularization to mine unlabeled information to complement unknown labels with the help of SR. It is verified to alleviate the negative dominance caused by the imbalance, and achieve state-of-the-art performance on Pascal VOC 2007, MS-COCO, VG-200, and OpenImages V3. Shouwen Wang, Xiang Xiang 0001, Zhigang Zeng |
ICCV | 3 |
| 2023 | Coupling Bracket Segmentation and Tooth Surface Reconstruction on 3D Dental Models
Yuwen Tan, Xiang Xiang 0001, Hongyi Jing, Shiyang Ye, Chaoran Xue |
MICCAI (6) | 2 |
| 2023 | Hierarchical Task-Incremental Learning with Feature-Space Initialization Inspired by Neural Collapse
Qinhao Zhou, Xiang Xiang 0001 |
Neural Process. Lett. | 2 |
| 2022 | Hierarchical Memory Learning for Fine-Grained Scene Graph Generation
Youming Deng, Yansheng Li 0001, Yongjun Zhang 0002, Xiang Xiang 0001, Jian Wang 0108, Jingdong Chen, Jiayi Ma 0001 |
ECCV (27) | 4 |
| 2022 | Coarse-To-Fine Incremental Few-Shot Learning
Xiang Xiang 0001, Yuwen Tan, Alan L. Yuille, Gregory D. Hager |
ECCV (31) | 1 |
| 2022 | Weakly Supervised Object Detection Based on Active Learning
Xiang Xiang 0001, Baochang Zhang 0001, Xuhui Liu, Jianying Zheng, Qinglei Hu |
Neural Process. Lett. | 2 |
| 2022 | Imbalanced regression for intensity series of pain expression from videos by regularizing spatio-temporal face nets
Xiang Xiang 0001, Feng Wang 0015, Yuwen Tan, Alan L. Yuille |
Pattern Recognit. Lett. | 1 |
| 2020 | Long-Short Graph Memory Network for Skeleton-based Action RecognitionabstractCurrent studies have shown the effectiveness of long short-term memory network (LSTM) for skeleton-based human action recognition in capturing temporal and spatial features of the skeleton sequence. Nevertheless, it still remains challenging for LSTM to extract the latent structural dependency among nodes. In this paper, we introduce a new long-short graph memory network (LSGM) to improve the capability of LSTM to model the skeleton sequence - a type of graph data. Our proposed LSGM can learn high-level temporal-spatial features end-to-end, enabling LSTM to extract the spatial information that is neglected but intrinsic to the skeleton graph data. To improve the discriminative ability of the temporal and spatial module, we use a calibration module termed as graph temporal-spatial calibration (GTSC) to calibrate the learned temporal-spatial features. By integrating the two modules into the same framework, we obtain a stronger generalization capability in processing dynamic graph data and achieve a significant performance improvement on the NTU and SYSU dataset. Experimental results have validated the effectiveness of our proposed LSGM+GTSC model in extracting temporal and spatial information from dynamic graph data.1 Junqin Huang, Zhenhuan Huang, Xiang Xiang 0001, Baochang Zhang 0001 |
WACV | 3 |
| 2018 | S3D: Stacking Segmental P3D for Action Quality AssessmentabstractAction quality assessment is crucial in areas of sports, surgery and assembly line where action skills can be evaluated. In this paper, we propose the Segment-based P3D-fused network S3D built-upon ED-TCN and push the performance on the UNLV-Dive dataset by a significant margin. We verify that segment-aware training performs better than full-video training which turns out to focus on the water spray. We show that temporal segmentation can be embedded with few efforts. Xiang Xiang 0001, Austin Reiter, Gregory D. Hager, Trac D. Tran |
ICIP | 1 |
| 2018 | Linear Disentangled Representation Learning for Facial ActionsabstractLimited annotated data available for the recognition of facial expression and particularly action units makes it hard to train a deep network which can learn disentangled invariant features. However, a supervised linear model is undemanding in terms of training data. In this paper, we propose an elegant linear model to untangle facial actions from expressive face videos which contain a mixture of linearly-representable attributes. Previous attempts require an explicit decoupling of identity and expression which is practically inexact. Instead, we exploit the low-rank property across frames to implicitly subtract the intrinsic neutral face, which are modeled jointly with sparse representation only on the residual expression components. On CK+, our one-shot C-HiSLR on raw-face pixel-intensities performs far more competitive than conventional shape+SVM models with landmark detection and two-stepped SRC of the same type yet applied on manually prepared expression components. It is also comparable with the piecewise linear model DCS and temporal models, such as CRF and Bayes nets. We apply it to action unit (AU) recognition on MPI-VDB achieving a decent performance. As expression is a mixture of AUs, the result gives hopes of approximating an expression using a piecewise linear model. Xiang Xiang 0001, Trac D. Tran |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2017 | Regularizing face verification nets for pain intensity regressionabstractLimited labeled data are available for the research of estimating facial expression intensities. For instance, the ability to train deep networks for automated pain assessment is limited by small datasets with labels of patient-reported pain intensities. Fortunately, fine-tuning from a data-extensive pre-trained domain, such as face verification, can alleviate this problem. In this paper, we propose a network that fine-tunes a state-of-the-art face verification network using a regularized regression loss and additional data with expression labels. In this way, the expression intensity regression task can benefit from the rich feature representations trained on a huge amount of data for face verification. The proposed regularized deep regressor is applied to estimate the pain expression intensity and verified on the widely-used UNBC-McMaster Shoulder-Pain dataset, achieving the state-of-the-art performance. A weighted evaluation metric is also proposed to address the imbalance issue of different pain intensities. Feng Wang 0015, Xiang Xiang 0001, Trac D. Tran, Austin Reiter, Gregory D. Hager, Harry Quon, Jian Cheng 0003, Alan L. Yuille |
ICIP | 2 |
| 2017 | Supervised hashing with jointly learning embedding and quantizationabstractCompared with unsupervised hashing, supervised hashing commonly illustrates better accuracy in many real applications by leveraging semantic (label) information. However, it is tough to solve the supervised hashing problem directly because it is essentially a discrete optimization problem. Some other works try to solve the discrete optimization problem directly using binary quadratic programming, but they are typically too complicated and time-consuming while some supervised hashing methods have to solve a relaxed continuous optimization problem by dropping the discrete constraints. However, these methods typically suffer from poor performance due to the errors caused by the relaxation manner. In this paper based on the general two-step framework: learning binary embedded codes and learning hash functions, we propose a new method to solve the problem introduced by relaxing the cost function. Inspired by the property of rotation invariance of learning embedding features, our method tries to jointly learn similarity-preserving representation and rotation transformation for better quantization alternatively. In experiments, our method shows significant improvement. Compared with the methods based on discrete optimization our methods obtains the competitive performance and even achieves the state-of-the-art performance in some image retrieval applications. Feng Wang 0015, Xiang Xiang 0001, Trac D. Tran |
ICIP | 3 |
| 2017 | NormFace: L2 Hypersphere Embedding for Face VerificationabstractThanks to the recent developments of Convolutional Neural Networks, the performance of face verification methods has increased rapidly. In a typical face verification method, feature normalization is a critical step for boosting performance. This motivates us to introduce and study the effect of normalization during training. But we find this is non-trivial, despite normalization being differentiable. We identify and study four issues related to normalization through mathematical analysis, which yields understanding and helps with parameter settings. Based on this analysis we propose two strategies for training using normalized features. The first is a modification of softmax loss, which optimizes cosine similarity instead of inner-product. The second is a reformulation of metric learning by introducing an agent vector for each class. We show that both strategies, and small variants, consistently improve performance by between 0.2% to 0.4% on the LFW dataset based on two models. This is significant because the performance of the two models on LFW dataset is close to saturation at over 98%. Feng Wang 0015, Xiang Xiang 0001, Jian Cheng 0003, Alan L. Yuille |
ACM Multimedia | 2 |
| 2015 | Hierarchical Sparse and Collaborative Low-Rank representation for emotion recognitionabstractIn this paper, we design a Collaborative-Hierarchical Sparse and Low-Rank (C-HiSLR) model that is natural for recognizing human emotion in visual data. Previous attempts require explicit expression components, which are often unavailable and difficult to recover. Instead, our model exploits the low-rank property to subtract neutral faces from expressive facial frames as well as performs sparse representation on the expression components with group sparsity enforced. For the CK+ dataset, C-HiSLR on raw expressive faces performs as competitive as the Sparse Representation based Classification (SRC) applied on manually prepared emotions. Our C-HiSLR performs even better than SRC in terms of true positive rate. Xiang Xiang 0001, Minh Dao, Gregory D. Hager, Trac D. Tran |
ICASSP | 1 |
| 2012 | Online Web-Data-Driven Segmentation of Selected Moving Objects in Videos
Xiang Xiang 0001, Jiebo Luo 0001 |
ACCV (2) | 1 |
| 2008 | Intelligent Target Tracking and Shooting System with Mean ShiftabstractTracking moving targets in sequence images is an essential key technology and one of the hot research topics in Computer Vision. This system is based on an embedded system platform named Embedded Star and makes full use of the OpenCV (Intel® open-source computer vision library)to implement and optimize the Mean Shift tracking algorithm. At last, it achieves the objective of real-time tracking and shooting of moving targets, and can be used in sports photography, real-time monitoring and so on. The test data has indicated that it can direct the camera through controlling the cloud terrace to track both rigid and non-rigid targets. Xiang Xiang 0001, Du Zeng |
ISPA | 1 |