VLDB 2026 Research / reviewers in the wild / expert
Yanyan Liang 0001
dblp:43/10437
· DBLP profile ↗
51ranked-venue papers
0as first author
46since 2021 · last 2026
0000-0002-5780-8540ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 21 since 2021Security and privacy · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PointMC: Multi-view Consistent Encoding and Center-Global Feature Fusion for Point Clouds UnderstandingabstractPoint cloud tasks have recently benefited from Mamba-based architecture, which leverage state space modeling to achieve strong performance. Previous studies have primarily focused on network design while overlooking the importance of position encoding and relying on coarse-grained geometric feature aggregation. The former leads to semantic ambiguity due to inconsistent spatial relationships, while the latter results in geometric feature dispersion by overlooking fine-grained local geometric details. To tackle the above problem, we propose a novel framework, PointMC, including Multi-view Consistent Learnable Position Encoding (MCLPE) and Center-Global Feature Fusion (CGFF), to provide semantically coherent positional guidance for inter-patch and enable fine-grained geometric structure aggregation within intra-patch regions. Specifically, the proposed MCLPE module is inspired by a spatial structure modeling mechanism guided by physical constraints, leverages multi-view virtual reconstruction and a learnable strategy to dynamically constrain spatial relationships along patch boundaries, thereby enhancing the semantic consistency and representational clarity across inter-patch regions. Furthermore, considering the lack of local structural information within each patch, the CGFF module employs a dual-guidance mechanism based on center and global structures to effectively promote the aggregation of local geometric features. Extensive experiments on multiple benchmark datasets validate the effectiveness of PointMC, consistently outperforming existing state-of-the-art methods, and demonstrating superior capability in capturing both inter-patch semantic consistency and intra-patch geometric details. Xinxing Yu, Ajian Liu 0001, Sunyuan Qiang, Hui Ma 0018, Yanyan Liang 0001 |
AAAI | 6 |
| 2026 | Distilling Future Temporal Knowledge with Masked Feature Reconstruction for 3D Object DetectionabstractCamera-based temporal 3D object detection has shown impressive results in autonomous driving, with offline models improving accuracy by using future frames. Knowledge distillation (KD) can be an appealing framework for transferring rich information from offline models to online models. However, existing KD methods overlook future frames, as they mainly focus on spatial feature distillation under strict frame alignment or on temporal relational distillation, thereby making it challenging for online models to effectively learn future knowledge. To this end, we propose a sparse query-based approach, Future Temporal Knowledge Distillation (FTKD), which effectively transfers future frame knowledge from an offline teacher model to an online student model. Specifically, we present a future-aware feature reconstruction strategy to encourage the student model to capture future features without strict frame alignment. In addition, we further introduce future-guided logit distillation to leverage the teacher's stable foreground and background context. FTKD is applied to two high-performing 3D object detection baselines, achieving up to 1.3 mAP and 1.3 NDS gains on the nuScenes dataset, as well as the most accurate velocity estimation, without increasing inference cost. Hu Zhu, Weihao Gu, Yang Yang 0062, Yanyan Liang 0001 |
AAAI | 6 |
| 2026 | ICPE-FAS: Instance and Category Prompts Engineering for Generalizable Face Anti-Spoofing
Ajian Liu 0001, Xun Lin, Hui Ma 0018, Xinxing Yu, Jiabao Guo, Zitong Yu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001 |
Int. J. Comput. Vis. | 10 |
| 2026 | OV-KFA: Open-vocabulary object detection via key feature alignment
Yunqing Jiang, Sunyuan Qiang, Wuchen Li, Huijia Zhao, Yanyan Liang 0001 |
Neurocomputing | 5 |
| 2026 | Context-aware knowledge distillation for anomaly detection
Ning Li 0035, Xuxin Lin, Ajian Liu 0001, Chaohao Jiang, Zhenwei Zhu, Yanyan Liang 0001 |
Knowl. Based Syst. | 6 |
| 2026 | SSEditor: Controllable mask-to-scene generation with diffusion model
Jiahao Pang, Zhiqiang Pu, Yanyan Liang 0001 |
Knowl. Based Syst. | 4 |
| 2026 | DGPDL: Domain-Guided Prompt Distribution Learning for Generalizable Face Anti-SpoofingabstractThe overfitting of domain signals results in poor domain generalization of face anti-spoofing. The current methods usually improve the diversity of source domains to alleviate this overfitting. However, this benefit is minimal, as even the most diverse domain signals will also be absent in the target domain. In this work, we propose a Domain-Guided Prompt Distribution Learning (DGPDL) built on Vision-Language Models like CLIP, which explores a unified representation of domain signals as a prompt across the source and target domain to alleviate the understanding bias caused by domain gaps. Specifically, we first define a learnable Domain-Specific Distribution (DSD) that covers as many domain elements as possible, such as image quality, color tone, camera settings, etc., which establish connections between different domains and linearly combinable prompt in any domain; Then, based on the style statistics of the given sample, we construct its optimal Domain-Specific Prompts (DSPs) from the defined DSD through the designed Prompt Assemble Attention (PAA) with the similarity matching; Finally, the assembled DSPs will act as carrier or agent to perform on both the vision and language branches, synergistically improving the model's recognition of domain signals. By using the prompt to represent domain signals uniformly, if the model can be robust to DSPs in the source domain, it should be applicable to target domain, as they share the same DSD. By representing domain signals as prompts rather than instantiation features, DGPDL effectively reduces the reliance on specific domain appearances. This design enables the model to dynamically adapt to unseen target domains without the need for retraining. Extensive experiments show that the DGPDL is effective and outperforms the state-of-the-art methods on several cross-domain benchmarks. Ajian Liu 0001, Xun Lin, Ruicong Zhi, Yanyan Liang 0001, Xinshan Zhu, Zhanchuan Cai, Jun Wan 0001, Sergio Escalera, Zhen Lei 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | Flexible Modal Mixture-of-Experts With Inter-Modal Knowledge Distillation for Face Anti-Spoofing
Hui Ma 0018, Ajian Liu 0001, Ning Li 0035, Boyun Wang, Hang Zou 0002, Yuan Zhang 0023, Jing Huang 0017, Zhiqiang Pu, Jun Wan 0001, Zhanchuan Cai, Zhen Lei 0001, Yanyan Liang 0001 |
IEEE Trans. Inf. Forensics Secur. | 13 |
| 2025 | Not All Frame Features are Equal: Video-to-4D Generation via Decoupling Dynamic-Static FeaturesabstractRecently, the generation of dynamic 3D objects from a video has shown impressive results. Existing methods directly optimize Gaussians using whole information in frames. However, when dynamic regions are interwoven with static regions within frames, particularly if the static regions account for a large proportion, existing methods often overlook information in dynamic regions and are prone to overfitting on static regions. This leads to producing results with blurry textures. We consider that decoupling dynamic-static features to enhance dynamic representations can alleviate this issue. Thus, we propose a dynamic-static feature decoupling module (DSFD). Along temporal axes, it regards the regions of current frame features that possess significant differences relative to reference frame features as dynamic features. Conversely, the remaining parts are the static features. Then, we acquire decoupled features driven by dynamic features and current frame features. Moreover, to further enhance the dynamic representation of decoupled features from different viewpoints and ensure accurate motion prediction, we design a temporal-spatial similarity fusion module (TSSF). Along spatial axes, it adaptively selects similar information of dynamic regions. Hinging on the above, we construct a novel approach, DS4D. Experimental results verify our method achieves state-of-the-art (SOTA) results in video-to-4D. In addition, the experiments on a real-world scenario dataset demonstrate its effectiveness on the 4D scene. Our code will be publicly available. Zhenwei Zhu, Ajian Liu 0001, Hui Ma 0018, Jian Nong, Yanyan Liang 0001 |
ICCV | 7 |
| 2025 | Stochastic Trajectory Prediction Under Unstructured ConstraintsabstractTrajectory prediction facilitates effective planning and decision-making, while constrained trajectory prediction integrates regulation into prediction. Recent advances in constrained trajectory prediction focus on structured constraints by constructing optimization objectives. However, handling unstructured constraints is challenging due to the lack of differentiable formal definitions. To address this, we propose a novel method for constrained trajectory prediction using a conditional generative paradigm, named Controllable Trajectory Diffusion (CTD). The key idea is that any trajectory corresponds to a degree of conformity to a constraint. By quantifying this degree and treating it as a condition, a model can implicitly learn to predict trajectories under unstructured constraints. CTD employs a pre-trained scoring model to predict the degree of conformity (i.e., a score), and uses this score as a condition for a conditional diffusion model to generate trajectories. Experimental results demonstrate that CTD achieves high accuracy on the ETH/UCY and SDD benchmarks. Qualitative analysis confirms that CTD ensures adherence to unstructured constraints and can predict trajectories that satisfy combinatorial constraints. Zhiqiang Pu, Shijie Wang 0006, Boyin Liu, Huimu Wang, Yanyan Liang 0001, Jianqiang Yi |
ICRA | 6 |
| 2025 | PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles
Tianshun Han, Benjia Zhou, Ajian Liu 0001, Yanyan Liang 0001, Zhen Lei 0001, Jun Wan 0001 |
ACM Multimedia | 4 |
| 2025 | FACNet: Feature Alignment Fast Point Cloud Completion NetworkabstractPoint cloud completion aims to infer complete point clouds based on partial 3D point cloud inputs. Various previous methods apply coarse-to-fine strategy networks for generating complete point clouds. However, such methods are not only relatively time-consuming but also cannot provide representative complete shape features based on partial inputs. In this paper, a novel feature alignment fast point cloud completion network (FACNet) is proposed to directly and efficiently generate the detailed shapes of objects. FACNet aligns high-dimensional feature distributions of both partial and complete point clouds to maintain global information about the complete shape. During its decoding process, the local features from the partial point cloud are incorporated along with the maintained global information to ensure complete and time-saving generation of the complete point cloud. Experimental results show that FACNet outperforms the state-of-the-art on PCN, Completion3D, and MVP datasets, and achieves competitive performance on ShapeNet-55 and KITTI datasets. Moreover, FACNet and a simplified version, FACNet-slight, achieve a significant speedup of 3–10 times over other state-of-the-art methods. Xinxing Yu, Chi-Chong Wong, Chi-Man Vong, Yanyan Liang 0001 |
Comput. Vis. Media | 5 |
| 2025 | Open-vocabulary object detection via Neighboring Region Attention Alignment
Sunyuan Qiang, Xianfei Li, Yanyan Liang 0001, Wenlong Liao |
Eng. Appl. Artif. Intell. | 3 |
| 2025 | Robust tracking via rethinking prediction head
Jian Nong, Yongjun Qi, Zhiyi Mo, Yanyan Liang 0001 |
Image Vis. Comput. | 5 |
| 2025 | LLM-DiffAug: Enhancing few-shot object detection via LLM-Guided diffusion augmentation
Yunqing Jiang, Sunyuan Qiang, Wuchen Li, Yanyan Liang 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Decisive vector guided column annotation
Xiaobo Wang 0001, Yanyan Liang 0001, Zhen Lei 0001 |
Pattern Recognit. | 3 |
| 2025 | Region-aware mutual relational knowledge distillation for semantic segmentation
Xuxin Lin, Hailun Liang, Benjia Zhou, Yanyan Liang 0001 |
Pattern Recognit. | 5 |
| 2025 | Collaborative Adapter Experts for Class-Incremental LearningabstractPre-trained models (PTMs) with parameter-efficient fine-tuning (PEFT) techniques have been extensively utilized in class-incremental learning (CIL) scenarios. However, they still remain susceptible to performance degradation as the individual PEFT module operates as an independent learning entity during the incremental process. To this end, this work proposes a novel class-incremental collaborative adapter experts (CICAE) model, which incorporates multiple adapters operating collaboratively to facilitate CIL. Specifically, our model primarily consists of two phases. Initially, multiple adapters are employed to establish a multi-expert system aimed at acquiring diverse incremental knowledge. Through the collaborative knowledge sharing (CKS) mechanism, the expertise of each adapter expert is transferable, promoting collaborative development and mutual advancement. Subsequently, with the category prototype distributions, collaborative classifier alignment (CCA) is proposed to further align the classifiers with the representation space in a cooperative manner. Extensive experiments on CIL benchmarks validate the superior performance of our model. Sunyuan Qiang, Xinxing Yu, Yanyan Liang 0001, Jun Wan 0001 |
IEEE Signal Process. Lett. | 3 |
| 2025 | C2RL: Content and Context Representation Learning for Gloss-Free Sign Language Translation and RetrievalabstractSign Language Representation Learning (SLRL) is crucial for a range of sign language-related downstream tasks such as Sign Language Translation (SLT) and Sign Language Retrieval (SLRet). Recently, many gloss-based and gloss-free SLRL methods have been proposed, showing promising performance. Among them, the gloss-free approach shows promise for strong scalability without relying on gloss annotations. However, it currently faces suboptimal solutions due to challenges in encoding the intricate, context-sensitive characteristics of sign language videos, mainly struggling to discern essential sign features using a non-monotonic video-text alignment strategy. Therefore, we introduce an innovative pretraining paradigm for gloss-free SLRL, called C2RL, in this paper. Specifically, rather than merely incorporating a non-monotonic semantic alignment of video and text to learn language-oriented sign features, we emphasize two pivotal aspects of SLRL: Implicit Content Learning (ICL) and Explicit Context Learning (ECL). ICL delves into the content of communication, capturing the nuances, emphasis, timing, and rhythm of the signs. In contrast, ECL focuses on understanding the contextual meaning of signs and converting them into equivalent sentences. Despite its simplicity, extensive experiments confirm that the joint optimization of ICL and ECL results in robust sign language representation and significant performance gains in gloss-free SLT and SLRet tasks. Notably, C2RL improves the BLEU-4 score by +5.3 on P14T, +10.6 on CSL-daily, +6.2 on OpenASL, and +1.3 on How2Sign. It also boosts the R@1 score by +8.3 on P14T, +14.4 on CSL-daily, and +5.9 on How2Sign. Additionally, we set a new baseline for the OpenASL dataset in the SLRet task. Benjia Zhou, Jun Wan 0001, Yibo Hu 0001, Hailin Shi, Yanyan Liang 0001, Zhen Lei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | FA3-CLIP: Frequency-Aware Cues Fusion and Attack-Agnostic Prompt Learning for Unified Face Attack DetectionabstractFacial recognition systems are vulnerable to physical (e.g., printed photos) and digital (e.g., DeepFake) face attacks. Existing methods struggle to simultaneously detect physical and digital attacks due to: 1) significant intra-class variations between these attack types, and 2) the inadequacy of spatial information alone to comprehensively capture live and fake cues. To address these issues, we propose a unified attack detection model termed Frequency-Aware and Attack-Agnostic CLIP (FA3-CLIP), which introduces attack-agnostic prompt learning to express generic live and fake cues derived from the fusion of spatial and frequency features, enabling unified detection of live faces and all categories of attacks. Specifically, the attack-agnostic prompt module generates generic live and fake prompts within the language branch to extract corresponding generic representations from both live and fake faces, guiding the model to learn a unified feature space for unified attack detection. Meanwhile, the module adaptively generates the live/fake conditional bias from the original spatial and frequency information to optimize the generic prompts accordingly, reducing the impact of intra-class variations. We further propose a dual-stream cues fusion framework in the vision branch, which leverages frequency information to complement subtle cues that are difficult to capture in the spatial domain. In addition, a frequency compression block is utilized in the frequency stream, which reduces redundancy in frequency features while preserving the diversity of crucial cues. We also establish new challenging protocols to facilitate unified face attack detection effectiveness. Experimental results on multiple benchmarks demonstrate that FA3-CLIP significantly improves performance, reducing ACER by over 1.2% on UniAttackData, and increasing AUC by more than 3% as well as reducing EER by over 4% on the JFSFDB dataset. Yongze Li, Ning Li 0035, Ajian Liu 0001, Hui Ma 0018, Xihong Chen, Zhiyao Liang, Yanyan Liang 0001, Jun Wan 0001, Zhen Lei 0001 |
IEEE Trans. Inf. Forensics Secur. | 8 |
| 2025 | Knowledge Distillation-Based Anomaly Detection via Adaptive Discrepancy OptimizationabstractKnowledge distillation has emerged as a primary solution for anomaly detection, leveraging feature discrepancies between teacher–student (T–S) networks to locate anomalies. However, previous approaches suffer from ambiguous feature discrepancies, which hinder effective anomaly detection due to two main challenges: 1) overgeneralization, where the student network excessively mimics teacher features in anomalous regions, and 2) semantic bias between T–S networks in normal regions. To address these issues, we propose an Adaptive Discrepancy Optimization (Ado) block. The Ado block adaptively calibrates feature discrepancies by reducing overgeneralization in anomalous regions and selectively aligning semantic features in normal regions via learnable feature offsets. This versatile block can be seamlessly integrated into various distillation-based methods. Experimental results demonstrate that the Ado block significantly enhances performance across 11 different knowledge distillation frameworks on two widely used datasets. Notably, when integrated with the Ado block, RD4AD achieves a 22% relative improvement in pixel-level PRO on the VisA dataset. In addition, a real-world keyboard inspection application further validates the effectiveness of the Ado block. Ning Li 0035, Ajian Liu 0001, Zhenwei Zhu, Xuxin Lin, Hui Ma 0018, Hongning Dai, Yanyan Liang 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2025 | PMMTalk$:$ Speech-Driven 3D Facial Animation From Complementary Pseudo Multi-Modal FeaturesabstractSpeech-driven 3D facial animation has improved a lot recently while most related works only utilize acoustic modality and neglect the influence of visual and textual cues, leading to unsatisfactory results in terms of precision and coherence. We argue that visual and textual cues are not trivial information. Therefore, we present a novel framework, namely PMMTalk, using complementaryPseudoMulti-Modal features for improving the accuracy of facial animation. The framework entails three modules: PMMTalk encoder, cross-modal alignment module, and PMMTalk decoder. Specifically, the PMMTalk encoder employs the off-the-shelf talking head generation architecture and speech recognition technology to extract visual and textual information from speech, respectively. Following this, the cross-modal alignment module aligns the audio-image-text features at temporal and semantic levels. Subsequently, the PMMTalk decoder is employed to predict lip-syncing facial blendshape coefficients. Contrary to prior methods, PMMTalk only requires an additional random reference face image but yields more accurate results. Additionally, it is artist-friendly as it seamlessly integrates into standard animation production workflows by introducing facial blendshape coefficients. Finally, given the scarcity of 3D talking face datasets, we introduce a large-scale3DChineseAudio-VisualFacialAnimation (3D-CAVFA) dataset. Extensive experiments and user studies show that our approach outperforms the state of the art. Codes and datasets are available at PMMTalk. Tianshun Han, Shengnan Gui, Baihui Li, Lijian Liu, Benjia Zhou, Ruicong Zhi, Yanyan Liang 0001, Jun Wan 0001 |
IEEE Trans. Multim. | 10 |
| 2024 | CFPL-FAS: Class Free Prompt Learning for Generalizable Face Anti-SpoofingabstractDomain generalization (DG) based Face Anti-Spoofing (FAS) aims to improve the model's performance on unseen domains. Existing methods either rely on domain labels to align domain-invariant feature spaces, or disentangle generalizable features from the whole sample, which inevitably lead to the distortion of semantic feature structures and achieve limited generalization. In this work, we make use of large-scale VLMs like CLIP and leverage the textual feature to dynamically adjust the classifier's weights for exploring generalizable visual features. Specifically, we propose a novel Class Free Prompt Learning (CFPL) paradigm for DG FAS, which utilizes two lightweight transformers, namely Content Q-Former (CQF) and Style Q-Former (SQF), to learn the different semantic prompts conditioned on content and style features by using a set of learnable query vectors, respectively. Thus, the generalizable prompt can be learned by two improvements: (1) A Prompt-Text Matched (PTM) supervision is introduced to ensure CQF learns visual representation that is most informative of the content description. (2) A Diversified Style Prompt (DSP) technology is proposed to diversify the learning of style prompts by mixing feature statistics between instance-specific styles. Finally, the learned text features modulate visual features to generalization through the designed Prompt Modulation (PM). Extensive experiments show that the CFPL is effective and outperforms the state-of-the-art methods on several cross-domain datasets. Ajian Liu 0001, Jianwen Gan, Jun Wan 0001, Yanyan Liang 0001, Jiankang Deng, Sergio Escalera, Zhen Lei 0001 |
CVPR | 5 |
| 2024 | CalibRBEV: Multi-Camera Calibration via Reversed Bird's-eye-view Representations for Autonomous DrivingabstractCamera calibration is crucial in computer vision tasks and applications, e.g., autonomous driving (AD). However, prevailing camera calibration models pose a time-consuming and labor-intensive off-board process in mass production settings, while simultaneously lacking exploration of real-world AD scenarios. To this end, inspired by recent advancements in bird's-eye-view (BEV) perception models, this paper proposes a novel multi-camera Calibration method via Reversed BEV representations for AD, termed CalibRBEV. Specifically, the proposed CalibRBEV model primarily comprises two stages. Initially, we innovatively reverse the BEV perception pipeline, reconstructing bounding boxes through an attention auto-encoder module to fully extract the latent reversed BEV representations. Subsequently, the obtained representations from encoder are interacted with the surrounding multi-view image features for further refinement and calibration parameters prediction. Extensive experimental results on nuScenes and Waymo datasets validate the effectiveness of our proposed model. Wenlong Liao, Sunyuan Qiang, Xianfei Li, Yanyan Liang 0001, Junchi Yan |
ACM Multimedia | 6 |
| 2024 | FM-CLIP: Flexible Modal CLIP for Face Anti-SpoofingabstractIn this work, borrowing a solution from the large-scale vision-language models (VLMs) instead of directly removing modality-specific signals from visual features, we propose a novel Flexible Modal CLIP (FM-CLIP) for flexible modal FAS, that can utilize text features to dynamically adjust visual features to be modality independent. In the visual branch, considering the huge visual differences of the same attack in different modalities, which makes it difficult for classifiers to flexibly identify subtle spoofing clues in different test modalities, we propose Cross-Modal Spoofing Enhancer (CMS-Enhancer). It includes a Frequency Extractor (FE) and Cross-Modal Interactor (CMI), aiming to map different modal attacks in a shared frequency space to reduce interference from modality-specific signals and enhance spoofing clues by leveraging cross-modal learning from the shared frequency space. In the text branch, we introduce a Language-Guided Patch Alignment (LGPA) based on prompt learning, which further guides the image encoder to focus on patch-level spoofing representations through dynamic weighting by text features. Thus, our FM-CLIP can flexibly test different modal samples by identifying and enhancing modality-agnostic spoofing cues. Finally, extensive experiments show that FM-CLIP is effective and outperforms state-of-the-art methods on multiple multi-modal datasets. Ajian Liu 0001, Hui Ma 0018, Junze Zheng, Haocheng Yuan, Xiaoyuan Yu, Yanyan Liang 0001, Sergio Escalera, Jun Wan 0001, Zhen Lei 0001 |
ACM Multimedia | 6 |
| 2024 | Coevolving with the Other You: Fine-Tuning LLM with Sequential Cooperative Multi-Agent Reinforcement LearningabstractReinforcement learning (RL) has emerged as a pivotal technique for fine-tuning large language models (LLMs) on specific tasks. However, prevailing RL fine-tuning methods predominantly rely on PPO and its variants. Though these algorithms are effective in general RL settings, they often exhibit suboptimal performance and vulnerability to distribution collapse when applied to the fine-tuning of LLMs. In this paper, we propose CORY, extending the RL fine-tuning of LLMs to a sequential cooperative multi-agent reinforcement learning framework, to leverage the inherent coevolution and emergent capabilities of multi-agent systems. In CORY, the LLM to be fine-tuned is initially duplicated into two autonomous agents: a pioneer and an observer. The pioneer generates responses based on queries, while the observer generates responses using both the queries and the pioneer’s responses. The two agents are trained together. During training, the agents exchange roles periodically, fostering cooperation and coevolution between them. Experiments evaluate CORY's performance by fine-tuning GPT-2 and Llama-2 under subjective and objective reward functions on the IMDB Review and GSM8K datasets, respectively. Results show that CORY outperforms PPO in terms of policy optimality, resistance to distribution collapse, and training robustness, thereby underscoring its potential as a superior methodology for refining LLMs in real-world applications. Zhiqiang Pu, Boyin Liu, Xiaolin Ai, Yanyan Liang 0001, Min Chen 0038 |
NeurIPS | 6 |
| 2024 | Adapt and Refine: A Few-Shot Class-Incremental Learner via Pre-Trained Models
Sunyuan Qiang, Zhu Xiong, Yanyan Liang 0001, Jun Wan 0001 |
PRCV (1) | 3 |
| 2024 | Agile Optimization Framework: A framework for tensor operator optimization in neural networkabstractIn recent years, with the gradual slowing of Moore’s Law and the development of deep learning , the demand for hardware performance of executing deep learning based applications has significantly increased. In this case, deep learning compilers have been proven to maximize hardware performance while keeping computational power constant, especially the end-to-end compiler Tensor Virtual Machine (TVM). TVM optimizes tensors by finding excellent parallel computing schemes, thereby achieving the goal of improving the performance of neural network inference. However, there is still untapped potential in current optimization methods. However, existing optimization methods based on the TVM, such as Genetic Algorithms Tuner (GA-Tuner), have failed to achieve a balance between optimization performance and optimization time. The intolerable duration of optimization detracts from TVM’s usability, rendering it challenging to extend into the scientific community. This paper introduces a novel deep learning compilation optimization framework base on TVM called Agile Optimization Framework (AOF), which incorporates a tuner based on the latest Beluga Whale Optimization Algorithm (BWO). The BWO is adept at tackling complex problems characterized by numerous local optima, making it particularly suitable for hardware compilation optimization scenarios. We further propose an Evolving Epsilon Strategy (EES), a search strategy that adaptively adjusts the balance between exploration and exploitation, thereby enhancing the effectiveness of the algorithm. Additionally, we developed a supervised Tuning Accelerator (TA) aimed at reducing the time required for optimization and enhancing efficiency. Comparative experiments demonstrate that AOF achieves 11.36%–66.20% improvement in performance and 30.30%–54.60% reduction in optimization time, significantly outperforming the control group. Mingwei Zhou, Xuxin Lin, Yanyan Liang 0001 |
Future Gener. Comput. Syst. | 3 |
| 2024 | CSDG-FAS: Closed-Space Domain Generalization for Face Anti-spoofing
Keyao Wang, Haixiao Yue, Yanyan Liang 0001, Mouxiao Huang, Junyu Han, Errui Ding, Jingdong Wang 0001 |
Int. J. Comput. Vis. | 4 |
| 2023 | Mixture Uniform Distribution Modeling and Asymmetric Mix Distillation for Class Incremental LearningabstractExemplar rehearsal-based methods with knowledge distillation (KD) have been widely used in class incremental learning (CIL) scenarios. However, they still suffer from performance degradation because of severely distribution discrepancy between training and test set caused by the limited storage memory on previous classes. In this paper, we mathematically model the data distribution and the discrepancy at the incremental stages with mixture uniform distribution (MUD). Then, we propose the asymmetric mix distillation method to uniformly minimize the error of each class from distribution discrepancy perspective. Specifically, we firstly promote mixup in CIL scenarios with the incremental mix samplers and incremental mix factor to calibrate the raw training data distribution. Next, mix distillation label augmentation is incorporated into the data distribution to inherit the knowledge information from the previous models. Based on the above augmented data distribution, our trained model effectively alleviates the performance degradation and extensive experimental results validate that our method exhibits superior performance on CIL benchmarks. Sunyuan Qiang, Jiayi Hou, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001 |
AAAI | 4 |
| 2023 | Long-Range Grouping Transformer for Multi-View 3D ReconstructionabstractNowadays, transformer networks have demonstrated superior performance in many computer vision tasks. In a multi-view 3D reconstruction algorithm following this paradigm, self-attention processing has to deal with intricate image tokens including massive information when facing heavy amounts of view input. The curse of information content leads to the extreme difficulty of model learning. To alleviate this problem, recent methods compress the token number representing each view or discard the attention operations between the tokens from different views. Obviously, they give a negative impact on performance. Therefore, we propose long-range grouping attention (LGA) based on the divide-and-conquer principle. Tokens from all views are grouped for separate attention operations. The tokens in each group are sampled from all views and can provide macro representation for the resided view. The richness of feature learning is guaranteed by the diversity among different groups. An effective and efficient encoder can be established which connects inter-view features using LGA and extract intra-view features using the standard self-attention layer. Moreover, a novel progressive upsampling decoder is also designed for voxel generation with relatively high resolution. Hinging on the above, we construct a powerful transformer-based network, called LRGT. Experimental results on ShapeNet verify our method achieves SOTA accuracy in multi-view reconstruction. Code is available at https://github.com/LiyingCV/Long-Range-Grouping-Transformer. Zhenwei Zhu, Xuxin Lin, Jian Nong, Yanyan Liang 0001 |
ICCV | 5 |
| 2023 | Gloss-free Sign Language Translation: Improving from Visual-Language PretrainingabstractSign Language Translation (SLT) is a challenging task due to its cross-domain nature, involving the translation of visual-gestural language to text. Many previous methods employ an intermediate representation, i.e., gloss sequences, to facilitate SLT, thus transforming it into a two-stage task of sign language recognition (SLR) followed by sign language translation (SLT). However, the scarcity of gloss-annotated sign language data, combined with the information bottleneck in the mid-level gloss representation, has hindered the further development of the SLT task. To address this challenge, we propose a novel Gloss-Free SLT based on Visual-Language Pretraining (GFSLT-VLP), which improves SLT by inheriting language-oriented prior knowledge from pre-trained models, without any gloss annotation assistance. Our approach involves two stages: (i) integrating Contrastive Language-Image Pre-training (CLIP) with masked self-supervised learning to create pre-tasks that bridge the semantic gap between visual and textual representations and restore masked sentences, and (ii) constructing an end-to-end architecture with an encoder-decoder-like structure that inherits the parameters of the pre-trained Visual Encoder and Text Decoder from the first stage. The seamless combination of these novel designs forms a robust sign language representation and significantly improves gloss-free sign language translation. In particular, we have achieved unprecedented improvements in terms of BLEU-4 score on the PHOENIX14T dataset (≥+5) and the CSL-Daily dataset (≥+3) compared to state-of-the-art gloss-free SLT methods. Furthermore, our approach also achieves competitive results on the PHOENIX14T dataset when compared with most of the gloss-based methods1. Benjia Zhou, Albert Clapés, Jun Wan 0001, Yanyan Liang 0001, Sergio Escalera, Zhen Lei 0001 |
ICCV | 5 |
| 2023 | UMIFormer: Mining the Correlations between Similar Tokens for Multi-View 3D ReconstructionabstractIn recent years, many video tasks have achieved breakthroughs by utilizing the vision transformer and establishing spatial-temporal decoupling for feature extraction. Although multi-view 3D reconstruction also faces multiple images as input, it cannot immediately inherit their success due to completely ambiguous associations between unstructured views. There is not usable prior relationship, which is similar to the temporally-coherence property in a video. To solve this problem, we propose a novel transformer network for Unstructured Multiple Images (UMIFormer). It exploits transformer blocks for decoupled intra-view encoding and designed blocks for token rectification that mine the correlation between similar tokens from different views to achieve decoupled interview encoding. Afterward, all tokens acquired from various branches are compressed into a fixed-size compact representation while preserving rich information for reconstruction by leveraging the similarities between tokens. We empirically demonstrate on ShapeNet and confirm that our decoupled learning method is adaptable for unstructured multiple images. Meanwhile, the experiments also verify our model outperforms existing SOTA methods by a large margin. Code will be available at https://github.com/GaryZhu1996/UMIFormer. Zhenwei Zhu, Ning Li 0035, Chaohao Jiang, Yanyan Liang 0001 |
ICCV | 5 |
| 2023 | Lazy Agents: A New Perspective on Solving Sparse Reward Problem in Multi-agent Reinforcement LearningabstractSparse reward remains a valuable and challenging problem in multi-agent reinforcement learning (MARL). This paper addresses this issue from a new perspective, i.e., lazy agents. We empirically illustrate how lazy agents damage learning from both exploration and exploitation. Then, we propose a novel MARL framework called Lazy Agents Avoidance through Influencing External States (LAIES). Firstly, we examine the causes and types of lazy agents in MARL using a causal graph of the interaction between agents and their environment. Then, we mathematically define the concept of fully lazy agents and teams by calculating the causal effect of their actions on external states using the do-calculus process. Based on definitions, we provide two intrinsic rewards to motivate agents, i.e., individual diligence intrinsic motivation (IDI) and collaborative diligence intrinsic motivation (CDI). IDI and CDI employ counterfactual reasoning based on the external states transition model (ESTM) we developed. Empirical results demonstrate that our proposed method achieves state-of-the-art performance on various tasks, including the sparse-reward version of StarCraft multi-agent challenge (SMAC) and Google Research Football (GRF). Our code is open-source and available at https://github.com/liuboyin/LAIES. Boyin Liu, Zhiqiang Pu, Yi Pan 0009, Jianqiang Yi, Yanyan Liang 0001 |
ICML | 5 |
| 2023 | Heterogeneous-graph Attention Reinforcement Learning for Football MatchesabstractFootball player's decision-making problem is quite challenging because of the essential feature of football matches: many players with different and complex cooperative or competitive relationships. To better leverage these relationships, this paper proposes a player-policy learning method with heterogeneous-graph attention reinforcement learning (PPL-HGARL) to enable an active player closest to the ball to learn effective policies for playing football matches. Specifically, a multi-head feature representation module is designed to reconstruct raw observations using prior expert knowledge. Furthermore, a heterogeneous-graph player-relation attention network is designed by use of graph among different roles, in order to model cooperative and competitive relations among players. The network learns proper state representation for the active player, making the player pay attention to other important players. Besides, an actor-critic algorithm is adopted to train the policies efficiently. Competing with rule-based opponents of different difficulty levels in the Google Research Football environment, the active player achieves excellent results, which validates the effectiveness and superiority of the proposed method. Shijie Wang 0006, Yi Pan 0009, Zhiqiang Pu, Jianqiang Yi, Yanyan Liang 0001 |
IJCNN | 5 |
| 2023 | A Unified Multimodal De- and Re-Coupling Framework for RGB-D Motion RecognitionabstractMotion recognition is a promising direction in computer vision, but the training of video classification models is much harder than images due to insufficient data and considerable parameters. To get around this, some works strive to explore multimodal cues from RGB-D data. Although improving motion recognition to some extent, these methods still face sub-optimal situations in the following aspects: (i) Data augmentation, i.e., the scale of the RGB-D datasets is still limited, and few efforts have been made to explore novel data augmentation strategies for videos; (ii) Optimization mechanism, i.e., the tightly space-time-entangled network structure brings more challenges to spatiotemporal information modeling; And (iii) cross-modal knowledge fusion, i.e., the high similarity between multimodal representations leads to insufficient late fusion. To alleviate these drawbacks, we propose to improve RGB-D-based motion recognition both from data and algorithm perspectives in this article. In more detail, firstly, we introduce a novel video data augmentation method dubbed ShuffleMix, which acts as a supplement to MixUp, to provide additional temporal regularization for motion recognition. Secondly, a Unified Multimodal De-coupling and multi-stage Re-coupling framework, termed UMDR, is proposed for video representation learning. Finally, a novel cross-modal Complement Feature Catcher (CFCer) is explored to mine potential commonalities features in multimodal information as the auxiliary fusion stream, to improve the late fusion results. The seamless combination of these novel designs forms a robust spatiotemporal representation and achieves better performance than state-of-the-art methods on four public motion datasets. Specifically, UMDR achieves unprecedented improvements of ↑ 4.5% on the Chalearn IsoGD dataset. Benjia Zhou, Pichao Wang, Jun Wan 0001, Yanyan Liang 0001, Fan Wang 0019 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | GARNet: Global-aware multi-view 3D reconstruction network and the cost-performance tradeoff
Zhenwei Zhu, Xuxin Lin, Yanyan Liang 0001 |
Pattern Recognit. | 5 |
| 2023 | FM-ViT: Flexible Modal Vision Transformers for Face Anti-SpoofingabstractThe availability of handy multi-modal (i.e., RGB-D) sensors has brought about a surge of face anti-spoofing research. However, the current multi-modal face presentation attack detection (PAD) has two defects: (1) The framework based on multi-modal fusion requires providing modalities consistent with the training input, which seriously limits the deployment scenario. (2) The performance of ConvNet-based model on high fidelity datasets is increasingly limited. In this work, we present a pure transformer-based framework, dubbed the Flexible Modal Vision Transformer (FM-ViT), for face anti-spoofing to flexibly target any single-modal (i.e., RGB) attack scenarios with the help of available multi-modal data. Specifically, FM-ViT retains a specific branch for each modality to capture different modal information and introduces the Cross-Modal Transformer Block (CMTB), which consists of two cascaded attentions named Multi-headed Mutual-Attention (MMA) and Fusion-Attention (MFA) to guide each modal branch to mine potential features from informative patch tokens, and to learn modality-agnostic liveness features by enriching the modal information of own CLS token, respectively. Experiments demonstrate that the single model trained based on FM-ViT can not only flexibly evaluate different modal samples, but also outperforms existing single-modal frameworks by a large margin, and approaches the multi-modal frameworks introduced with smaller FLOPs and model parameters. Ajian Liu 0001, Zichang Tan, Zitong Yu, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001, Stan Z. Li, Guodong Guo |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2022 | Decoupling and Recoupling Spatiotemporal Representation for RGB-D-based Motion RecognitionabstractDecoupling spatiotemporal representation refers to decomposing the spatial and temporal features into dimension-independent factors. Although previous RGB-D-based motion recognition methods have achieved promising performance through the tightly coupled multi-modal spatiotemporal representation, they still suffer from (i) optimization difficulty under small data setting due to the tightly spatiotemporal-entangled modeling; (ii) information redundancy as it usually contains lots of marginal information that is weakly relevant to classification; and (iii) low interaction between multi-modal spatiotemporal information caused by insufficient late fusion. To alleviate these drawbacks, we propose to decouple and recouple spatiotemporal representation for RGB-D-based motion recognition. Specifically, we disentangle the task of learning spatiotemporal representation into 3 sub-tasks: (1) Learning high-quality and dimension independent features through a decoupled spatial and temporal modeling network. (2) Recoupling the decoupled representation to establish stronger space-time dependency. (3) Introducing a Cross-modal Adaptive Posterior Fusion (CAPF) mechanism to capture cross-modal spatiotemporal information from RGB-D data. Seamless combination of these novel designs forms a robust spatiotemporal representation and achieves better performance than state-of-the-art methods on four public motion datasets. Our code is available at https://github.com/damo-cv/MotionRGBD. Benjia Zhou, Pichao Wang, Jun Wan 0001, Yanyan Liang 0001, Fan Wang 0019, Zhen Lei 0001, Hao Li 0030, Rong Jin 0001 |
CVPR | 4 |
| 2022 | Disentangling Facial Pose and Appearance Information for Face Anti-spoofingabstractFace Anti-spoofing aims to determine whether the captured face from a face recognition system is real or fake. However, the facial pose and local significant spoofing traces (i.e., the boundary and reflection spot in presentation attack instruments) seriously affects the performance and stability of the current algorithms. Due to they regard the face image as an indivisible unit, and process it holistically, rarely consider excluding these liveness-irrelated factors. Unlike it, we design a Pose-Independent Face Anti-Spoofing (PIFAS) framework to disentangle face into an appearance information and a pose code to capture liveness and liveness-irrelated features, respectively. Specifically, the PIFAS consists of an Unsupervised Pose Switching (UPS) module and a Mutual Information Averaged Defense (MIAD) module, which are used to control the facial pose and suppress the local significant attack traces by averaging the local and global knowledge. Extensive experimental evaluations on multiple face anti-spoofing datasets verify that the proposed method can improve the generalization and stabilize the performance of each testing video through alleviating the interference from liveness-irrelated factors. Ajian Liu 0001, Jun Wan 0001, Yanyan Liang 0001 |
ICPR | 5 |
| 2022 | MA-ViT: Modality-Agnostic Vision Transformers for Face Anti-SpoofingabstractThe existing multi-modal face anti-spoofing (FAS) frameworks are designed based on two strategies: halfway and late fusion. However, the former requires test modalities consistent with the training input, which seriously limits its deployment scenarios. And the latter is built on multiple branches to process different modalities independently, which limits their use in applications with low memory or fast execution requirements. In this work, we present a single branch based Transformer framework, namely Modality-Agnostic Vision Transformer (MA-ViT), which aims to improve the performance of arbitrary modal attacks with the help of multi-modal data. Specifically, MA-ViT adopts the early fusion to aggregate all the available training modalities’ data and enables flexible testing of any given modal samples. Further, we develop the Modality-Agnostic Transformer Block (MATB) in MA-ViT, which consists of two stacked attentions named Modal-Disentangle Attention (MDA) and Cross-Modal Attention (CMA), to eliminate modality-related information for each modal sequences and supplement modality-agnostic liveness features from another modal sequences, respectively. Experiments demonstrate that the single model trained based on MA-ViT can not only flexibly evaluate different modal samples, but also outperforms existing single-modal frameworks by a large margin, and approaches the multi-modal frameworks introduced with smaller FLOPs and model parameters. Ajian Liu 0001, Yanyan Liang 0001 |
IJCAI | 2 |
| 2022 | Contrastive Context-Aware Learning for 3D High-Fidelity Mask Face Presentation Attack DetectionabstractFace presentation attack detection (PAD) is essential to secure face recognition systems primarily from high-fidelity mask attacks. Most existing 3D mask PAD benchmarks suffer from several drawbacks: 1) a limited number of mask identities, types of sensors, and a total number of videos; 2) low-fidelity quality of facial masks. Basic deep models and remote photoplethysmography (rPPG) methods achieved acceptable performance on these benchmarks but still far from the needs of practical scenarios. To bridge the gap to real-world applications, we introduce a large-scale High-Fidelity Mask dataset, namely HiFiMask. Specifically, a total amount of 54,600 videos are recorded from 75 subjects with 225 realistic masks by 7 new kinds of sensors. Along with the dataset, we propose a novel Contrastive Context-aware Learning (CCL) framework. CCL is a new training methodology for supervised PAD tasks, which is able to learn by leveraging rich contexts accurately (e.g., subjects, mask material and lighting) among pairs of live faces and high-fidelity mask attacks. Extensive experimental evaluations on HiFiMask and three additional 3D mask datasets demonstrate the effectiveness of our method. The codes and dataset will be released soon. Ajian Liu 0001, Zitong Yu, Jun Wan 0001, Anyang Su, Zichang Tan, Sergio Escalera, Junliang Xing, Yanyan Liang 0001, Guodong Guo, Zhen Lei 0001, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 10 |
| 2022 | RVFace: Reliable Vector Guided Softmax Loss for Face RecognitionabstractFace recognition has witnessed significant progress with the advances of deep convolutional neural networks (CNNs), and the central task of which is how to improve the feature discrimination. To this end, several margin-based (e.g., angular, additive and additive angular margins) softmax loss functions have been proposed to increase the feature margin between different classes. However, despite great achievements have been made, they mainly suffer from four issues: 1) They are based on the assumption of well-cleaned training sets, without considering the consequence of noisy labels inherently existing in most of face recognition datasets; 2) They ignore the importance of informative (e.g., semi-hard) features mining for discriminative learning; 3) They encourage the feature margin only from the perspective of ground truth class, without realizing the discriminability from other non-ground truth classes; and 4) They set the feature margin between different classes to be same and fixed, which may not adapt the situation of unbalanced data in different classes very well. To cope with these issues, this paper develops a novel loss function, which explicitly estimates the noisy labels to drop them and adaptively emphasizes the semi-hard feature vectors from the remaining reliable ones to guide the discriminative feature learning. Thus we can address all the above issues and achieve more discriminative features for face recognition. To the best of our knowledge, this is the first attempt to inherit the advantages of feature-based noisy labels detection, feature mining and feature margin into a unified loss function. Extensive experimental results on a variety of face recognition benchmarks have demonstrated the effectiveness of our method over state-of-the-art alternatives. Our source code is available at http://www.cbsr.ia.ac.cn/users/xiaobowang/. Xiaobo Wang 0001, Yanyan Liang 0001, Liang Gu, Zhen Lei 0001 |
IEEE Trans. Image Process. | 3 |
| 2022 | Frequency Feature Pyramid Network With Global-Local Consistency Loss for Crowd-and-Vehicle Counting in Congested ScenesabstractContext prediction plays a crucial role in implementing autonomous driving applications. As one of important context-prediction tasks, crowd-and-vehicle counting is critical for achieving real-time traffic and crowd analysis, consequently facilitating decision-making processes for autonomous vehicles. However, the completion of crowd-and-vehicle counting also faces challenges, such as large-scale variations, imbalanced data distribution, and insufficient local patterns. To tackle these challenges, we put forth a novel frequency feature pyramid network (FFPNet) in this paper. Our proposed FFPNet extracts the multi-scale information by frequency feature pyramid module, which can tackle the issue of large-scale variations. Meanwhile, the frequency feature pyramid module uses different frequency branches to obtain different scale information. We also adopt the attention mechanism to strength the extraction of different scale information. Moreover, we devise a novel loss function, namely global-local consistency loss, to address the existing problems of imbalanced data distribution and insufficient local patterns. Furthermore, we conduct extensive experiments on six datasets to evaluate our proposed FFPNet. It is worth mentioning that we also construct a novel crowd-and-vehicle dataset (CROVEH), which is the only dataset that contains both crowd-and-vehicle annotations. The experimental results show that FFPNet achieves the best performance on different backbones, e.g., 52.69 mean absolute error (MAE) on P2PNet with FFP module. The codes are available at:https://github.com/MUST-AI-Lab/FFPNet. Xiaoyuan Yu, Yanyan Liang 0001, Xuxin Lin, Jun Wan 0001, Tian Wang 0001, Hongning Dai |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Face Anti-Spoofing via Adversarial Cross-Modality TranslationabstractFace Presentation Attack Detection (PAD) approaches based on multi-modal data have been attracted increasingly by the research community. However, they require multi-modal face data consistently involved in both the training and testing phases. It would severely limit the applicability due to the most Face Anti-spoofing (FAS) systems are only equipped with Visible (VIS) imaging devices, i.e., RGB cameras. Therefore, how to use other modality (i.e., Near-Infrared (NIR)) to assist the performance improvement of VIS-based PAD is significant for FAS. In this work, we first discuss the big gap of performances among different modalities even though the same backbone network is applied. Then, we propose a novel Cross-modal Auxiliary (CMA) framework for the VIS-based FAS task. The main trait of CMA is that the performance can be greatly improved with the help of other modality while no other modality is required in the testing stage. The proposed CMA consists of a Modality Translation Network (MT-Net) and a Modality Assistance Network (MA-Net). The former aims to close the visible gap between different modalities via a generative model that maps inputs from one modality (i.e., RGB) to another (i.e., NIR). The latter focuses on how to use the translated modality (i.e., target modality) and RGB modality (i.e., source modality) together to train a discriminative PAD model. Extensive experiments are conducted to demonstrate that the proposed framework can push the state-of-the-art (SOTA) performances on both multi-modal datasets (i.e., CASIA-SURF, CeFA, and WMCA) and RGB-based datasets (i.e., OULU-NPU, and SiW). Ajian Liu 0001, Zichang Tan, Jun Wan 0001, Yanyan Liang 0001, Zhen Lei 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2021 | Adaptive cross-fusion learning for multi-modal gesture recognitionabstractGesture recognition has attracted significant attention because of its wide range of potential applications. Although multi-modal gesture recognition has made significant progress in recent years, a popular method still is simply fusing prediction scores at the end of each branch, which often ignores complementary features among different modalities in the early stage and does not fuse the complementary features into a more discriminative feature. This paper proposes an Adaptive Cross-modal Weighting (ACmW) scheme to exploit complementarity features from RGB-D data in this study. The scheme learns relations among different modalities by combining the features of different data streams. The proposed ACmW module contains two key functions: (1) fusing complementary features from multiple streams through an adaptive one-dimensional convolution; and (2) modeling the correlation of multi-stream complementary features in the time dimension. Through the effective combination of these two functional modules, the proposed ACmW can automatically analyze the relationship between the complementary features from different streams, and can fuse them in the spatial and temporal dimensions. Extensive experiments validate the effectiveness of the proposed method, and show that our method outperforms state-of-the-art methods on IsoGD and NVGesture. Benjia Zhou, Jun Wan 0001, Yanyan Liang 0001, Guodong Guo |
Virtual Real. Intell. Hardw. | 3 |
| 2020 | Task-Oriented Feature-Fused Network With Multivariate Dataset for Joint Face AnalysisabstractDeep multitask learning for face analysis has received increasing attentions. From literature, most existing methods focus on optimizing a main task by jointly learning several auxiliary tasks. It is challenging to consider the performance of each task in a multitask framework due to the following reasons: 1) different face tasks usually rely on different levels of semantic features; 2) each task has different learning convergence rate, which could affect the whole performance when joint training; and 3) multitask model needs rich label information for efficient training, but existing facial datasets provide limited annotations. To address these issues, we propose a task-oriented feature-fused network (TFN) for simultaneously solving face detection, landmark localization, and attribute analysis. In this network, a task-oriented feature-fused block is designed to learn task-specific feature combinations; then, an alternative multitask training scheme is presented to optimize each task with considering of their different learning capacities. We also present a large-scale face dataset called JFA in support of proposed method, which provides multivariate labels, including face bounding box, 68 facial landmarks, and 3 attribute labels (i.e., apparent age, gender, and ethnicity). The experimental results suggest that the TFN outperforms several multitask models on the JFA dataset. Furthermore, our approach achieves competitive performances on WIDER FACE and 300W dataset, and obtains state-of-the-art results for gender recognition on the MORPH II dataset. Xuxin Lin, Jun Wan 0001, Yiliang Xie, Chi Lin 0002, Yanyan Liang 0001, Guodong Guo, Stan Z. Li |
IEEE Trans. Cybern. | 6 |
| 2019 | Abnormal gesture recognition based on multi-model fusion strategy
Chi Lin 0002, Xuxin Lin, Yiliang Xie, Yanyan Liang 0001 |
Mach. Vis. Appl. | 4 |
| 2019 | Region-Based Context Enhanced Network for Robust Multiple Face AlignmentabstractThe recent studies for face alignment have involved developing an isolated algorithm on well-cropped face images. It is difficult to obtain the expected input by using an off-the-shelf face detector in practical applications. In this paper, we attempt to bridge between face detection and face alignment by establishing a novel joint multi-task model, which allows us to simultaneously detect multiple faces and their landmarks on a given scene image. In contrast to the pipeline-based framework by cascading separate models, we aim to propose an end-to-end convolutional network by sharing and transform feature representations between the task-specific modules. To learn a robust landmark estimator for unconstrained face alignment, three types of context enhanced blocks are designed to encode feature maps with multi-level context, multi-scale context, and global context. In the post-processing step, we develop a shape reconstruction algorithm based on point distribution model to refine the landmark outliers. Extensive experiments demonstrate that our results are robust for the landmark location task and insensitive to the location of estimated face regions. Furthermore, our method significantly outperforms recent state-of-the-art methods on several challenging datasets including 300 W, AFLW, and COFW. Xuxin Lin, Yanyan Liang 0001, Jun Wan 0001, Chi Lin 0002, Stan Z. Li |
IEEE Trans. Multim. | 2 |
| 2018 | Large-Scale Isolated Gesture Recognition Using a Refined Fused Model Based on Masked Res-C3D Network and Skeleton LSTMabstractIn this paper, we focus on large-scale isolated gesture recognition for RGB-D videos. We develop a novel ensemble method to explore deep spatio-temporal features using 3D Convolutional Neural Networks (CNNs) with residual architecture (Res-C3D) and build a time-series model with skeleton information based on Long Short Term Memory network (LSTM). First, relative positions and angles of different keypoints are extracted and used to build time-series model in LSTM. Obtaining the skeleton information (keypoints) of body and reserving arm regions with discarding other parts, masked Res-C3D is obtained, which decreases the effect of the background and other variations, as gestures are mainly derived from the arm or hand movements. Moreover, the weights of each voting sub-classifier being of advantage to a certain class in our ensemble model are adaptively obtained by training in place of fixed weights. Our experimental results show that the proposed method has obtained a state-of-the-art performance with accuracy 0.6842 in the IsoGD dataset. Chi Lin 0002, Jun Wan 0001, Yanyan Liang 0001, Stan Z. Li |
FG | 3 |
| 2007 | Applications of Complete Orthogonal V-system with Multiresolution PropertyabstractV-system, a new class of complete orthogonal system in L2[0,1], is the generalization of the well-known Haar system and consists of not only smooth functions but also discontinuous functions at multi-levels. Therefore, the V-system can be applied to represent the information with both continuous and discontinuous signals, which is fundamentally different from that of Fourier orthogonal system. In this paper, we first describe the multiresolution property of the V-system, then point out that a graphics group can be reconstructed precisely by finite terms of the V-series, finally discuss the classification and recognition of graphics groups using the V-descriptor. It shows from the experiments that the V-system can be applied to the fields of reconstruction of graphics group and pattern recognition. Xiaochun Wang, Yanyan Liang 0001, Hui Ma 0018, Ruixia Song |
CAD/Graphics | 2 |