VLDB 2026 Research / reviewers in the wild / expert
Yuanyuan Liu 0004
dblp:97/2119-4 · also Yuan-Yuan Liu 0004
· DBLP profile ↗
65ranked-venue papers
23as first author
54since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 13 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 8 first-author · 18 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 13 since 2021Databases, data management, data science and information retrieval · 11 · 3 first-author · 10 since 2021Computer networks · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PLA-MGRA: Multi-Granularity and Relation-Aware Learning for Efficient and Generalizable Protein-Ligand Binding Affinity PredictionabstractProtein-Ligand Affinity (PLA) prediction quantifies the interaction strength to guide rational drug design. Existing approaches typically analyze interaction at a single granularity and overlook tightly coupled relationships between protein and ligand in both structure and functionality, consequently yielding suboptimal representations, leading to significant performance drops in real-world scenarios. To address this problem, we propose PLA-MGRA, a minimalist and effective PLA prediction framework. Specifically, PLA-MGRA captures both fine-grained atomic details and coarse grained functional semantics within the 3D structure of protein–ligand complexes, through multi-granularity learning. To further parse the coupled protein–ligand relationships, we design relation-aware learning to enhance the binding nature of representations. Extensive experiments demonstrate that our method achieves state-of-the-art performance on multiple protein–ligand affinity prediction benchmarks, while also offering generalizability and interpretability. Shunfan Li, Jiangkai Long, Xin Zou 0001, Chang Tang, Yuanyuan Liu 0004, Xiao He 0010 |
AAAI | 5 |
| 2026 | When Genes Speak: A Semantic-Guided Framework for Spatially Resolved Transcriptomics Data ClusteringabstractSpatial transcriptomics enables gene expression profiling with spatial context, offering unprecedented insights into the tissue microenvironment. However, most computational models treat genes as isolated numerical features, ignoring the rich biological semantics encoded in their symbols. This prevents a truly deep understanding of critical biological characteristics. To overcome this limitation, we present SemST, a semantic-guided deep learning framework for spatial transcriptomics data clustering. SemST leverages Large Language Models (LLMs) to enable genes to "speak" through their symbolic meanings, transforming gene sets within each tissue spot into biologically informed embeddings. These embeddings are then fused with the spatial neighborhood relationships captured by Graph Neural Networks (GNNs), achieving a coherent integration of biological function and spatial structure. We further introduce the Fine-grained Semantic Modulation (FSM) module to optimally exploit these biological priors. The FSM module learns spot-specific affine transformations that empower the semantic embeddings to perform an element-wise calibration of the spatial features, thus dynamically injecting high-order biological knowledge into the spatial context. Extensive experiments on public spatial transcriptomics datasets show that SemST achieves state-of-the-art clustering performance. Crucially, the FSM module exhibits plug-and-play versatility, consistently improving the performance when integrated into other baseline methods. Jiangkai Long, Yanran Zhu, Chang Tang, Kun Sun 0002, Yuanyuan Liu 0004 |
AAAI | 5 |
| 2026 | SGAT: Learning Feature Matching with Singularity-enhanced Graph Attention NetworkabstractThe task of image feature matching aims to establish correct correspondences between images from two different views. While approaches based on attention mechanisms have demonstrated remarkable advancements in image feature matching, they still encounter substantial limitations. Specifically, current graph attention network approaches face performance bottlenecks in complex scenarios, such as low-texture regions or occlusions. This limitation stems from the self-attention mechanism, which, when lacking effective guidance, can lead to divergent attention weights or incorrect focus on regions with low discriminability, resulting in matching failures in low-texture environments. Inspired by how humans focus on distinctive regions when performing cross-view matching, we enhance attention to singular points in images that are salient, unique and have high cross-view matching potential during information aggregation, thereby improving matching capability. To realize the aforementioned strategies, we develop a novel Singularity-enhanced Graph Attention Network (SGAT). SGAT leverages Co-potentiality and Multi-Scale Singularity as prior guidance, and designs a Singularity-aware Attention mechanism and a Co-potentiality Guided Attention mechanism , specifically enhancing the perception of singularity and matching potential during feature interaction. Experimental results on multiple datasets, including ScanNet1500, demonstrate that our method outperforms current state-of-the-art sparse matching methods. In particular, the improvement is most pronounced in complex scenarios such as low-texture environments, significantly enhancing the accuracy and robustness of image matching and its downstream tasks. Kun Sun 0002, Chang Tang, Yuanyuan Liu 0004, Xin Li 0005 |
AAAI | 4 |
| 2026 | Unsupervised multimodal emotion-unified representation learning with dual-level language-driven cross-modal emotion alignment
Shaoze Feng, Qiyin Zhou, Yuanyuan Liu 0004, Kejun Liu, Chang Tang |
Pattern Recognit. | 3 |
| 2026 | Mutual Calibration Network for Multi-View ClusteringabstractMulti-view clustering, which uses information from multiple views to partition data into distinct clusters, has garnered significant attention. Existing MLP-based and GCN-based algorithms primarily focus on enhancing performance by extracting node attribute and graph structure features, and then integrating them for clustering. However, features extracted by MLP and GCN may differ in quality due to the high sparsity and noise in multi-view data. Directly integrating these features can lead to feature contamination. To address this issue, we propose a novel mutual calibration network for multi-view clustering (McMVC). Specifically, the attribute and structural features from multi-view data are integrated separately using the attention fusion module. A classifier is then employed to obtain high-confidence pseudo-labels for these fused features. We design a Cluster-level Mutual Calibration (CLMC) module that uses the pseudo-labels to mutually calibrate view features while maintaining compact cluster structures. During training, the attribute and structural features are not directly integrated to avoid feature contamination but are jointly optimized by reliable class information. Concurrently, we construct a Centroid-level Contrastive Calibration (CLCC) module to map view features into their inner centroid space and learn more discriminative centroid representations. Our approach outperforms state-of-the-art methods in multi-view data clustering, as demonstrated by extensive experiments on six real-world benchmark datasets. The source code is available at https://github.com/YuangXiao/ McMVC. Yuang Xiao, Chang Tang, Weiqing Yan, Yuanyuan Liu 0004, Xinwang Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Task-Aware Information Decoupling for Multimodal ClusteringabstractMultimodal clustering (MMC) overcomes the limitations of unimodal methods by integrating information from multiple sources, but the complexity of heterogeneous information coupling hinders effective feature extraction. Critically, existing MMC paradigms primarily focus on capturing consensus through coarse-grained cross-modal alignment. However, such task-agnostic strategies overlook the differences in the utility of feature information across varying task environments. In the absence of task-centric guidance, models often struggle to effectively distinguish task-relevant critical information from task-irrelevant redundant noise during the disentanglement process, leading to information confusion in the representation space. To address this challenge, we propose a deep disentangled multimodal clustering method guided by information theory, named DRLMMC, which employs a tripartite information optimization mechanism to achieve deep disentanglement of cross-modal representations. 1) We design modality-specific encoders to construct nonlinear mapping spaces, transforming the reconstruction mechanism of autoencoders into an information-theoretic mutual information (MI) constraint problem, preserving the unique features of different modalities; 2) To establish cross-modal semantic associations, it constructs a cross-modal shared information extraction module, and, based on an information-theoretic framework, designs an optimization objective function to progressively align multimodal feature subspaces through MI maximization and contrastive learning, capturing task-relevant invariant features across modalities; 3) A unique information dynamic perception module is proposed, which employs a conditional MI projection network combined with learning distribution regularization to adaptively extract and enhance modality-specific task-relevant unique information. Experimental results demonstrate that DRLMMC outperforms existing state-of-the-art methods on multimodal benchmark datasets, exhibiting excellent generalization ability. Notably, it achieves precise disentanglement of cross-omics features in multi-omics analysis, offering a novel methodological approach for handling complex biomedical data. Zixiao Jin, Chang Tang, Chuankun Li, Yuanyuan Liu 0004, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2026 | Contrastive and Dual Adversarial Representation Learning for Multi-View ClusteringabstractMulti-View Clustering (MVC) has gained increasing attention due to its ability to effectively leverage the complementary information of multi-view data. Despite the success of existing MVC methods in many real-world applications, they often overlook the discrepancy of view-specific latent distribution and struggle to ensure the completeness of the multi-view data. To address these challenges and harness the powerful feature extraction capability of deep networks, we propose a novel Contrastive and Dual Adversarial Representation Learning method for Multi-view Clustering, termed as CDARL, to solve multi-view clustering problems with both complete and incomplete multi-view data. Specifically, CDARL employs alternating adversarial and contrastive learning to align the view-specific representations, driving them into the same semantic latent space to minimize the discrepancy in view-specific distributions. In addition, a consensus latent representation is learned by an adaptive fusion block that integrates information from multiple views. The consensus representation is further refined through adversarial learning modeling the transformation of the standard Gaussian distribution to the original data distribution. Moreover, the proposed method incorporates an imputation strategy designed to handle the incomplete multi-view data clustering task. This strategy utilizes both reconstructed samples and cross-view neighbors to impute missing views from the latent space and the original space, thereby preserving clustering information, which ensures the quality and feasibility of the imputed samples. Experimental results on six widely used datasets have verified the competitiveness of the proposed CDARL method against state-of-the-art methods in MVC problems with complete and incomplete multi-view data. Code is available athttps://github.com/xywy220/CDARL-MVC. Yanwanyu Xi, Chang Tang, Junjie Huang 0001, Xingchen Hu 0001, Yuanyuan Liu 0004, Xinwang Liu 0002 |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2025 | LRGR: Self-Supervised Incomplete Multi-View Clustering via Local Refinement and Global RealignmentabstractIncomplete Multi-View Clustering (IMVC) aims to explore comprehensive representations from multiple views with missing samples. Recent studies have revealed that IMVC methods benefit from Graph Convolutional Network (GCN) in achieving robust feature imputation and effective representation learning. Despite these notable improvements, GCN imputation methods often cause a distribution shift between the imputed and original representations, particularly when the neighbors of the imputed nodes are assigned to different groups. Moreover, GCN learning methods tend to produce homogeneous imputed representations, which blur cluster boundaries and hinder effective discriminative clustering. To remedy these challenges, the Local Refinement and Global Realignment (LRGR) Self-supervised model is proposed for incomplete multi-view clustering, which includes two stages. In the first stage, a local imputed refinement module is designed to enhance the versatility of imputed representations through cross-view contrastive learning guided by view-specific prototypes. In the second stage, a global realignment module is introduced to achieve semantic consistency across views, alleviating distribution shifts by leveraging pseudo-labels and their corresponding confidence scores as guidance. Experiments on five widely used multi-view datasets demonstrate the competitiveness and superiority of our method compared to state-of-the-art approaches. Yanwanyu Xi, Chang Tang, Xingchen Hu 0001, Yuanyuan Liu 0004, Xinwang Liu 0002 |
IJCAI | 5 |
| 2025 | Spatially Resolved Transcriptomics Data Clustering with Tailored Spatial-scale ModulationabstractSpatial transcriptomics, comprising spatial location and high-throughput gene expression information, provides revolutionary insights into disease discovery and cellular evolution. Spatial transcriptomic clustering, which pinpoints distinct spatial domains within tissues, reveals cellular interactions and enhances our understanding of the intricate architecture of tissues. Existing methods typically construct spatial graphs using a static radius based on spatial coordinates, which hinders the accurate identification of spatial domains and complicates the precise partitioning of boundary nodes within clusters. To address this issue, we introduce a novel spatially resolved transcriptomics data clustering network (TSstc). Specifically, we employ a tailored spatial-scale modulation approach, constructing different spatial graphs incrementally as the radius of the spatial domain expands, and a Spatiality-Aware Sampling (SAS) strategy is proposed to aggregate node representations by considering the spatial dependencies between spots. We then use GCN encoders to learn gene embedding with gene graph and multiple spatial embeddings with spatial graphs. During training, we incorporate cross-view correlation-based tailored spatial regularization constraints to preserve high-quality neighbor relationships across spatial embeddings at different scales. Finally, a zero-inflated negative binomial model is utilized to capture the global probability distribution of gene expression profiles. Extensive experimental results demonstrate that our approach surpasses existing state-of-the-art methods in clustering tasks and related downstream applications. Yuang Xiao, Yanran Zhu, Chang Tang, Yuanyuan Liu 0004, Kun Sun 0002, Xinwang Liu 0002 |
IJCAI | 5 |
| 2025 | ClothHMR: 3D Mesh Recovery of Humans in Diverse Clothing from Single ImageabstractWith 3D data rapidly emerging as an important form of multimedia information, 3D human mesh recovery technology has also advanced accordingly. However, current methods mainly focus on handling humans wearing tight clothing and perform poorly when estimating body shapes and poses under diverse clothing, especially loose garments. To this end, we make two key insights: (1) tailoring clothing to fit the human body can mitigate the adverse impact of clothing on 3D human mesh recovery, and (2) utilizing human visual information from large foundational models can enhance the generalization ability of the estimation. Based on these insights, we propose ClothHMR, to accurately recover 3D meshes of humans in diverse clothing. ClothHMR primarily consists of two modules: clothing tailoring (CT) and FHVM-based mesh recovering (MR). The CT module employs body semantic estimation and body edge prediction to tailor the clothing, ensuring it fits the body silhouette. The MR module optimizes the initial parameters of the 3D human mesh by continuously aligning the intermediate representations of the 3D mesh with those inferred from the foundational human visual model (FHVM). ClothHMR can accurately recover 3D meshes of humans wearing diverse clothing, precisely estimating their body shapes and poses. Experimental results demonstrate that ClothHMR significantly outperforms existing state-of-the-art methods across benchmark datasets and in-the-wild images. Additionally, a web application for online fashion and shopping powered by ClothHMR is developed, illustrating that ClothHMR can effectively serve real-world usage scenarios. The code and model for ClothHMR are available at: https://github.com/starVisionTeam/ClothHMR. Yunqi Gao, Leyuan Liu 0001, Yuhan Li 0009, Changxin Gao, Yuanyuan Liu 0004, Jingying Chen 0001 |
ICMR | 5 |
| 2025 | Find True Collaborators: Banzhaf Index-based Cross View Alignment for Partially View-aligned ClusteringabstractPartially view-aligned clustering (PVC) has emerged as a critical area in multi-view clustering, addressing the inherent instance misalignment across views during data collection. The primary challenge of PVC is accurately establishing correspondences between cross-view samples. The Banzhaf index in cooperative game theory serves as an effective tool for modeling complex relationships between multi-view samples by quantifying the marginal contributions of coalition members to collaborative benefits. To this end, we propose a Banzhaf Index-driven cross-view aligNment method, dubbed BIN, which systematically evaluates each view sample's contribution to joint decision-making within a game-theoretic framework. This approach overcomes the limitations of existing PVC methods reliant on prior alignment information and enhances the robustness of multi-view matching. Specifically, we model multi-view samples as players in a cooperative game and quantify their interactions using a payoff model. Simultaneously, we propose a dual-loss constraint: (1) Banzhaf gain loss, which dynamically captures the marginal contribution of key cross-view sample pairs to reinforce associations; (2) contrast loss, which applies exclusion constraints in the feature space to suppress interference from weakly correlated samples. Together, these losses form an effective optimization mechanism. This game-theoretic approach adaptively learns sample correspondences without pre-alignment and ensures robust matching in complex misalignment scenarios. Extensive experiments demonstrate that our method achieves competitive performance against eight state-of-the-art PVC algorithms. Shanghui Deng, Chang Tang, Kun Sun 0002, Yuanyuan Liu 0004, Xinwang Liu 0002 |
ACM Multimedia | 5 |
| 2025 | Sample-Cohesive Pose-Aware Contrastive Facial Representation LearningabstractAbstract Self-supervised facial representation learning (SFRL) methods, especially contrastive learning (CL) methods, have been increasingly popular due to their ability to perform face understanding without heavily relying on large-scale well-annotated datasets. However, analytically, current CL-based SFRL methods still perform unsatisfactorily in learning facial representations due to their tendency to learn pose-insensitive features, resulting in the loss of some useful pose details. This could be due to the inappropriate positive/negative pair selection within CL. To conquer this challenge, we propose a Pose-disentangled Contrastive Facial Representation Learning (PCFRL) framework to enhance pose awareness for SFRL. We achieve this by explicitly disentangling the pose-aware features from non-pose face-aware features and introducing appropriate sample calibration schemes for better CL with the disentangled features. In PCFRL, we first devise a pose-disentangled decoder with a delicately designed orthogonalizing regulation to perform the disentanglement; therefore, the learning on the pose-aware and non-pose face-aware features would not affect each other. Then, we introduce a false-negative pair calibration module to overcome the issue that the two types of disentangled features may not share the same negative pairs for CL. Our calibration employs a novel neighborhood-cohesive pair alignment method to identify pose and face false-negative pairs, respectively, and further help calibrate them to appropriate positive pairs. Lastly, we devise two calibrated CL losses, namely calibrated pose-aware and face-aware CL losses, for adaptively learning the calibrated pairs more effectively, ultimately enhancing the learning with the disentangled features and providing robust facial representations for various downstream tasks. In the experiments, we perform linear evaluations on four challenging downstream facial tasks with SFRL using our method, including facial expression recognition, face recognition, facial action unit detection, and head pose estimation. Experimental results show that PCFRL outperforms existing state-of-the-art methods by a substantial margin, demonstrating the importance of improving pose awareness for SFRL. Our evaluation code and model will be available at https://github.com/fulaoze/CV/tree/main . Yuanyuan Liu 0004, Shaoze Feng, Yibing Zhan, Dapeng Tao, Zijing Chen, Zhe Chen 0013 |
Int. J. Comput. Vis. | 1 |
| 2025 | Noise-Resistant Multimodal Transformer for Emotion Recognition
Yuanyuan Liu 0004, Haoyu Zhang 0001, Yibing Zhan, Zijing Chen, Guanghao Yin, Zhe Chen 0013 |
Int. J. Comput. Vis. | 1 |
| 2025 | Degradation-adaptive attack-robust self-supervised facial representation learning
Yuanyuan Liu 0004, Chang Tang, Kun Sun 0002, Yibing Zhan, Zhe Chen 0013 |
Neurocomputing | 2 |
| 2025 | Feature Contrast Difference and Enhanced Network for RGB-D Indoor Scene Classification in Internet of ThingsabstractThe era of smart connectivity spawned by the Internet of Things (IoT) has made the need to achieve environmental perception and understanding of different scenarios increasingly urgent. Among the many scenarios, indoor scene classification has attracted much attention because of its relevance to the daily lives of people, ranging from comfort regulation in living spaces to the optimal allocation of resources in offices, and a variety of approaches for this task have emerged. However, increasing accuracy remains a crucial objective due to the complexity and disorder of indoor scenes. Therefore, we propose a feature contrast difference and enhanced network for RGB-D indoor scene classification, FCDENet. First, the red, green, and blue and Depth images express different information. Therefore, we built a feature contrast difference module for the first two low-level features to extend the receptive fields of the different features, utilizing differential contrast to complement each other. Second, the high-level feature semantic information is abstract. Therefore, we introduced information cluster blocks, which are used to aggregate feature points with similar attributes into compact clusters after being parsed by an initial frequency transform, enabling instantiated representations of the semantic information. Finally, to further enhance the integrated features, we introduced a wavelet transform block in the cross-layer decoding process. In contrast to conventional decoding methods, we employed a wavelet transform for initial denoising cross-layer features and used multiple pooling structures to supplement local information, gradually weighting to achieve higher prediction accuracy. Extensive experiments on two typical indoor datasets, NYUDv2 and SUN RGB-D, show that our results exhibit excellent performance. In addition, to better demonstrate the reliability of the method, we conducted generalizability experiments on other datasets, and the proposed method provides a robust solution to the challenges of multiple scenarios in the era of IoT smart connectivity. The code is available athttps://github.com/XUEXIKUAIL/FCDENet. Wujie Zhou, Bitao Jian, Yuanyuan Liu 0004 |
IEEE Internet Things J. | 3 |
| 2025 | Discriminative Binary Multi-View Clustering
Yun-Ning You, Chang Tang, Xinwang Liu 0002, Yuanyuan Liu 0004, Xian-Ju Li, Liang-Xiao Jiang |
J. Comput. Sci. Technol. | 5 |
| 2025 | Beyond boundaries: Hierarchical-contrast unsupervised temporal action localization with high-coupling feature learning
Yuanyuan Liu 0004, Leyuan Liu 0001, Wujie Zhou, Chang Tang |
Pattern Recognit. | 1 |
| 2025 | Leveraging Eye Movement for Instructing Robust Video-Based Facial Expression RecognitionabstractVideo-based facial expression recognition (VFER) is challenging due to variations caused by cultural background and expression camouflage. To tackle these problems, researchers introduced eye movement signals to complement visual information. However, existing methods either require expensive devices to capture high-quality eye movements or can only extract low-quality eye movements visually, making them ineffective in the real world. To address this, we propose an eye movement-instructed VFER (EM-VFER) that leverages high-quality eye movements to instruct the visual learning, obtaining robust performance without requiring costly devices during inference. Specifically, our EM-VFER operates in two stages: the high-quality eye movement pre-training stage and the eye movement-instructed video fine-tuning stage. In the pre-training, we compile an Eye-behavior-aided Multimodal Emotion Recognition (EMER) dataset and use it to train a multimodal Transformer. During the fine-tuning, we propose a novel progressive eye movement-instructed learning to take better advantage of the prior knowledge about high-quality eye movement signals from EMER. The instructed fine-tuning model could then make more robust predictions on downstream facial expression datasets. We evaluate our approach on three macroexpression datasets (DFEW, MAFW and Aff-wild2) and two micro-expression datasets (CASME III and CASME II). The results demonstrate that EM-VFER significantly outperforms existing methods. The code will be available. Yuanyuan Liu 0004, Kejun Liu, Zijing Chen, Zhe Chen 0013, Chang Tang, Jingying Chen 0001, Shiguang Shan |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Enhancing RGB-D Mirror Segmentation With a Neighborhood-Matching and Demand-Modal Adaptive Network Using Knowledge DistillationabstractRecent breakthroughs in computer vision have led to remarkable progress in the areas of autonomous vehicles and robotics. However, ordinary objects such as mirrors pose unique challenges to computer vision systems owing to occlusion, reflection, and distortion. Moreover, existing deep learning models suffer from issues such as excessive parameters and high computational complexity, making it challenging to implement numerous studies offline. To address these issues, we propose an innovative solution: a neighborhood-matching and demand-modal adaptive network using knowledge distillation (KD), called NDANet-S$^{\ast }$, specifically designed for red-green-blue depth mirror segmentation. NDANet-S$^{\ast }$operates by iteratively matching detailed and semantic difference between neighborhood features during the encoding phase. It then complements information across different modalities through demand-modal adaptation, enhancing heteromodal cross-complementation during the KD stage. In the decoding phase, semantic enhancement features and iterative encoding features are deeply integrated, forming a strong foundation for multistage progressive knowledge transfer in the KD process. Furthermore, we introduce a multistage teacher-assisted KD scheme, guided by sample complexity, to work synergistically with the mirror segmentation model. This innovative scheme includes a sample complexity rater, heterogeneous cross-complementarity, and hierarchical progressive knowledge transfer. Experimental evaluations on publicly available datasets indicate that NDANet-S$^{\ast }$significantly enhances segmentation accuracy while preserving a consistent number of parameters. Additionally, it achieves state-of-the-art performance in mirror segmentation. The source code for our model is publicly available and can be accessed at:https://github.com/2021nihao/NMDANet. Note to Practitioners—This study presents a neighborhood-matching and demand-modal adaptive network for RGB-D mirror segmentation, incorporating knowledge distillation (KD). The combination of the KD framework and segmentation network enables sample complexity discrimination, cross-modal distillation, and multilevel distillation (guided by the former) to achieve targeted distillation. The newly proposed NDANet-S$^{\ast }$surpasses current state-of-the-art methods, with a reduction of 93.2% in parameters and 87.65% in floating-point operations (FLOPs) compared to NDANet-T. Wujie Zhou, Yuanyuan Liu 0004, Ting Luo 0001 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2025 | Asymmetric Dual-Encoder Network With Clustering and Mutual Contrast Loss for the Semantic Segmentation of Remote-Sensing ImagesabstractIn recent years, the semantic segmentation of multimodal remote-sensing images using convolutional methods has received significant attention. Owing to the localized nature of convolutional operations, existing methods use the attention method to obtain global relations. However, effectively complementing and eliminating the global and local relations between two modalities has become a major challenge. In addition, category imbalance often affects the model performance as a result of the high resolution of remote-sensing images. In this study, we propose an asymmetric dual encoder network with clustering mutual contrast loss. Specifically, we use a convolutional neural network and Transformer as the backbone to extract two modal features in parallel to learn the local and global information, respectively. Next, our proposed multimodal hierarchical interaction module and dynamic weight inspired block efficiently fuse the multimodal features to complement the local and global information. The fused features are fed into our proposed local context extraction module and global context extraction module. Furthermore, to address the challenge presented by feature class imbalances, we apply a clustering algorithm to classify each pixel, which is subsequently inter-supervised using the inter-contrast loss. Extensive experiments on benchmark datasets show that the proposed model is extremely effective in the semantic segmentation of remotely sensed images and that it outperforms current state-of-the-art networks both quantitatively and qualitatively. Our code and results can be found at https://github.com/LYZ00918/AMCNet. Wujie Zhou, Yangzhen Li, Yuanyuan Liu 0004 |
IEEE Trans. Big Data | 3 |
| 2025 | scSPAF: Cell Similarity Purified Adaptive Fusion Network for Single-Cell Multi-Omics ClusteringabstractThe rapid advancement of single-cell sequencing technology has generated vast amounts of multi-omics data, presenting unprecedented opportunities for single-cell multi-omics clustering analysis. However, existing single-cell clustering algorithms focus on extracting shared representations, overlooking the interactions and correlations among cells. This oversight inevitably leads to biased or confounded cell clustering results. In this paper, we propose a cell similarity purified adaptive fusion network for single-cell multi-omics clustering, named scSPAF, which adopts a multi-level fusion approach to thoroughly explore the consistency and complementarity of omics data. Specifically, we design a cell similarity purification module to accurately incorporate neighborhood information among cells into cell features, thereby purifying the latent representation that reflects cell correlations. In addition, we align attribute features from different omics to extract the consistent representation across omics. Simultaneously, by employing an adaptive fusion mechanism, we integrate representations of omics-specific and the consistent representation across omics to generate more discriminative representations of omics, further enhancing the clustering performance. Experimental results obtained from six real-world datasets demonstrate the superiority of the scSPAF algorithm when compared with other state-of-the-art methods. Shanghui Deng, Chang Tang, Xinwang Liu 0002, Yuanyuan Liu 0004, Shan An |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | MASDG: Multiview Augmented Single-Source Domain Generalization Method for Robust Remote Sensing Building ExtractionabstractDespite advances in deep learning for remote sensing building extraction (RSBE), Multi-target Domain RSBE (MD-RSBE) remains challenging, as it requires transferring knowledge from a labeled source domain to multiple unlabeled target domains, with domain shifts in texture, style, and semantics. Existing domain adaptation (DA) and generalization (DG) methods face significant limitations: DA requires target-domain training, while DG needs multi-source training, leading to high training costs and low generalization in practical MD-RSBE scenarios. To address this, we propose a Multi-view Augmented Single-source Domain Generalization (MASDG) method, which effectively mitigates domain shifts across RS source and target domains for robust MD-RSBE performance by enriching the diversity of the source domain through multi-view augmentation and enforcing semantic consistency. Specifically, MASDG consists of three key components: Texture-level Domain Augmentation (TDA) module, Style-level Domain Augmentation (SDA) module and Semantic-invariant Representation Learning (SRL). To mitigate texture-level domain shift, TDA first introduces parameter-optimized multi-layer random convolution to modify the texture of source image, generating texture-augmented image pairs for simulating real-world texture diversity across various RS domains. Then, with each image pair from TDA, SDA employs two paralleled encoders, namely the general feature encoder and the batch-guided style encoder, to formulate multi-view building features, further mitigating style-level domain shift. Finally, SRL ensures semantic-invariant representation learning via a dual mechanism, including multi-view segmentation loss and semantic consistency loss. The former generates predictions from diverse feature views (original, texture-augmented, style-augmented, etc.), while the latter performs semantic alignment by minimizing distribution discrepancies among predictions, bridging semantic inconsistency to enable robust segmentation. Extensive experiments across three different MD-RSBE settings with 7 different target domains demonstrate that our MASDG outperforms existing state-of-the-art methods by a significant margin. Yunjiao Liu, Yuanyuan Liu 0004, Kejun Liu, Chang Tang, Wujie Zhou, Zhe Chen 0013, Wei Xiang 0001, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | PMTSeg: Prompt-Driven Multimodal Transformer for Task-Adapted Remote Sensing Image SegmentationabstractMultimodal remote sensing image segmentation (MRSIS) is important for intelligent remote sensing image (RS) interpretation, which encompasses three distinct tasks: semantic segmentation, instance segmentation, and panoptic segmentation. Existing methods typically address individual tasks with specialized models, limiting generalization and real-world applicability. Multi-task learning approaches have introduced separated task heads to unify tasks, yet we identify two key challenges when directly applying them to MRSIS: (1) the modality gap, arising from semantic discrepancies and granularity discrepancies across RS modalities, and (2) the task gap, due to varying preferences in learning different segmentation tasks. To overcome these challenges, we propose PMTSeg—a novel Prompt-driven Multimodal Transformer for task-adapted MRSIS. PMTSeg integrates three key components: (1) Task-common Multimodal Affinity Approximation (TMAA), (2) Task-common Multi-scale Semantic Fusion (TMSF), and (3) a unified Prompt-driven Segmentation Head (PSH). First, TMAA addresses the modality gap by approximating inter-modal affinity matrices, extracting task-common features across modalities and aligning semantic information. Then, TMSF further integrates these features using the scale-matched fusion at multiple scales to produce enriched, multi-scale task-common features. Moreover, to address the task gap, the PSH leverages task-adapted text prompts and task-adapted contrastive loss to model relationships across tasks, enabling adaptive optimization for robust and universal MRSIS performance. Extensive experiments on three MRSIS datasets—VALID, SEMCITY TOULOUSE, and UBCV2—demonstrate that PMTSeg significantly surpasses state-of-the-art methods in all three segmentation tasks, offering a unified and accurate solution to MRSIS. Kejun Liu, Xuesong Yan 0001, Yuanyuan Liu 0004, Chang Tang, Yibing Zhan, Wujie Zhou, Hongyan Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | HLMamba: Hybrid Lightweight Mamba-Based Fusion Network for Dense Prediction of Remote Sensing Images
Wujie Zhou, Penghan Yang, Yuanyuan Liu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Emotion-Oriented Cross-Modal Prompting and Alignment for Human-Centric Emotional Video CaptioningabstractHuman-centric Emotional Video Captioning (H-EVC) aims to generate fine-grained, emotion-related sentences for human-based videos, enhancing the understanding of human emotions and facilitating human-computer emotional interaction. However, existing video captioning methods often overlook subtle emotional clues and interactions in videos. As a result, the generated captions frequently lack emotional information. To address this, we proposeEmotion-orientedCross-modalPrompting andAlignment (ECPA), which improves HEVC accuracy by modeling fine-grained visual-textual emotion clues. Using large foundation models, ECPA introduces two learnable prompting strategies: visual emotion prompting (VEP) and textual emotion prompting (TEP), along with an emotion-oriented cross-modal alignment (ECA) module. VEP uses two levels of visual prompts,i.e., emotion recognition (ER) and action unit (AU), to focus on both coarse and fine visual emotional features. TEP devise two-level learnable textual prompts,i.e., sentence-level emotional tokens and word-level masked tokens to capture global and local textual emotion representations. ECA introduces another two levels of emotion-oriented prompt alignment learning mechanisms: the ER-sentence level and the AU-word level alignment losses. Both enhance the model's ability to capture and integrate both global and local cross-modal emotion semantics, thereby enabling the generation of fine-grained emotional linguistic descriptions in video captioning. Experiments show ECPA significantly outperforms state-of-the-art methods on various H-EVC datasets (relative improvements of 9.98%, 5.72%, 4.46%, 24.52% on MAFW, and 12.82%, 20.27%, 4.22%, 5.01% on EmVidCap across four evaluation metrics) and supports zero-shot tasks on MSVD and MSRVTT, demonstrating strong applicability and generalization. Yu Wang 0246, Yuanyuan Liu 0004, Shunping Zhou, Chang Tang, Wujie Zhou, Zhe Chen 0013 |
IEEE Trans. Multim. | 2 |
| 2025 | Smooth Multiple Kernel k-Means via Underlying Graph FilteringabstractClustering has attracted more and more attention as one of the most fundamental techniques in the field of unsupervised learning. To deal with nonlinear problems, clustering methods have been extended to the kernel version. As a traditional kernel clustering algorithm, multiple kernel k-means (MKKM) aims to learn clustering results from a consensus kernel obtained by combining a set of predefined kernels optimally. However, we observe that the existing MKKM algorithm and its variants insufficiently consider the noise that existed in kernel space and the underlying structure of kernelized data points. To this end, we propose a novel smooth MKKM via underlying graph filtering (SMKKM-UGF) to learn the smooth representations of kernelized data points through their nearby nodes in the underlying graph. In particular, different from the common graph filter, we jointly update the graph filter while learning the smooth kernel, so that the graph filter can be guaranteed to adapt to the updating kernel space constantly. Besides, an iterative algorithm with proven convergence is designed to solve the resultant optimization problem. Extensive experiments have been performed on numerous benchmark datasets, whose results prove the superiority of the proposed SMKKM-UGF compared to the other state-of-the-art clustering methods. The demo code of this work is publicly available at https://github.com/wqyang23/SMKKM-UGF.git. Wenqi Yang, Chang Tang, Xinwang Liu 0002, Guanghui Yue 0001, Yuanyuan Liu 0004, Changqing Zhang 0002, En Zhu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Multiattentive Perception and Multilayer Transfer Network Using Knowledge Distillation for RGB-D Indoor Scene ParsingabstractScene parsing has gained wide attention in the field of computer vision, with emerging methods and techniques providing superior solutions. Although some methods have improved performance, they tend to neglect the number of model parameters and computational size, which makes achieving real-time operation in practical applications challenging. To address these limitations, we propose a multiattentive perception and multilayer transfer network that employs knowledge distillation (MPMTNet-KD), which is generated by a student network (MPMTNet-S) under the guidance of a teacher network (MPMTNet-T) with the aid of our proposed multilayer transfer knowledge distillation (KD) methods. To capture complete information from different modalities, a multiattentive perception module (MAPM) is introduced to mine features from various perspectives, and hetero-oriented sensing (HOS) convolution is utilized to integrate cross-layer features in a single and holistic manner. Importantly, we introduce multilayer transfer KD to explore the different knowledge types between layers, as well as intraclass and interclass correlations. In addition, we use the discrete cosine transform (DCT) approach combined with filtering during the KD process to mitigate noise that may be induced by the depth map, thereby improving the depth information and further enhancing the knowledge transfer effect. We conducted comprehensive experiments on two challenging indoor benchmark datasets, namely NYUDv2 and SUN RGB-D. Compared with existing methods, the proposed MPMTNet-KD reduces the number of parameters from 125.8 M in MPMTNet-T to 28.3 M in MPMTNet-S, achieving a mean intersection over union (mIoU) of 54.9% in the indoor scene parsing task. MPMTNet-KD was also evaluated on two additional public datasets, namely MFNet and PST900, to demonstrate its generalization capacity. The source code is available at https://github.com/XUEXIKUAIL/MPMTNet. Wujie Zhou, Bitao Jian, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | VS: Reconstructing Clothed 3D Human from Single Image via Vertex ShiftabstractVarious applications require high-fidelity and artifact free 3D human reconstructions. However, current implicit function-based methods inevitably produce artifacts while existing deformation methods are difficult to reconstruct high-fidelity humans wearing loose clothing. In this paper, we propose a two-stage deformation method named Vertex Shift (VS) for reconstructing clothed 3D humans from single images. Specifically, VS first stretches the estimated SMPL-X mesh into a coarse 3D human model using shift fields inferred from normal maps, then refines the coarse 3D human model into a detailed 3D human model via a graph convolutional network embedded with implicit-function-learned features. This “stretch-refine” strategy addresses large deformations required for reconstructing loose clothing and delicate deformations for recovering intricate and detailed surfaces, achieving high-fidelity reconstructions that faithfully convey the pose, clothing, and surface details from the input images. The graph convolutional network's ability to exploit neighborhood vertices coupled with the advantages inherited from the deformation methods ensure VS rarely produces artifacts like distortions and non-human shapes and never produces artifacts like holes, broken parts, and dismembered limbs. As a result, VS can reconstruct highfidelity and artifact-less clothed 3D humans from single images, even under scenarios of challenging poses and loose clothing. Experimental results on three benchmarks and two in-the-wild datasets demonstrate that VS significantly outperforms current state-of-the-art methods. The code and models of VS are available for research purposes at https://github.com/starVisionTeam/VS. Leyuan Liu 0001, Yuhan Li 0009, Yunqi Gao, Changxin Gao, Yuanyuan Liu 0004, Jingying Chen 0001 |
CVPR | 5 |
| 2024 | Open-Set Video-based Facial Expression Recognition with Human Expression-sensitive PromptingabstractIn Video-based Facial Expression Recognition (V-FER), models are typically trained on closed-set datasets with a fixed number of known classes. However, these models struggle with unknown classes common in real-world scenarios. In this paper, we introduce a challenging Open-set Video-based Facial Expression Recognition (OV-FER) task, aiming to identify both known and new, unseen facial expressions. While existing approaches use large-scale vision-language models like CLIP to identify unseen classes, we argue that these methods may not adequately capture the subtle human expressions needed for OV-FER. To address this limitation, we propose a novel Human Expression-Sensitive Prompting (HESP) mechanism to significantly enhance CLIP's ability to model video-based facial expression details effectively. Our proposed HESP comprises three components: 1) a textual prompting module with learnable prompts to enhance CLIP's textual representation of both known and unknown emotions, 2) a visual prompting module that encodes temporal emotional information from video frames using expression-sensitive attention, equipping CLIP with a new visual modeling ability to extract emotion-rich information, and 3) an open-set multi-task learning scheme that promotes interaction between the textual and visual modules, improving the understanding of novel human emotions in video sequences. Extensive experiments conducted on four OV-FER task settings demonstrate that HESP can significantly boost CLIP's performance (a relative improvement of 17.93% on AUROC and 106.18% on OSCR) and outperform other state-of-the-art open-set video understanding methods by a large margin. Code is available at https://github.com/cosinehuang/HESP. Yuanyuan Liu 0004, Yibing Zhan, Zijing Chen, Zhe Chen 0013 |
ACM Multimedia | 1 |
| 2024 | Token-disentangling Mutual Transformer for multimodal emotion recognition
Guanghao Yin, Yuanyuan Liu 0004, Haoyu Zhang 0001, Fang Fang 0008, Chang Tang, Liangxiao Jiang |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | DGPINet-KD: Deep Guided and Progressive Integration Network With Knowledge Distillation for RGB-D Indoor Scene AnalysisabstractSignificant advancements in RGB-D semantic segmentation have been made owing to the increasing availability of robust depth information. Most researchers have combined depth with RGB data to capture complementary information in images. Although this approach improves segmentation performance, it requires excessive model parameters. To address this problem, we propose DGPINet-KD, a deep-guided and progressive integration network with knowledge distillation (KD) for RGB-D indoor scene analysis. First, we used branching attention and depth guidance to capture coordinated, precise location information and extract more complete spatial information from the depth map to complement the semantic information for the encoded features. Second, we trained the student network (DGPINet-S) with a well-trained teacher network (DGPINet-T) using a multilevel KD. Third, an integration unit was developed to explore the contextual dependencies of the decoding features and to enhance relational KD. Comprehensive experiments on two challenging indoor benchmark datasets, NYUDv2 and SUN RGB-D, demonstrated that DGPINet-KD achieved improved performance in indoor scene analysis tasks compared with existing methods. Notably, on the NYUDv2 dataset, DGPINet-KD (DGPINet-S with KD) achieves a pixel accuracy gain of 1.7% and a class accuracy gain of 2.3% compared with DGPINet-S. In addition, compared with DGPINet-T, the proposed DGPINet-KD (DGPINet-S with KD) utilizes significantly fewer parameters (29.3M) while maintaining accuracy. The source code is available at https://github.com/XUEXIKUAIL/DGPINet. Wujie Zhou, Bitao Jian, Meixin Fang, Xiena Dong, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Multitarget Domain Adaptation Building Instance Extraction of Remote Sensing Imagery With Domain-Common Approximation LearningabstractDeep learning-based building instance extraction on remote sensing imagery (RSI) has achieved tremendous success under the large-scale labeled training data. However, multi-target domain adaptation building instance extraction (MD-BIE) is still a challenge task that involves transferring knowledge from a source domain to multiple unlabeled target domains, which poses various semantic gaps between and within multiple domains,e.g., style, illumination, resolution, density, scale, etc. Most current methods for single-target domain adaptation are not applicable to the more realistic MD-BIE task. To this end, we propose a novel Domain-common Approximation Learning (DAL) for both modelling intra-domain and inter-domain adaptation, thus obtaining robust MD-BIE. DAL contains three main modules: multi-domain style transfer (MST), multi-domain feature approximation (MFA), and multi-domain cascaded instance extraction (MCIE). To alleviate the semantic gaps between multiple domains for inter-domain adaptation, we first employ the MST to learn multiple target-domain-like features that preserve both the styles of target domains and the content of the source domain, and then use the MFA to approximate these features towards a central domain-common space, thus producing domain-common semantic representations. Moreover, we develop the MCIE with hierarchical extraction losses for intra-domain adaptation to extract precise building instance contours from the domain-common semantic representations, further eliminating the potential gaps within multiple domains. By co-learning these three modules in an end-to-end manner, the DAL bridges the semantic gaps between and within multiple domains. Extensive experiments on different popular MD-BIS tasks (SAB → Crowd & WHU, Crowd → SAB & WHU, SAB → Crowd & SAB & WHU and SAB → WHU) show that our DAL outperforms the current methods by a significant margin. Fayong Zhang, Kejun Liu, Yuanyuan Liu 0004, Wujie Zhou, Hongyan Zhang 0001, Lizhe Wang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | MSTNet-KD: Multilevel Transfer Networks Using Knowledge Distillation for the Dense Prediction of Remote-Sensing ImagesabstractRecently, methods based on convolutional neural networks have achieved good results in the dense prediction of remote-sensing images, particularly when employing normalized digital surface models. However, most existing methods use multiscale convolution and attention methods to mine multimodal feature information without considering the differences and complementarities between the two features. Moreover, previous studies have prioritized model segmentation performance and ignored parametric issues, which makes it difficult to deploy the model in practical applications. To address this challenge, we designed a multilevel semantic transfer network (MSTNet) for the dense prediction of remote-sensing images using a knowledge-distillation approach to adaptively select useful semantic information for the transfer network. We designed a multilevel semantic knowledge alignment distillation framework (MSKA) to enable a compact student model to learn the semantic information extracted from a complex model. The MSKA framework comprises three main components: cross-layer semantic alignment, dynamic semantic aggregation, and softening learning for semantic information transfer and predictive label softening. Experiments on the Vaihingen and Potsdam datasets showed that the student network employing the MSKA framework achieved excellent segmentation performance with only 8.88M parameters and 2.09 gigaFLOPs in terms of computational costs compared with current state-of-the-art methods. The source code and results are available at https://github.com/LYZ00918/MSKANet. Wujie Zhou, Yangzhen Li, Juan Huan, Yuanyuan Liu 0004, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Remote Sensing Image Scene Classification via Graph Template Enhancement and Supplementation Network With Dual-Teacher Knowledge DistillationabstractImage-processing techniques used for remote sensing images (RSIs) have addressed the issues associated with scene classification. However, two challenges remain; first, most bi-modal methods integrate the extracted features indiscriminately while neglecting certain more relevant features, and bi-modal feature qualities may vary in different scenarios. Second, previous studies have improved the feature extraction performance at the expense of computational speed and complexity, hindering the wide application of improved versions. To address these issues, we constructed a graph template enhancement and supplementation network (GTESNet) with dual-teacher knowledge distillation (KD), GTESNet-S$^{\ast }$, to emphasize and manage the extracted features adaptively and compress the model. First, we developed a graph template enhancement (GTE) module that combined long-range contextual information with clustered template features. Second, an exchange-correlation fusion (ECF) module was introduced to allocate feature distribution dynamically and achieve integration. Third, we constructed feedback complementary decoders (FCDs) consisting of two subdecoders with a cascade connection that used a feedback mechanism to supplement the final output. A dual-teacher distillation method that utilized simple confidence perceptions to assign weights to both teachers was implemented during KD. Furthermore, we introduced clustered templates and graph-relation distillations (GRDs), allowing the student GTESNet-S to learn the clustering ability and semantic association process of the teacher GTESNet-T. Finally, we introduced simple progressive decoder distillation (PDD) to facilitate the learning of GTESNet-S. Extensive experiments on two publicly available datasets demonstrated that the GTESNet outperformed similar state-of-the-art (SOTA) methods. The code is available at:https://github.com/MAXHAN22/GTESNet. Wujie Zhou, Penghan Yang, Yuanyuan Liu 0004, Runmin Cong, Qiuping Jiang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Multi-View Adaptive Fusion Network for Spatially Resolved Transcriptomics Data ClusteringabstractSpatial transcriptomics technology fully leverages spatial location and gene expression information for spatial clustering tasks. However, existing spatial clustering methods primarily concentrate on utilizing the complementary features between spatial and gene expression information, while overlooking the discriminative features during the integration process. Consequently, the discriminative capability of node representation in the gene expression features is limited. Besides, most existing methods lack a flexible combination mechanism to adaptively integrate spatial and gene expression information. To this end, we propose an end-to-end deep learning method named MAFN for spatially resolved transcriptomics data clustering via a multi-view adaptive fusion network. Specifically, we first adaptively learn inter-view complementary features from spatial and gene expression information. To improve the discriminative capability of gene expression nodes by utilizing spatial information, we employ two GCN encoders to learn intra-view specific features and design a Cross-view Correlation Reduction (CCR) strategy to filter the irrelevant information. Moreover, considering the distinct characteristics of each view, a Cross-view Attention Module (CAM) is utilized to adaptively fuse the multi-view features. Extensive experimental results demonstrate that the proposed MAFN achieves competitive performance in spatial domain identification compared to other state-of-the-art ones. Yanran Zhu, Xiao He 0010, Chang Tang, Xinwang Liu 0002, Yuanyuan Liu 0004, Kunlun He |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2024 | Personalized Federated Mutual Learning for Unsupervised Camera-Aware Person Re-IdentificationabstractPerson re-identification (ReID) is essential for enhancing security and tracking in multi-camera surveillance systems. To achieve effective ReID performance across diverse datasets, the Federated Unsupervised Person Re-identification via Camera-aware Clustering (FedUCA) approach has made strides in utilizing distributed datasets while ensuring data privacy. Nevertheless, its uniform model may not adequately cater to the specific characteristics of each participant’s data, given the diversity in camera perspectives and client-specific data variances, thus obtaining degraded results. To address this issue, we propose an advanced framework, Personalized Federated Dual-Model Learning for Camera-Aware Person Re-Identification (PerFedDual), which introduces knowledge-sharing techniques inherent to mutual learning for FedUCA with the camera-centric clustering process. PerFedDual supports a dual-model training approach that creates a cooperative learning space that improves both the global model and client-specific models by exchanging knowledge both ways. The methodology adopted is precisely adjusted to the distinctive data environment of each client, ensuring the protection of privacy while simultaneously enhancing the accuracy and flexibility of ReID models in diverse camera configurations. The empirical evaluation reveals that PerFedDual outperforms FedUCA and alternative federated learning strategies, highlighting the benefits of our technique that leverages collective intelligence to enhance unsupervised ReID. Jiabei Liu, Weiming Zhuang, Yuanyuan Liu 0004, Yonggang Wen 0001, Jun Huang 0007, Wei Lin 0016 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | PCDNF: Revisiting Learning-Based Point Cloud Denoising via Joint Normal FilteringabstractPoint cloud denoising is a fundamental and challenging problem in geometry processing. Existing methods typically involve direct denoising of noisy input or filtering raw normals followed by point position updates. Recognizing the crucial relationship between point cloud denoising and normal filtering, we re-examine this problem from a multitask perspective and propose an end-to-end network called PCDNF for joint normal filtering-based point cloud denoising. We introduce an auxiliary normal filtering task to enhance the network's ability to remove noise while preserving geometric features more accurately. Our network incorporates two novel modules. First, we design a shape-aware selector to improve noise removal performance by constructing latent tangent space representations for specific points, taking into account learned point and normal features as well as geometric priors. Second, we develop a feature refinement module to fuse point and normal features, capitalizing on the strengths of point features in describing geometric details and normal features in representing geometric structures, such as sharp edges and corners. This combination overcomes the limitations of each feature type and better recovers geometric information. Extensive evaluations, comparisons, and ablation studies demonstrate that the proposed method outperforms state-of-the-art approaches in both point cloud denoising and normal filtering. Zheng Liu 0004, Yaowu Zhao, Sijing Zhan, Yuanyuan Liu 0004, Renjie Chen 0001, Ying He 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Pose-disentangled Contrastive Learning for Self-supervised Facial RepresentationabstractSelf-supervised facial representation has recently attracted increasing attention due to its ability to perform face understanding without relying on large-scale annotated datasets heavily. However, analytically, current contrastive-based self-supervised learning (SSL) still performs unsatisfactorily for learning facial representation. More specifically, existing contrastive learning (CL) tends to learn pose-invariant features that cannot depict the pose details of faces, compromising the learning performance. To conquer the above limitation of CL, we propose a novel Pose-disentangled Contrastive Learning (PCL) method for general self-supervised facial representation. Our PCL first devises a pose-disentangled decoder (PDD) with a delicately designed orthogonalizing regulation, which disentangles the pose-related features from the face-aware features; therefore, pose-related and other pose-unrelated facial information could be performed in individual subnetworks and do not affect each other's training. Furthermore, we introduce a pose-related contrastive learning scheme that learns pose-related information based on data augmentation of the same image, which would deliver more effective face-aware representation for various downstream tasks. We conducted linear evaluation on four challenging downstream facial understanding tasks, i.e., facial expression recognition, face recognition, AU detection and head pose estimation. Experimental results demonstrate that PCL significantly outperforms cuttingedge SSL methods. Our Code is available at https://github.com/DreamMr/PCL. Yuanyuan Liu 0004, Wenbin Wang 0001, Yibing Zhan, Shaoze Feng, Kejun Liu, Zhe Chen 0013 |
CVPR | 1 |
| 2023 | Learning Language-guided Adaptive Hyper-modality Representation for Multimodal Sentiment AnalysisabstractThough Multimodal Sentiment Analysis (MSA) proves effective by utilizing rich information from multiple sources (e.g., language, video, and audio), the potential sentiment-irrelevant and conflicting information across modalities may hinder the performance from being further improved.To alleviate this, we present Adaptive Language-guided Multimodal Transformer (ALMT), which incorporates an Adaptive Hyper-modality Learning (AHL) module to learn an irrelevance/conflict-suppressing representation from visual and audio features under the guidance of language features at different scales.With the obtained hypermodality representation, the model can obtain a complementary and joint representation through multimodal fusion for effective MSA.In practice, ALMT achieves state-of-the-art performance on several popular datasets (e.g., MOSI, MOSEI and CH-SIMS) and an abundance of ablation demonstrates the validity and necessity of our irrelevance/conflict suppression mechanism. Haoyu Zhang 0001, Yu Wang 0246, Guanghao Yin, Kejun Liu, Yuanyuan Liu 0004, Tianshu Yu 0001 |
EMNLP | 5 |
| 2023 | APSL: Action-positive separation learning for unsupervised temporal action localization
Yuanyuan Liu 0004, Fayong Zhang, Wenbin Wang 0001, Yu Wang 0246, Kejun Liu, Ziyuan Liu 0005 |
Inf. Sci. | 1 |
| 2023 | Secure map legends based on just noticeable distortion and watermark bit recovery
Lin Zhou 0017, Xiaohui Yuan 0001, Yuanyuan Liu 0004, Zhanlong Chen |
Multim. Tools Appl. | 3 |
| 2023 | Joint spatial and scale attention network for multi-view facial expression recognition
Yuanyuan Liu 0004, Jiyao Peng, Jiabei Zeng, Shiguang Shan |
Pattern Recognit. | 1 |
| 2023 | Expression snippet transformer for robust video-based facial expression recognition
Yuanyuan Liu 0004, Wenbin Wang 0001, Chuanxu Feng, Haoyu Zhang 0001, Zhe Chen 0013, Yibing Zhan |
Pattern Recognit. | 1 |
| 2023 | Utilizing Bounding Box Annotations for Weakly Supervised Building Extraction From Remote-Sensing ImagesabstractImage-level weakly supervised semantic segmentation (WSSS) methods have greatly facilitated the extraction of buildings from remote sensing (RS) images. However, the lack of the locations and extents of individual buildings in image-level labels results in some limitations of the methods, especially in the cases of cluttered backgrounds, diverse building shapes and sizes. By utilizing bounding box annotations, a novel WSSS model is developed to improve building extraction from RS images in this paper. Specifically, during the training phase, a multiscale feature retrieval (MFR) module is designed to learn multiscale building features and suppress the background noise inside the bounding box. In the inference phase, multiscale class activation maps (CAM) are generated from multiscale features to achieve accurate building localization. Finally, a pseudo mask generation and correction (PGC) module refines the CAMs to generate and correct the building pseudo masks. Experiments are conducted to examine the proposed model in three datasets, namely, the WHU aerial building dataset, the CrowdAI building dataset, and a self-annotated building dataset. Experimental results demonstrate that the proposed method outperforms baselines, achieving 76.99%, 75.51% and 67.35% in terms of IoU scores on the three challenging datasets, respectively. This paper provides a methodological reference for the application of weakly supervised learning on RS images. Daoyuan Zheng, Shengwen Li, Fang Fang 0008, Bo Wan 0006, Yuanyuan Liu 0004 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2023 | DSM-Assisted Unsupervised Domain Adaptive Network for Semantic Segmentation of Remote Sensing Imagery
Shunping Zhou, Shengwen Li, Daoyuan Zheng, Fang Fang 0008, Yuanyuan Liu 0004, Bo Wan 0006 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | A Dual Channel Intent Evolution Network for Predicting Period-Aware Travel Intentions at FliggyabstractFliggy of Alibaba group is one of the largest online travel platform (OTPs) in China, which provides travel products and travel experiences for tens of millions of online users by the personalized recommendation system (RS). User's future travel intent prediction is one key problem in travel scenario, which decides where and what to recommend, e.g., traveling to a surrounding city or a distant city. Such travel intent prediction problem has a lot of important applications, e.g., to push a notification with surrounding scenic spots recommendation to a user with intent to travel around, or to enable personalized promotion strategies to users with different intents. Existing studies on user's intent are largely sub-optimal for users' travel intent prediction at OTPs, since they rarely pay attentions to the characteristics of the travel industry, namely, user behavior sparsity due to low frequency of travel, spatial-temporal periodicity patterns, and the correlations between user's online and offline behaviors. In this paper, to address these challenges, we propose a dual channel intent evolution network based online-offline periodicity-aware network, DCIEN, for user's future travel intent prediction. In particular, it consists of two basic components including 1) Spatial-temporal Intent Patterns Network(ST-IPN), which exploits users' periodic intent patterns from offline data based on convolutional neural networks; 2) Periodicity-aware Intent Evolution Network(PA-IEN), which captures user's instant intent from online behaviors data and the interactions between online and offline intents. Extensive offline and online experiments on a real-world OTP demonstrate the superior performance of DCIEN over state-of-the-art methods. Wanjie Tao, Zhang-Hua Fu, Liangyue Li, Zulong Chen, Hong Wen 0002, Yuanyuan Liu 0004, Qijie Shen |
CIKM | 6 |
| 2022 | MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the WildabstractDynamic facial expression recognition (FER) databases provide important data support for affective computing and applications. However, most FER databases are annotated with several basic mutually exclusive emotional categories and contain only one modality, e.g., videos. The monotonous labels and modality cannot accurately imitate human emotions and fulfill applications in the real world. In this paper, we propose MAFW, a large-scale multi-modal compound affective database with 10,045 video-audio clips in the wild. Each clip is annotated with a compound emotional category and a couple of sentences that describe the subjects' affective behaviors in the clip. For the compound emotion annotation, each clip is categorized into one or more of the 11 widely-used emotions, i.e., anger, disgust, fear, happiness, neutral, sadness, surprise, contempt, anxiety, helplessness, and disappointment. To ensure high quality of the labels, we filter out the unreliable annotations by an Expectation Maximization (EM) algorithm, and then obtain 11 single-label emotion categories and 32 multi-label emotion categories. To the best of our knowledge, MAFW is the first in-the-wild multi-modal database annotated with compound emotion annotations and emotion-related captions. Additionally, we also propose a novel Transformer-based expression snippet feature learning method to recognize the compound emotions leveraging the expression-change relations among different emotions and modalities. Extensive experiments on MAFW database show the advantages of the proposed method over other state-of-the-art methods for both uni- and multi-modal FER. Our MAFW database is publicly available from https://mafw-database.github.io/MAFW. Yuanyuan Liu 0004, Chuanxu Feng, Wenbin Wang 0001, Guanghao Yin, Jiabei Zeng, Shiguang Shan |
ACM Multimedia | 1 |
| 2022 | Clip-aware expressive feature learning for video-based facial expression recognition
Yuanyuan Liu 0004, Chuanxu Feng, Xiaohui Yuan 0001, Lin Zhou 0017, Wenbin Wang 0001, Zhongwen Luo |
Inf. Sci. | 1 |
| 2022 | ConGNN: Context-consistent cross-graph neural network for group emotion recognition in the wild
Yu Wang 0246, Shunping Zhou, Yuanyuan Liu 0004, Fang Fang 0008, Haoyue Qian |
Inf. Sci. | 3 |
| 2021 | Hierarchical Domain-Consistent Network For Cross-Domain Object DetectionabstractCross-domain object detection is a very challenging task due to multi-level domain shift in an unseen domain. To address the problem, this paper proposes a hierarchical domain-consistent network (HDCN) for cross-domain object detection, which effectively suppresses pixel-level, image-level, as well as instance-level domain shift via jointly aligning three-level features. Firstly, at the pixel-level feature alignment stage, a pixel-level subnet with foreground-aware attention learning and pixel-level adversarial learning is proposed to focus on local foreground transferable information. Then, at the image-level feature alignment stage, global domain-invariant features are learned from the whole image through image-level adversarial learning. Finally, at the instance-level alignment stage, a prototype graph convolution network is conducted to guarantee distribution alignment of instances by minimizing the distance of prototypes with the same category but from different domains. Moreover, to avoid the non-convergence problem during multi-level feature alignment, a domain-consistent loss is proposed to harmonize the adaptation training process. Comprehensive results on various cross-domain detection tasks demonstrate the broad applicability and effectiveness of the proposed approach. Yuanyuan Liu 0004, Fang Fang 0008, Zhanghua Fu, Zhanlong Chen |
ICIP | 1 |
| 2021 | Synthesizing location semantics from street view images to improve urban land-use classificationabstractLand-use maps are instrumental to inform urban planning and environmental research. Street view images (SVIs) have shown great potential for automated land-use classification for land-use mapping. However, previous studies overlooked SVI-derived location contextual information that may help improve land-use classification. This study proposes a novel land-use classification method that synthesizes location semantics from SVIs to account for contextual information from SVIs, land parcels and roads around the SVIs. The proposed method first generates land-use scene images (LUSIs) by using an SVI-derived straightforward algorithm. The LUSIs are then relocated to land parcels by using a displacement strategy and classified into land-use types by using a deep learning network. This study determines the land-use types of land parcels with classified LUSIs. Two case studies, consisting of LUSIs for five land-use types, show that introducing location semantics of SVIs can remarkably improve the classification accuracy of land-use types. Fang Fang 0008, Yafang Yu, Shengwen Li, Zejun Zuo, Yuanyuan Liu 0004, Bo Wan 0006, Zhongwen Luo |
Int. J. Geogr. Inf. Sci. | 5 |
| 2021 | TSingNet: Scale-aware and context-rich feature learning for traffic sign detection and recognition in the wild
Yuanyuan Liu 0004, Jiyao Peng, Jing-Hao Xue, Yongquan Chen, Zhang-Hua Fu |
Neurocomputing | 1 |
| 2021 | Dynamic multi-channel metric network for joint pose-aware and identity-invariant facial expression recognition
Yuanyuan Liu 0004, Fang Fang 0008, Yongquan Chen, Rui Huang 0001, Run Wang 0002, Bo Wan 0006 |
Inf. Sci. | 1 |
| 2021 | Multiscale U-Shaped CNN Building Instance Extraction Framework With Edge Constraint for High-Spatial-Resolution Remote Sensing ImageryabstractBuilding extraction based on high-resolution remote sensing imagery has been widely used in automatic surveying and mapping. However, few methods have been developed for building instance extraction, i.e., extracting each building's footprint separately, which is required in a number of applications, such as the smallest unit of a cadastral database. In building instance extraction, there are two challenges: 1) buildings with various scales exist in the imagery and 2) precise building footprints are difficult to extract due to the blurry boundaries. In this article, to solve these problems, a multiscale U-shaped convolutional neural network building instance extraction framework with edge constraint (EMU-CNN) for high-spatial-resolution remote sensing imagery is proposed. The proposed framework consists of three components: 1) a multiscale fusion U-shaped network (MFUN); 2) a region proposal network (RPN); and 3) an edge-constrained multitask network (ECMN). First, in the proposed method, the MFUN includes three parallel branches to learn multiple building features with different scales. The RPN then detects the positions of the building instances, even for buildings that are connected with each other. Moreover, according to the instance positions, the ECMN is proposed to extract a precise mask and suppress overfitting. The experiments conducted on a self-annotated data set and two public data sets (the ISPRS Vaihingen semantic labeling contest data set and the WHU aerial image data set) show that the EMU-CNN method can achieve excellent performance and shows great robustness at different scales. Yuanyuan Liu 0004, Ailong Ma, Yanfei Zhong, Fang Fang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2019 | Visual Focus of Attention and Spontaneous Smile Recognition Based on Continuous Head Pose Estimation by Cascaded Multi-Task LearningabstractMulti-person Visual focus of attention (M-VFOA) and spontaneous smile (SS) recognition are important for persons’ behavior understanding and analysis in class. Recently, promising results have been reported using special hardware in constrained environment. However, M-VFOA and SS remain challenging problems in natural and crowd classroom environment, e.g. various poses, occlusion, expressions, illumination and poor image quality, etc. In this study, a robust and un-invasive M-VFOA and SS recognition system has been developed based on continuous head pose estimation in the natural classroom. A novel cascaded multi-task Hough forest (CM-HF) combined with weighted Hough voting and multi-task learning is proposed for continuous head pose estimation, tip of the nose location and SS recognition, which improves accuracies of recognition and reduces the training time. Then, M-VFOA can be recognized based on estimated head poses, environmental cues and prior states in the natural classroom. Meanwhile, SS is classified using CM-HF with local cascaded mouth-eyes areas normalized by the estimated head poses. The method is rigorously evaluated for continuous head pose estimation, multi-person VFOA recognition, and SS recognition on some public available datasets and real-class video sequences. Experimental results show that our method reduces training time greatly and outperforms the state-of-the-art methods for both performance and robustness with an average accuracy of 83.5% on head pose estimation, 67.8% on M-VFOA recognition and 97.1% on SS recognition in challenging environments. Yuanyuan Liu 0004, Xingmei Li, Fang Fang 0008, Fayong Zhang, Jingying Chen 0001, Zhizhong Zeng |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2018 | Multi-Channel Pose-Aware Convolution Neural Networks for Multi-View Facial Expression RecognitionabstractAlthough tremendous strides have been made in facial expression recognition(FER), recognizing facial expressions in non-frontal views remains an open challenge due to the limited access to large scale training data with various poses. To make full use of the limited data, we propose a novel multi-channel pose-aware convolution neural network (MPCNN) that consists of three parts: the multi-channel feature extraction, jointly multi-scale feature fusion, and the pose-aware recognition. The feature extraction part has 3 sub-CNNs and it learns convolutional features from different features. The joint fusion part fuses multi-scale features to enhance high-level feature representation in a hierarchical way. The fused features are fed to the pose-aware recognition part that includes pose-specific recognition branches and a pose estimation sub-network. According to the estimated pose, MPCNN finally classifies the facial expression through a conditional weighted combination of the pose-specific recognition branches. MPCNN is end-to-end trainable by minimizing the joint loss of pose and expression recognition. We evaluated the proposed method on two public multi-view FER datasets (BU-3DFE and KDEF) and a FER dataset in the wild (SFEW). The experimental results demonstrate that MPCNN outperforms the state-of-the-art FER methods with both within-dataset and cross-dataset settings. Yuanyuan Liu 0004, Jiabei Zeng, Shiguang Shan, Zhuo Zheng |
FG | 1 |
| 2018 | Urban Land-Use Classification From PhotographsabstractLand-use (LU) classification of urban areas is conventionally achieved via field survey or remote sensing technologies, which is labor-intensive and time-consuming. With the wide development of social networks such as microblog and ubiquitous network access, images are captured by residents and tourists. In this letter, we propose a method for an automatic urban LU classification using geotagged images from public venues. Our method identifies the LU type depicted in those images that are extrapolated to the local regions bounded by street blocks. Experiments were conducted with geotagged photographs and Open Street Map of an urban area in London, U.K. It was demonstrated that the proposed method achieved overall 76.5% accuracy across five LU types. More importantly, our method demonstrated a greater performance in dealing with a mixture of LU types. Fang Fang 0008, Xiaohui Yuan 0001, Yuanyuan Liu 0004, Zhongwen Luo |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2018 | Student engagement study based on multi-cue detection and recognition in an intelligent learning environment
Yuanyuan Liu 0004, Jingying Chen 0001, Mulan Zhang, Chuan Rao |
Multim. Tools Appl. | 1 |
| 2018 | Conditional convolution neural network enhanced random forest for facial expression recognition
Yuanyuan Liu 0004, Xiaohui Yuan 0001, Xi Gong, Zhong Xie, Fang Fang 0008, Zhongwen Luo |
Pattern Recognit. | 1 |
| 2017 | Urban function zoning using geotagged photos and openstreetmapabstractUrban function zoning is of great importance for urban structure optimization, urban resource allocation, and urban development planning. Since citizens usually act as a network of motion sensors of the city, their activities could reflect the environment around them. We considered taking advantage of VGI data to classify urban function zones. In this paper, we proposed a framework for automated urban function zoning which is based on VGI geo-tagged photos and OpenStreetMap (OSM) data. Through combining the high-level image features of geo-tagged photos with the road network data, we obtained the functional zoning map of the study area. The experiment result shows the effectiveness of the framework we proposed. Fang Fang 0008, Xiaohui Yuan 0001, Zhongwen Luo, Yuanyuan Liu 0004, Bo Wan 0006, Yishi Zhao |
IGARSS | 5 |
| 2017 | A Performance Evaluation Model for Taxi Cruising Path Recommendation System
Huimin Lv, Fang Fang 0008, Yishi Zhao, Yuanyuan Liu 0004, Zhongwen Luo |
PAKDD (2) | 4 |
| 2017 | Multi-level structured hybrid forest for joint head detection and pose estimation
Yuanyuan Liu 0004, Zhong Xie, Xiaohui Yuan 0001, Jingying Chen 0001, Wu Song |
Neurocomputing | 1 |
| 2016 | Multi-person Visual Focus of Attention from Head Pose on a Natural Classroom
Yuanyuan Liu 0004, Leyuan Liu 0001, Jingying Chen 0001, Chunyan Su, Kun Zhang 0031 |
ICPRAM | 1 |
| 2016 | Robust head pose estimation using Dirichlet-tree distribution enhanced random forests
Yuanyuan Liu 0004, Jingying Chen 0001, Zhiming Su, Zhenzhen Luo, Nan Luo, Leyuan Liu 0001, Kun Zhang 0031 |
Neurocomputing | 1 |
| 2014 | Dirichlet-tree Distribution Enhanced Random Forests for Head Pose EstimationabstractHead pose estimation is important in human-machine interfaces. However, illumination variation, occlusion and low image resolution make the estimation task difficult. Hence, a Dirichlet-tree distribution enhanced Random Forests approach (D-RF) is proposed in this paper to estimate head pose efficiently and robustly under various conditions. First, PCA based sub-features space from Gabor features and histogram distributions of the facial patches are extracted to eliminate the influence of occlusion and noise. Then, the D-RF is proposed to estimate the head pose in a coarse-to-fine way. In order to improve the discrimination capability of the approach, an adaptive Gaussian mixture model is introduced in the tree distribution. The proposed method has been evaluated with different data sets spanning from -90° to 90° in vertical and horizontal directions under various conditions. The experimental results demonstrate the approachâs robustness and efficiency. Yuanyuan Liu 0004, Jingying Chen 0001, Leyuan Liu 0001, Yujiao Gong, Nan Luo |
ICPRAM | 1 |