VLDB 2026 Research / reviewers in the wild / expert
Zhen Cui 0001
dblp:59/8491-1
· DBLP profile ↗
165ranked-venue papers
15as first author
96since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 102 · 11 first-author · 58 since 2021Graphics, computer vision, multimedia, augmented reality and games · 92 · 9 first-author · 52 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 10 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Computer networks · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From Noisy Candidates to Reliable Grounding in Weakly-Supervised Referring Expression ComprehensionabstractWeakly-Supervised Referring Expression Comprehension (WREC) aims to ground natural language expressions in image regions, using only image-text pairs without bounding-box annotations. However, the absence of explicit localization supervision results in noisy training signals and unreliable candidate selection. We observe that modern query-based detectors exhibit a high-recall, low-precision behavior in WREC: although the top-ranked query often localizes inaccurately, the ground-truth region is frequently recalled among high-confidence candidates. Motivated by this observation, we propose a Noisy-to-Reliable Grounding (NRG) framework that progressively transforms noisy high-recall candidates into reliable grounding supervision. Given an paired image-text data, we leverage a large Vision-Language Model (VLM) to generate confidence-aware soft pseudo labels, which provide robust semantic and spatial guidance under weak supervision. To disentangle true positives from noisy candidates, we introduce a contrastive candidate mining module that jointly exploits the detector’s prediction and VLM-derived cues to identify positive and hard negative queries, progressively enhancing grounding discriminability. Furthermore, the confidence-aware pseudo labels are integrated to construct a regression objective, enabling effective localization learning through low-rank adaptation of a pretrained open-vocabulary detector. Extensive experiments on RefCOCO, RefCOCO+, and RefCOCOg datasets demonstrate that the proposed NRG consistently outperforms existing WREC methods. Ziqi Gu, Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu |
ICMR | 4 |
| 2026 | Progressive subgoal-aggregated long-sequence decision-making with large language models
Yiliang Liu, Zhen Cui 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | An efficient dominance decomposition-based deep graph evolutionary algorithm for the expensive multi-objective optimization
Xing Cai, Tong Zhang 0021, Zhen Cui 0001 |
Expert Syst. Appl. | 3 |
| 2026 | Minimum Description Length-Driven Fragment Mining for Pretraining Molecule Property Prediction ModelabstractMolecular fragments play a crucial role in molecular property prediction. However, most existing deep learning approaches rely heavily on expert-defined substructural patterns, limiting their ability to identify novel or latent fragments. This constraint reduces the generalizability and applicability of molecular fragments in molecular representation learning. In this study, we propose the Molecular Multi-view Pre-training model with Adaptive Fragment Mining (MMP-AFM), a unified framework that facilitates the seamless integration of molecular structural information. MMP-AFM formulates fragment discovery as a combinatorial optimization problem, using description length as the objective to enable adaptive extraction of molecular fragments and dynamic construction of a fragment library. Additionally, we design a molecular multi-view self-supervised pretraining framework that aligns features from the fragment, global, and data augmentation views, ensuring a comprehensive integration of molecular substructural information. Finally, the MMP-AFM is applied to both molecular classification and regression tasks. Experimental results demonstrate that MMP-AFM consistently outperforms existing methods across multiple tasks, highlighting its broad applicability and efficiency. Xing Cai, Tong Zhang 0021, Yide Qiu, Baotong Su, Zhen Cui 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 6 |
| 2026 | Training-Free Controllable Text-Guided Video EditingabstractDecomposition-based text-guided video editing paradigm aims to utilize the layered neural atlas model to decompose the input video into foreground and background parts and edit the video in a divide-and-conquer manner, which is meaningful and improves the controllability of editing. However, they may suffer from some limitations: 1) High computational cost of per-video training (i.e, 7∼8 hours for training a single atlas model). 2) Foreground object deformation is restricted by the foreground opacity value. 3) Restricted flexibility in manipulating multiple objects. In this paper, we propose TraFrCo, aTraining-Free Controllable Text-guided Video Editingframework to mitigate these challenges. Instead of training complex atlas models, our method leverages pre-trained segmentation to rapidly decompose videos into foreground and background parts. This allows users to perform independent edits on foreground objects using existing video diffusion editing models without affecting the environment. To ensure visual consistency, we introduce a training-free mechanism that effectively propagates information across frames to fill missing background regions caused by the segmentation-derived foreground masks and reconstructs the scene behind moving objects. Finally, the edited components are seamlessly composited by re-predicting the new foreground masks. In contrast to prior works, TraFrCo enables efficient, fine-grained manipulation of video content without the burden of training. Experimental results verify that our TraFrCo consistently reduces the costs of decomposing video and achieves superior text-guided video editing performance. Codes and video demos will be released at https://github.com/mdswyz/TraFrCo. Yuanzhi Wang, Yong Li 0032, Zhen Cui 0001, Jian Yang 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Boosting Few-Shot Continual Learning via Self-Adaptive EvolutionabstractFew-shot continual learning (FSCL) has attracted increasing attention for real-world applications, where models must continuously adapt to new classes with only a few labeled samples while retaining prior knowledge. These abilities are essential in dynamic environments where data availability is often sparse and nonstationary. However, traditional FSCL methods are largely confined to closed data spaces, which limits their generalizability when diverse and evolving distributions are involved. Inspired by the paradigm of human lifelong learning, we propose a new self-adaptive evolution framework for FSCL that enables continuous interaction with and adaptation to external environments. To exploit latent knowledge in large-scale models, we use an adaptive diffusion-based generator that not only implicitly captures the distribution of new few-shot samples but also produces more high-quality samples. To mitigate the inevitable variability in generation quality, we also use a reinforced sample selection module, comprising a generated sample explorer and a selection evaluator, which explicitly guides the retained distributions toward alignment with the large-scale models. Integrated with the continual model, these components are optimized in an iterative self-adaptive evolution framework, ensuring stable knowledge retention while improving adaptability to newly emerging classes. We validate our approach through experiments on three benchmarks, revealing its effectiveness in exploiting external distributions and achieving notable performance improvements. Ziqi Gu, Chunyan Xu, Yuanzhi Wang, Cao Han, Di Xia, Zhen Cui 0001 |
IEEE Trans. Image Process. | 7 |
| 2026 | ADAN: Adversarial Distribution Alignment Network for Multi-View Semi-Supervised ClassificationabstractMulti-view learning aims to integrate multi-source information for a comprehensive data representation, which has gained widespread attention in image processing. Each view contains view-specific noise and joint features associated with other views, and thus exploring the specificity and consistency among views is a typical solution to deal with multi-view data for learning discriminative representations. In this paper, we present a theory-induced model, termed Adversarial Distribution Alignment Network (ADAN), which learns view-invariant features and alleviate the negative impact of view-specific noise. We first demonstrate the necessity of suppressing view-specific noise and capturing view-invariant features inspired by the theory of view generalization, and then derive two collaborative modules: a feature disentangler and an adversarial alignment module. In detail, the feature disentanglement separates view-specific noise and view-invariant features by minimizing the mutual information between them. Following this, a negative entropy is proposed to suppress the negative impact of view-specific noise. Meanwhile, the adversarial module uses the adversarial technique that can fit more complex data conformed to different distributions to adaptively align cross-view features so that features encoded in different views converge. Substantial experiments are constructed on multi-view datasets, demonstrating that ADAN can achieve more promising performance compared to other superior methods. Code is available at https://github.com/huangsuj/ADANet. Sujia Huang, Lele Fu, Zhaoliang Chen, Tong Zhang 0021, Xiaoli Li 0016, Zhen Cui 0001 |
IEEE Trans. Image Process. | 6 |
| 2026 | Graph Probabilistic Pooling: From Bernoulli to Poisson DistributionabstractGraph pooling is crucial for enlarging the receptive field and reducing computational costs in deep graph representation learning. In this work, we propose a simple but effective graph probabilistic pooling (GP-Pool) framework to facilitate graph feature learning. Instead of either deterministic selection or random dropping, we design a probabilistic subgraph sampling to reach an expected distribution by deducing a variational bound. Accordingly, a Bernoulli graph pooling (BernPool) is first derived to sample nodes together with the local structures, for which a learnable reference set is introduced to encode nodes into a latent expressive probability space. Hereby, the resultant BernPool captures salient graph substructures while possessing much diversity on sampled nodes due to its nondeterministic manner. For more controllable pooling, we derive the Poisson-distributed version (aka PoissonPool) from BernPool to explicitly cut the node quantity with less variables in variational learning. Furthermore, considering the complementarity of node sampling and clustering, we propose a hybrid graph pooling (HGP) paradigm to combine a compact subgraph (via BernPool/PoissonPool) and a coarsening graph (via clustering), to retain both representative substructures and global topology. Extensive experiments on multiple public graph classification datasets demonstrate that our GP-Pool is superior to various graph pooling methods and achieves state-of-the-art performance. Guangbu Liu, Tong Zhang 0021, Chuanwei Zhou, Cheng Long 0001, Zhen Cui 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | Re-Attentional Controllable Video Diffusion EditingabstractEditing videos with textual guidance has garnered popularity due to its streamlined process which mandates users to solely edit the text prompt corresponding to the source video. Recent studies have explored and exploited large-scale text-to-image diffusion models for text-guided video editing, resulting in remarkable video editing capabilities. However, they may still suffer from some limitations such as mislocated objects, incorrect number of objects. Therefore, the controllability of video editing remains a formidable challenge. In this paper, we aim to challenge the above limitations by proposing a Re-Attentional Controllable Video Diffusion Editing (ReAtCo) method. Specially, to align the spatial placement of the target objects with the edited text prompt in a training-free manner, we propose a Re-Attentional Diffusion (RAD) to refocus the cross-attention activation responses between the edited text prompt and the target video during the denoising stage, resulting in a spatially location-aligned and semantically high-fidelity manipulated video. In particular, to faithfully preserve the invariant region content with less border artifacts, we propose an Invariant Region-guided Joint Sampling (IRJS) strategy to mitigate the intrinsic sampling errors w.r.t the invariant regions at each denoising timestep and constrain the generated content to be harmonized with the invariant region content. Experimental results verify that ReAtCo consistently improves the controllability of video diffusion editing and achieves superior video editing performance. Yuanzhi Wang, Yong Li 0032, Xin Liu 0011, Zhen Cui 0001, Antoni B. Chan |
AAAI | 6 |
| 2025 | Multi-clue Consistency Learning to Bridge Gaps Between General and Oriented Object in Semi-supervised DetectionabstractWhile existing semi-supervised object detection (SSOD) methods perform well in general scenes, they encounter challenges in handling oriented objects in aerial images. We experimentally find three gaps between general and oriented object detection in semi-supervised learning: 1) Sampling inconsistency: the common center sampling is not suitable for oriented objects with larger aspect ratios when selecting positive labels from labeled data. 2) Assignment inconsistency: balancing the precision and localization quality of oriented pseudo-boxes poses greater challenges which introduces more noise when selecting positive labels from unlabeled data. 3) Confidence inconsistency: there exists more mismatch between the predicted classification and localization qualities when considering oriented objects, affecting the selection of pseudo-labels. Therefore, we propose a Multi-clue Consistency Learning (MCL) framework to bridge gaps between general and oriented objects in semi-supervised detection. Specifically, considering various shapes of rotated objects, the Gaussian Center Assignment is specially designed to select the pixel-level positive labels from labeled data. We then introduce the Scale-aware Label Assignment to select pixel-level pseudo-labels instead of unreliable pseudo-boxes, which is a divide-and-rule strategy suited for objects with various scales. The Consistent Confidence Soft Label is adopted to further boost the detector by maintaining the alignment of the predicted results. Comprehensive experiments on DOTA-v1.5 and DOTA-v1.0 benchmarks demonstrate that our proposed MCL can achieve state-of-the-art performance in the semi-supervised oriented object detection task. Chunyan Xu, Xiang Li 0041, YuXuan Li, Ziqi Gu, Zhen Cui 0001 |
AAAI | 7 |
| 2025 | Scene Graph-Grounded Image GenerationabstractWith the beneft of explicit object-oriented reasoning capabilities of scene graphs, scene graph-to-image generation has made remarkable advancements in comprehending object coherence and interactive relations. Recent state-of-the-arts typically predict the scene layouts as an intermediate representation of a scene graph before synthesizing the image. Nevertheless, transforming a scene graph into an exact layout may restrict its representation capabilities, leading to discrepancies in interactive relationships (such as standing on, wearing, or covering) between the generated image and the input scene graph. In this paper, we propose a Scene Graph-Grounded Image Generation (SGG-IG) method to mitigate the above issues. Specifcally, to enhance the scene graph representation, we design a masked auto-encoder module and a relation embedding learning module to integrate structural knowledge and contextual information of the scene graph with a mask self-supervised manner. Subsequently, to bridge the scene graph with visual content, we introduce a spatial constraint and image-scene alignment constraint to capture the fne-grained visual correlation between the scene graph symbol representation and the corresponding image representation, thereby generating semantically consistent and high-quality images. Extensive experiments demonstrate the effectiveness of the method both quantitatively and qualitatively. Fuyun Wang, Tong Zhang 0021, Yuanzhi Wang, Xin Liu 0011, Zhen Cui 0001 |
AAAI | 6 |
| 2025 | Distribution Prototype Diffusion Learning for Open-set Supervised Anomaly DetectionabstractIn Open-set Supervised Anomaly Detection (OSAD), the existing methods typically generate pseudo anomalies to compensate for the scarcity of observed anomaly samples, while overlooking critical priors of normal samples, leading to less effective discriminative boundaries. To address this issue, we propose a Distribution Prototype Diffusion Learning (DPDL) method aimed at enclosing normal samples within a compact and discriminative distribution space. Specifically, we construct multiple learnable Gaussian prototypes to create a latent representation space for abundant and diverse normal samples and learn a Schrödinger bridge to facilitate a diffusive transition toward these prototypes for normal samples while steering anomaly samples away. Moreover, to enhance inter-sample separation, we design a dispersion feature learning way in hyper-spherical space, which benefits the identification of out-of-distribution anomalies. Experimental results demonstrate the effectiveness and superiority of our proposed DPDL, achieving state-of-the-art performance on 9 public datasets. Fuyun Wang, Tong Zhang 0021, Yuanzhi Wang, Yide Qiu, Xin Liu 0011, Zhen Cui 0001 |
CVPR | 7 |
| 2025 | STDD: Spatio-Temporal Dual Diffusion for Video GenerationabstractDiffusion probabilistic model is becoming the cornerstone of data generation, especially generating high-quality images. As an extension, video diffusion generation is in urgent need of a principled temporal-sequence diffusion way, while the spatial-domain diffusion dominates most video diffusion methods. In this work, we propose an explicit Spatio-Temporal Dual Diffusion (STDD) method by principledly extending the standard diffusion model to a spatio-temporal diffusion model for joint spatial and temporal noise propagation/reduction. Mathematically, an analysable dual diffusion process is derived to accumulate information in temporal sequence as well as spatial domain. Correspondingly, we theoretically derive a spatio-temporal probabilistic reverse diffusion process and propose an accelerated sampling way to reduce the inference cost. In principle, the spatio-temporal dual diffusion enables the information of previous frames to be transferred to the current frame, which thus could be beneficial for video consistency. Extensive experiments demonstrate that our proposed STDD is more competitive over the state-of-the-art methods in the task of video generation/prediction as well as text-to-video generation. Shuaizhen Yao, Xin Liu 0011, Zhen Cui 0001 |
CVPR | 5 |
| 2025 | M3Rec: Selective State Space Models with Mixture-of-Modality Experts for Multi-Modal Sequential RecommendationabstractThe rapid growth of multimedia-sharing platforms drives the development of recommender systems. While traditional ID-based methods for mining user behavior signals are well-studied, research into multimodal sequential recommendation remains nascent. Current approaches face three critical challenges: (1) inadequate modeling of user preferences across diverse modalities, (2) ineffective capture of user action sequence dependencies hinders representation learning of preferences, and (3) inefficiency in Transformer-based models due to the quadratic complexity of attention mechanisms. To address these issues, we propose M3Rec, a Mamba-based selective state space model incorporating Mixture-of-Modality experts for Multimodal sequential recommendation. M3Rec strengthens the modeling of user action sequence dependencies through shared Mamba blocks across modalities and employs modality experts to extract modality-specific user preferences. The shared Mamba blocks efficiently model long-term user preferences with fast inference and linear scalability through hardware-aware parallelism, enhancing ID-based sequence signals and filtering out non-action-dependent redundant information. This enables more accurate modeling of user preferences across heterogeneous data. Extensive experiments on three public datasets validate the model’s effectiveness. The implementation is released at https://github.com/Xu107/M3Rec-main. Tong Zhang 0021, Fuyun Wang, Zhen Cui 0001 |
ICASSP | 6 |
| 2025 | LLM-Assisted Semantic Guidance for Sparsely Annotated Remote Sensing Object DetectionabstractSparse annotation in remote sensing object detection poses significant challenges due to dense object distributions and category imbalances. Although existing Dense Pseudo-Label methods have demonstrated substantial potential in pseudo-labeling tasks, they remain constrained by selection ambiguities and inconsistencies in confidence estimation.In this paper, we introduce an LLM-assisted semantic guidance framework tailored for sparsely annotated remote sensing object detection, exploiting the advanced semantic reasoning capabilities of large language models (LLMs) to distill high-confidence pseudo-labels.By integrating LLM-generated semantic priors, we propose a Class-Aware Dense Pseudo-Label Assignment mechanism that adaptively assigns pseudo-labels for both unlabeled and sparsely labeled data, ensuring robust supervision across varying data distributions. Additionally, we develop an Adaptive Hard-Negative Reweighting Module to stabilize the supervised learning branch by mitigating the influence of confounding background information. Extensive experiments on DOTA and HRSC2016 demonstrate that the proposed method outperforms existing single-stage detector-based frameworks, significantly improving detection performance under sparse annotations. Chunyan Xu, Zhen Cui 0001 |
ICCV | 4 |
| 2025 | Semantic Discrepancy-Aware Detector for Image Forgery IdentificationabstractWith the rapid advancement of image generation techniques, robust forgery detection has become increasingly imperative to ensure the trustworthiness of digital media. Recent research indicates that the learned semantic concepts of pre-trained models are critical for identifying fake images. However, the misalignment between the forgery and semantic concept spaces hinders the model's forgery detection performance. To address this problem, we propose a novel Semantic Discrepancy-aware Detector (SDD) that leverages reconstruction learning to align the two spaces at a fine-grained visual level. By exploiting the conceptual knowledge embedded in the pre-trained vision language model, we specifically design a semantic token sampling module to mitigate the space shifts caused by features irrelevant to both forgery traces and semantic concepts. A concept-level forgery discrepancy learning module, built upon a visual reconstruction paradigm, is proposed to strengthen the interaction between visual semantic concepts and forgery traces, effectively capturing discrepancies under the concepts' guidance. Finally, the low-level forgery feature enhancemer integrates the learned concept level forgery discrepancies to minimize redundant forgery information. Experiments conducted on two standard image forgery datasets demonstrate the efficacy of the proposed SDD, which achieves superior results compared to existing methods. The code is available at https://github.com/wzy1111111/SSD. Minghang Yu, Chunyan Xu, Zhen Cui 0001 |
ICCV | 4 |
| 2025 | CLIP-driven Few-Shot Continual LearningabstractFew-Shot continual learning (FSCL) has garnered significant attention and has made notable progress in recent years. Drawing inspiration from the continuous interaction of humans with their environment in the lifelonglearning [1], we propose a novel CLIP-driven Few-Shot Continual Learning (C-FSCL) framework that aims to progressively enhance the continual network by leveraging the capabilities of a foundational CLIP model. When the continual model encounters new-coming data, we introduce a hierarchical feature-aware alignment module to better generalize knowledge gained from the foundational CLIP model, allowing it to apply learned concepts to new tasks more effectively. Recognizing the instability of inter-class structural relationships, which leads to catastrophic forgetting, we further introduce a relation-aware structure alignment module to align the consistent relationships from the CLIP model, thereby enhancing the reliability of the continual model’s knowledge structure. Experiments demonstrate the feasibility of CLIP-driven FSCL task and its remarkable performance on three benchmarks. Ziqi Gu, Chunyan Xu, Zhen Cui 0001 |
ICME | 3 |
| 2025 | Going Beyond Consistency: Target-oriented Multi-view Graph Neural NetworkabstractMulti‐view learning has emerged as a pivotal research area driven by the growing heterogeneity of real‐world data, and graph neural network-based models, modeling multi-view data as multi-view graphs, have achieved remarkable performance by revealing its deep semantics. However, by assuming cross‐view consistency, most approaches collect not only task-relevant (determinative) semantics but also symbiotic yet task-irrelevant (incidental) factors are collected to obscure model inference. Furthermore, these approaches often lack rigorous theoretical analysis that bridges training data to test data. To address these issues, we propose Target-oriented Graph Neural Network (TGNN), a novel framework that goes beyond traditional consistency by prioritizing task-relevant information, ensuring alignment with the target. Specifically, TGNN employs a class-level dual-objective loss to minimize the classification similarity between determinative and incidental factors, accentuating the former while suppressing the latter during model inference. Meanwhile, to ensure consistency between the learned semantics and predictions in representation learning, we introduce a penalty term that aims to amplify the divergence between these two types of factors. Furthermore, we derive an upper bound on the loss discrepancy between training and test data, providing formal guarantees for generalization to test domains. Extensive experiments conducted on three types of multi-view datasets validate the superiority of TGNN. Sujia Huang, Lele Fu, Shuman Zhuang, Yide Qiu, Zhen Cui 0001, Tong Zhang 0021 |
IJCAI | 6 |
| 2025 | Learn and Ensemble Bridge Adapters for Multi-domain Task Incremental LearningabstractMulti-domain task incremental learning (MTIL) demands models to master domain-specific expertise while preserving generalization capabilities.
Inspired by human lifelong learning, which relies on revisiting, aligning, and integrating past experiences, we propose a Learning and Ensembling Bridge Adapters (LEBA) framework.
To facilitate cohesive knowledge transfer across domains, specifically, we propose a continuous-domain bridge adaptation module, leveraging the distribution transfer capabilities of Schrödinger bridge for stable progressive learning.
To strengthen memory consolidation, we further propose a progressive knowledge ensemble strategy that revisits past task representations via a diffusion model and dynamically integrates historical adapters.
For efficiency, LEBA maintains a compact adapter pool through similarity-based selection and employs learnable weights to align replayed samples with current task semantics.
Together, these components effectively mitigate catastrophic forgetting and enhance generalization across tasks.
Extensive experiments across multiple benchmarks validate the effectiveness and superiority of LEBA over state-of-the-art methods. Ziqi Gu, Chunyan Xu, Xin Liu 0011, Yide Qiu, Zhen Cui 0001 |
NeurIPS | 6 |
| 2025 | Value Diffusion Reinforcement LearningabstractModel-free reinforcement learning (RL) combined with diffusion models has achieved significant progress in addressing complex continuous control tasks. However, a persistent challenge in RL remains the accurate estimation of Q-values, which critically governs the efficacy of policy optimization. Although recent advances employ parametric distributions to model value distributions for enhanced estimation accuracy, current methodologies predominantly rely on unimodal Gaussian assumptions or quantile representations. These constraints introduce distributional bias between the learned and true value distributions, particularly in some tasks with a nonstationary policy, ultimately degrading performance. To address these limitations, we propose value diffusion reinforcement learning (VDRL), a novel model-free online RL method that utilizes the generative capacity of diffusion models to represent multimodal value distributions. The core innovation of VDRL lies in the use of the variational loss of diffusion-based value distribution, which is theoretically proven to be a tight lower bound for the optimization objective under the KL-divergence measurement. Furthermore, we introduce double value diffusion learning with sample selection to enhance training stability and further improve value estimation accuracy. Extensive experiments conducted on the MuJoCo benchmark demonstrate that VDRL significantly outperforms some SOTA model-free online RL baselines, showcasing its effectiveness and robustness. Xiaoliang Hu, Fuyun Wang, Tong Zhang 0021, Zhen Cui 0001 |
NeurIPS | 4 |
| 2025 | One for All: Universal Topological Primitive Transfer for Graph Structure LearningabstractThe non-Euclidean geometry inherent in graph structures fundamentally impedes cross-graph knowledge transfer. Drawing inspiration from texture transfer in computer vision, we pioneer topological primitives as transferable semantic units for graph structural knowledge. To address three critical barriers - the absence of specialized benchmarks, aligned semantic representations, and systematic transfer methodologies - we present G²SN-Transfer, a unified framework comprising: (i) TopoGraph-Mapping that transforms non-Euclidean graphs into transferable sequences via topological primitive distribution dictionaries; (ii) G²SN, a dual-stream architecture learning text-topology aligned representations through contrastive alignment; and (iii) AdaCross-Transfer, a data-adaptive knowledge transfer mechanism leveraging cross-attention for both full-parameter and parameter-frozen scenarios. Particularly, G²SN is a dual-stream sequence network driven by ordinary differential equations, and our theoretical analysis establishes the convergence guarantee of G²SN. We construct STA-18, the first large-scale benchmark with aligned topological primitive-text pairs across 18 diverse graph datasets. Comprehensive evaluations demonstrate that G²SN achieves state-of-the-art performance on four structural learning tasks (average 3.2\% F1-score improvement), while our transfer method yields consistent enhancements across 13 downstream tasks (5.2\% average gains) including 10 large-scale graph datasets. The datasets and code are available at https://anonymous.4open.science/r/UGSKT-C10E/. Yide Qiu, Tong Zhang 0021, Xing Cai, Zhen Cui 0001 |
NeurIPS | 5 |
| 2025 | UniHG: A Large-scale Universal Heterogeneous Graph Dataset and Benchmark for Representation Learning and Cross-Domain TransferringabstractIrregular data in the real world are usually organized as heterogeneous graphs consisting of multiple types of nodes and edges. However, current heterogeneous graph research confronts three fundamental challenges: i) Benchmark Deficiency, ii) Semantic Disalignment, and iii) Propagation Degradation. In this paper, we construct a large-scale, universal, and joint multi-domain heterogeneous graph dataset named UniHG to facilitate heterogeneous graph representation learning and cross-domain knowledge mining. Overall, UniHG contains 77.31 million nodes and 564 million directed edges with thousands of labels and attributes, which is currently the largest universal heterogeneous graph dataset available to the best of our knowledge. To perform effective learning and provide comprehensively benchmarks on UniHG , two key measures are taken, including i) the semantic alignment strategy for multi-attribute entities, which projects the feature description of multi-attribute nodes and edges into a common embedding space to facilitate information aggregation; ii) proposing the novel Heterogeneous Graph Decoupling (HGD) framework with a specifically designed Anisotropy Feature Propagation (AFP) module for learning effective multi-hop anisotropic propagation kernels. These two strategies enable efficient information propagation among a tremendous number of multi-attribute entities and meanwhile mine multi-attribute association adaptively through the multi-hop aggregation in large-scale heterogeneous graphs. Comprehensive benchmark results demonstrate that our model significantly outperforms existing methods with an accuracy improvement of 28.93\%. And the UniHG can facilitate downstream tasks, achieving an NDCG@20 improvement rate of 11.48\% and 11.71\%. The UniHG dataset and benchmark codes have been released at https://anonymous.4open.science/r/UniHG-AA78. Yide Qiu, Tong Zhang 0021, Shaoxiang Ling, Xing Cai, Ziqi Gu, Zhen Cui 0001 |
NeurIPS | 6 |
| 2025 | Diffusion Dynamic Model for Unsupervised Reinforcement Learning
Xiaoliang Hu, Zhen Cui 0001, Luying Wu, Tong Zhang 0021 |
PRCV (2) | 3 |
| 2025 | DBB-Det: High-Precision Oriented Object Detection via Deformable Bounding Box Representation
Yiliang Liu, Zhen Cui 0001 |
PRCV (17) | 4 |
| 2025 | MPDS: A Movie Posters Dataset for Image Generation with Diffusion Model
Tong Zhang 0021, Fuyun Wang, Xin Liu 0011, Zhen Cui 0001 |
PRCV (2) | 6 |
| 2025 | Boosting Graph Convolution with Disparity-induced Structural RefinementabstractGraph Neural Networks (GNNs) have expressed remarkable capability in processing graph-structured data. Recent studies have found that most GNNs rely on the homophily assumption of graphs, leading to unsatisfactory performance on heterophilous graphs. While certain methods have been developed to address heterophilous links, they lack more precise estimation of high-order relationships between nodes. This could result in the aggregation of excessive interference information during message propagation, thus degrading the representation ability of learned features. In this work, we propose a Disparity-induced Structural Refinement (DSR) method that enables adaptive and selective message propagation in GNN, to enhance representation learning in heterophilous graphs. We theoretically analyze the necessity of structural refinement during message passing grounded in the derivation of error bound for node classification. To this end, we design a disparity score that combines both features and structural information at the node level, reflecting the connectivity degree of hopping neighbor nodes. Based on the disparity score, we can adjust the aggregation of neighbor nodes, thereby mitigating the impact of irrelevant information during message passing. Experimental results demonstrate that our method achieves competitive performance, mostly outperforming advanced methods on both homophilous and heterophilous datasets. Sujia Huang, Yueyang Pi, Tong Zhang 0021, Zhen Cui 0001 |
WWW | 5 |
| 2025 | A Generalized Contour Vibration Model for Building Extraction
Chunyan Xu, Shuaizhen Yao, Zhen Cui 0001, Jian Yang 0003 |
Int. J. Comput. Vis. | 4 |
| 2025 | Correction: A Generalized Contour Vibration Model for Building Extraction
Chunyan Xu, Shuaizhen Yao, Zhen Cui 0001, Jian Yang 0003 |
Int. J. Comput. Vis. | 4 |
| 2025 | Adaptive exploration for few-shot incremental learning
Cao Han, Ziqi Gu, Chunyan Xu, Zhen Cui 0001 |
Knowl. Based Syst. | 4 |
| 2025 | Instance-Consistent Fair Face RecognitionabstractThe fairness of face recognition (FR) is a challenging issue to numerous FR algorithms in the modern pluralistic and egalitarian society. In this work, we propose an instance-consistent fair face recognition (IC-FFR) method by fulfilling complete instance fairness on false positive rate (FPR) and true positive rate (TPR). In view of the misalignment of testing and training metrics, not yet considered by the current fair FR algorithms, in theory, we inspect the correlation between the testing metrics (FPR and TPR) and the label classification loss, and we derive a high-probability consistency of unfairness penalties from FPR and TPR to the softmax loss. According to the theoretical analysis, we further develop an instance-consistent fairness solution by introducing customized instance margins, which well preserve consistent FPR and TPR of all instances during the label classification in training. To encourage more fine-grained fairness evaluation, we contribute a dataset called national faces in the world (NFW) to measure the fairness of individuals and countries. Extensive experiments on our NFW as well as the RFW and BFW benchmarks demonstrate the effectiveness and superiority of our method compared to those state-of-the-art fair FR methods. Yong Li 0032, Zhen Cui 0001, Pengcheng Shen, Shiguang Shan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Collaborative contrastive learning for cross-domain gaze estimation
Lifan Xia, Yong Li 0032, Zhen Cui 0001, Chunyan Xu, Antoni B. Chan |
Pattern Recognit. | 4 |
| 2025 | Fragment-Driven Progressive Alternating Diffusion for De Novo Molecular DesignabstractHigh reliability and creativity remain key goals for AI-driven de novo molecule design. In this work, we propose a fragment-driven progressive alternating diffusion (FDPAD) framework in a coarse-to-fine generation mode. By modeling molecules as fragment-structured graphs, FDPAD entails a progressive discrete diffusion process by randomly walking some sequences of fragment-structured units (FSU), thereby mitigating combinatorial complexities and facilitating the synthesis of intricate macroscopic structures. To delve deeper internal structures of FSU, we design two distinct diffusion processes: the conditioned fragment diffusion (CFD) and the inter-fragment bond diffusion (IBD). In CFD, a string-based diffusion probability model is proposed to enrich the diversity of fragments, leveraging the partially-generated molecule as condition. And in IBD, a graph-based diffusion model upon bond-related atom graph is proposed to boost the prediction of intricate chemical bond connections among molecular fragments. Through the interleaving of CFD and IBD processes, our model outperforms state-of-the-art algorithms in de novo molecular generation, particularly in generating novel and unique molecules. Xing Cai, Tong Zhang 0021, Yide Qiu, Zhen Cui 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | Deciphering the Structural Code of Proteins With Deep Graph LearningabstractDeciphering the three-dimensional structure of proteins remains a grand challenge in biology and medicine, as it holds the key to understanding their biological functions and facilitating drug discovery. In this paper, we introduce DECIPHER (Deep Encoding of Cellular Interactions and Protein HiErarchical Representation), a novel deep graph learning framework for protein structure prediction. By representing proteins as graphs, where residues and atoms serve as nodes and their interactions form edges, we capture the intricate spatial relationships within these complex biomolecules. Our framework consists of two complementary modules: 1) a general protein structure prediction module that employs residue and atomic graphs to predict backbone and side-chain conformations, respectively, and utilizes SE(3) transformation for structure optimization; and 2) an antibody-specific structure prediction module that incorporates a dual-track network architecture to model sequence co-evolution and structural template information, coupled with a physics-based energy optimization process. Through extensive experiments on multiple benchmark datasets, we demonstrate that our approach significantly outperforms state-of-the-art methods, setting new standards for accuracy and efficiency in protein structure prediction. By deciphering the structural code of proteins, our work paves the way for accelerated research on protein function and opens up new avenues for rational drug design and discovery. Xiaoyi Yin, Xin Liu 0011, Zhen Cui 0001, Tong Zhang 0021 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | Decoupled Doubly Contrastive Learning for Cross-Domain Facial Action Unit DetectionabstractDespite the impressive performance of current vision-based facial action unit (AU) detection approaches, they are heavily susceptible to the variations across different domains and the cross-domain AU detection methods are under-explored. In response to this challenge, we propose a decoupled doubly contrastive adaptation (D2CA) approach to learn a purified AU representation that is semantically aligned for the source and target domains. Specifically, we decompose latent representations into AU-relevant and AU-irrelevant components, with the objective of exclusively facilitating adaptation within the AU-relevant subspace. To achieve the feature decoupling, D2CA is trained to disentangle AU and domain factors by assessing the quality of synthesized faces in cross-domain scenarios when either AU or domain attributes are modified. To further strengthen feature decoupling, particularly in scenarios with limited AU data diversity, D2CA employs a doubly contrastive learning mechanism comprising image and feature-level contrastive learning to ensure the quality of synthesized faces and mitigate feature ambiguities. This new framework leads to an automatically learned, dedicated separation of AU-relevant and domain-relevant factors, and it enables intuitive, scale-specific control of the cross-domain facial image synthesis. Extensive experiments demonstrate the efficacy of D2CA in successfully decoupling AU and domain factors, yielding visually pleasing cross-domain synthesized facial images. Meanwhile, D2CA consistently outperforms state-of-the-art cross-domain AU detection approaches, achieving an average F1 score improvement of 6%-14% across various cross-domain scenarios. Yong Li 0032, Menglin Liu, Zhen Cui 0001, Yi Ding 0012, Yuan Zong, Wenming Zheng, Shiguang Shan, Cuntai Guan |
IEEE Trans. Image Process. | 3 |
| 2025 | TVDO: Tchebycheff Value-Decomposition Optimization for Multiagent Reinforcement LearningabstractIn cooperative multiagent reinforcement learning (MARL), centralized training with decentralized execution (CTDE) has recently attracted more attention due to the physical demand. However, the most dilemma therein is the inconsistency between jointly-trained policies and individually executed actions. In this article, we propose a factorized Tchebycheff value-decomposition optimization (TVDO) method to overcome the trouble of inconsistency. In particular, a nonlinear Tchebycheff aggregation function is formulated to realize the global optimum by tightly constraining the upper bound of individual action-value bias, which is inspired by the Tchebycheff method of multiobjective optimization (MOO). We theoretically prove that, under no extra limitations, the factorized value decomposition with Tchebycheff aggregation satisfies the sufficiency and necessity of individual-global-max (IGM), which guarantees the consistency between the global and individual optimal action-value function. Empirically, in the climb and penalty game, we verify that TVDO precisely expresses the global-to-individual value decomposition with a guarantee of policy consistency. Meanwhile, we evaluate TVDO in the StarCraft multiagent challenge (SMAC) benchmark, and extensive experiments demonstrate that TVDO achieves a significant performance superiority over some SOTA MARL baselines. Xiaoliang Hu, Zhen Cui 0001, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | MMHCL: Multi-Modal Hypergraph Contrastive Learning for RecommendationabstractThe burgeoning presence of multimodal content-sharing platforms propels the development of personalized recommender systems. Previous works usually suffer from data sparsity and cold-start problems and may fail to adequately explore semantic user–product associations from multimodal data. To address these issues, we propose a novel Multi-Modal Hypergraph Contrastive Learning (MMHCL) framework for user recommendation. For a comprehensive information exploration from user–product relations, we construct two hypergraphs, i.e., a user-to-user (u2u) hypergraph and an item-to-item (i2i) hypergraph, to mine shared preferences among users and intricate multimodal semantic resemblance among items, respectively. This process yields denser second-order semantics that are fused with first-order user–item interaction as complementary to alleviate the data sparsity issue. Then, we design a contrastive feature enhancement paradigm by applying synergistic contrastive learning. By maximizing/minimizing the mutual information between second-order (e.g., shared preference pattern for users) and first-order (information of selected items for users) embeddings of the same/different users and items, the feature distinguishability can be effectively enhanced. Compared with using sparse primary user–item interaction only, our MMHCL obtains denser second-order hypergraphs and excavates more abundant shared attributes to explore the user–product associations, which to a certain extent alleviates the problems of data sparsity and cold-start. Extensive experiments have comprehensively demonstrated the effectiveness of our method. Our code is publicly available at https://github.com/Xu107/MMHCL . Tong Zhang 0021, Fuyun Wang, Xin Liu 0011, Zhen Cui 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | ICPR 2024 Competition on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Han Wang 0049, Furui Chen, Silei Liu, Xiaomin Huang, Shining Wang, Ying Li 0017, Peng Wang 0015, Shiyong Peng, Xiaokai Bi, Renbin Zou, Wenjing Deng, Zhen Cui 0001 |
ICPR (34) | 27 |
| 2024 | Progressive Exploration-Conformal Learning for Sparsely Annotated Object Detection in Aerial ImagesabstractThe ability to detect aerial objects with limited annotation is pivotal to the development of real-world aerial intelligence systems. In this work, we focus on a demanding but practical sparsely annotated object detection (SAOD) in aerial images, which encompasses a wider variety of aerial scenes with the same number of annotated objects. Although most existing SAOD methods rely on fixed thresholding to filter pseudo-labels for enhancing detector performance, adapting to aerial objects proves challenging due to the imbalanced probabilities/confidences associated with predicted aerial objects. To address this problem, we propose a novel Progressive Exploration-Conformal Learning (PECL) framework to address the SAOD task, which can adaptively perform the selection of high-quality pseudo-labels in aerial images. Specifically, the pseudo-label exploration can be formulated as a decision-making paradigm by adopting a conformal pseudo-label explorer and a multi-clue selection evaluator. The conformal pseudo-label explorer learns an adaptive policy by maximizing the cumulative reward, which can decide how to select these high-quality candidates by leveraging their essential characteristics and inter-instance contextual information. The multi-clue selection evaluator is designed to evaluate the explorer-guided pseudo-label selections by providing an instructive feedback for policy optimization. Finally, the explored pseudo-labels can be adopted to guide the optimization of aerial object detector in a closed-looping progressive fashion. Comprehensive evaluations on two public datasets demonstrate the superiority of our PECL when compared with other state-of-the-art methods in the sparsely annotated aerial object detection task. Zihan Lu, Chunyan Xu, Xiangwei Zheng 0001, Zhen Cui 0001 |
NeurIPS | 5 |
| 2024 | MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image GenerationabstractRecently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote sensing (RS) images that are tremendously different from general images in terms of scale and perspective remains a formidable challenge due to the lack of a comprehensive remote sensing image generation dataset with various modalities, ground sample distances (GSD), and scenes. In this paper, we propose a Multi-modal, Multi-GSD, Multi-scene Remote Sensing (MMM-RS) dataset and benchmark for text-to-image generation in diverse remote sensing scenarios. Specifically, we first collect nine publicly available RS datasets and conduct standardization for all samples. To bridge RS images to textual semantic information, we utilize a large-scale pretrained vision-language model to automatically output text prompts and perform hand-crafted rectification, resulting in information-rich text-image pairs (including multi-modal images). In particular, we design some methods to obtain the images with different GSD and various environments (e.g., low-light, foggy) in a single sample. With extensive manual screening and refining annotations, we ultimately obtain a MMM-RS dataset that comprises approximately 2.1 million text-image pairs. Extensive experimental results verify that our proposed MMM-RS dataset allows off-the-shelf diffusion models to generate diverse RS images across various modalities, scenes, weather conditions, and GSD. The dataset is available at https://github.com/ljl5261/MMM-RS. Jialin Luo, Yuanzhi Wang, Ziqi Gu, Yide Qiu, Shuaizhen Yao, Fuyun Wang, Chunyan Xu, Zhen Cui 0001 |
NeurIPS | 10 |
| 2024 | An attribution graph-based interpretable method for CNNs
Xiangwei Zheng 0001, Chunyan Xu, Xuanchi Chen, Zhen Cui 0001 |
Neural Networks | 5 |
| 2024 | Wasserstein Discriminant Dictionary Learning for Graph RepresentationabstractMining discriminative graph topological information plays an important role in promoting graph representation ability. However, it suffers from two main issues: (1) the difficulty/complexity of computing global inter-class/intra-class scatters, commonly related to mean and covariance of graph samples, for discriminant learning; (2) the huge complexity and variety of graph topological structure that is rather challenging to robustly characterize. In this paper, we propose the Wasserstein Discriminant Dictionary Learning (WDDL) framework to achieve discriminant learning on graphs with robust graph topology modeling, and hence facilitate graph-based pattern analysis tasks. Considering the difficulty of calculating global inter-class/intra-class scatters, a reference set of graphs (aka graph dictionary) is first constructed by generating representative graph samples (aka graph keys) with expressive topological structure. Then, a Wasserstein Graph Representation (WGR) process is proposed to project input graphs into a succinct dictionary space through the graph dictionary lookup. To further achieve discriminant graph learning, a Wasserstein discriminant loss (WD-loss) is defined on the graph dictionary, in which the graph keys are optimizable, to make the intra-class keys more compact and inter-class keys more dispersed. Hence, the calculation of global Wasserstein metric (W-metric) centers can be bypassed. For sophisticated topology mining in the WGR process, a joint-Wasserstein graph embedding module is constructed to model both between-node and between-edge relationships across inputs and graph keys by encapsulating both the Wasserstein metric (between cross-graph nodes) and proposed novel Kron-Gromov-Wasserstein (KGW) metric (between cross-graph adjacencies). Specifically, the KGW-metric comprehensively characterizes the cross-graph connection patterns with the Kronecker operation, then adaptively captures those salient patterns through connection pooling. To evaluate the proposed framework, we study two graph-based pattern analysis problems, i.e. graph classification and cross-modal retrieval, with the graph dictionary flexibly adjusted to cater to these two tasks. Extensive experiments are conducted to comprehensively compare with existing advanced methods, as well as dissect the critical component of our proposed architecture. The experimental results validate the effectiveness of the WDDL framework. Tong Zhang 0021, Guangbu Liu, Zhen Cui 0001, Wei Liu 0005, Wenming Zheng, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | OLCH: Online Label Consistent Hashing for streaming cross-modal retrieval
Shu-Juan Peng, Jinhan Yi, Xin Liu 0011, Yiu-Ming Cheung, Zhen Cui 0001, Taihao Li |
Pattern Recognit. | 5 |
| 2024 | Intra-Inter Graph Representation Learning for Protein-Protein Binding Sites PredictionabstractGraph neural networks have drawn increasing attention and achieved remarkable progress recently due to their potential applications for a large amount of irregular data. It is a natural way to represent protein as a graph. In this work, we focus on protein-protein binding sites prediction between the ligand and receptor proteins. Previous work just simply adopts graph convolution to learn residue representations of ligand and receptor proteins, then concatenates them and feeds the concatenated representation into a fully connected layer to make predictions, losing much of the information contained in complexes and failing to obtain an optimal prediction. In this paper, we present Intra-Inter Graph Representation Learning for protein-protein binding sites prediction (IIGRL). Specifically, for intra-graph learning, we maximize the mutual information between local node representation and global graph summary to encourage node representation to embody the global information of protein graph. Then we explore fusing two separate ligand and receptor graphs as a whole graph and learning affinities between their residues/nodes to propagate information to each other, which could effectively capture inter-protein information and further enhance the discrimination of residue pairs. Extensive experiments on multiple benchmarks demonstrate that the proposed IIGRL model outperforms state-of-the-art methods. Wenting Zhao 0001, Gongping Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | Object Knowledge Distillation for Joint Detection and Tracking in Satellite VideosabstractExisting mainstream MOT methods can be categorised into two frameworks including two-stage and one-stage ones. Two-stage ones divide MOT task into object detection and association tasks which usually achieve high accuracy. One-stage ones train a joint model to achieve both detection and tracking. So their advantage usually lies in the high tracking efficiency. In this paper, we inherit the advantages of the two types frameworks and propose the object knowledge distilled joint detection and tracking framework (OKD-JDT) to achieve accurate as well as efficient tracking. Firstly, the performance of two-stage methods largely depends on the highly performed detection network. So, we treat the detection network as the teacher network to guide the discriminative object feature learning in one-stage methods by using knowledge distillation. Then, in distillation learning, we design the adaptive attention learning to learn the discriminative features from teacher network to student network. In addition, with the similar appearance and uniform moving behaviour of objects in satellite videos, we propose to use joint center point distance and intersection-over-onion (IOU) to generate tracklets. Experiments on JiLin-1 satellite videos with different objects demonstrate the effectiveness and the state-of-the-art performance of the proposed method. Wenjing Deng, Zhen Cui 0001, Jia Liu 0020, Licheng Jiao |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Cross-Attention Regression Flow for Defect DetectionabstractDefect detection from images is a crucial and challenging topic of industry scenarios due to the scarcity and unpredictability of anomalous samples. However, existing defect detection methods exhibit low detection performance when it comes to small-size defects. In this work, we propose a Cross-Attention Regression Flow (CARF) framework to model a compact distribution of normal visual patterns for separating outliers. To retain rich scale information of defects, we build an interactive cross-attention pattern flow module to jointly transform and align distributions of multi-layer features, which is beneficial for detecting small-size defects that may be annihilated in high-level features. To handle the complexity of multi-layer feature distributions, we introduce a layer-conditional autoregression module to improve the fitting capacity of data likelihoods on multi-layer features. By transforming the multi-layer feature distributions into a latent space, we can better characterize normal visual patterns. Extensive experiments on four public datasets and our collected industrial dataset demonstrate that the proposed CARF outperforms state-of-the-art methods, particularly in detecting small-size defects. Tianchu Guo, Bin Luo 0008, Zhen Cui 0001, Jian Yang 0003 |
IEEE Trans. Image Process. | 4 |
| 2024 | Relation-Aggregated Cross-Graph Correlation Learning for Fine-Grained Image-Text RetrievalabstractFine-grained image-text retrieval has been a hot research topic to bridge the vision and languages, and its main challenge is how to learn the semantic correspondence across different modalities. The existing methods mainly focus on learning the global semantic correspondence or intramodal relation correspondence in separate data representations, but which rarely consider the intermodal relation that interactively provide complementary hints for fine-grained semantic correlation learning. To address this issue, we propose a relation-aggregated cross-graph (RACG) model to explicitly learn the fine-grained semantic correspondence by aggregating both intramodal and intermodal relations, which can be well utilized to guide the feature correspondence learning process. More specifically, we first build semantic-embedded graph to explore both fine-grained objects and their relations of different media types, which aim not only to characterize the object appearance in each modality, but also to capture the intrinsic relation information to differentiate intramodal discrepancies. Then, a cross-graph relation encoder is newly designed to explore the intermodal relation across different modalities, which can mutually boost the cross-modal correlations to learn more precise intermodal dependencies. Besides, the feature reconstruction module and multihead similarity alignment are efficiently leveraged to optimize the node-level semantic correspondence, whereby the relation-aggregated cross-modal embeddings between image and text are discriminatively obtained to benefit various image-text retrieval tasks with high retrieval performance. Extensive experiments evaluated on benchmark datasets quantitatively and qualitatively verify the advantages of the proposed framework for fine-grained image-text retrieval and show its competitive performance with the state of the arts. Shu-Juan Peng, Xin Liu 0011, Yiu-Ming Cheung, Xing Xu 0001, Zhen Cui 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Edit Temporal-Consistent Videos with Image Diffusion ModelabstractLarge-scale text-to-image (T2I) diffusion models have been extended for text-guided video editing, yielding impressive zero-shot video editing performance. Nonetheless, the generated videos usually show spatial irregularities and temporal inconsistencies as the temporal characteristics of videos have not been faithfully modeled. In this article, we propose an elegant yet effective Temporal-Consistent Video Editing (TCVE) method to mitigate the temporal inconsistency challenge for robust text-guided video editing. In addition to the utilization of a pretrained T2I 2D Unet for spatial content manipulation, we establish a dedicated temporal Unet architecture to faithfully capture the temporal coherence of the input video sequences. Furthermore, to establish coherence and interrelation between the spatial-focused and temporal-focused components, a cohesive spatial-temporal modeling unit is formulated. This unit effectively interconnects the temporal Unet with the pretrained 2D Unet, thereby enhancing the temporal consistency of the generated videos while preserving the capacity for video content manipulation. Quantitative experimental results and visualization results demonstrate that TCVE achieves state-of-the-art performance in both video temporal consistency and video editing capability, surpassing existing benchmarks in the field. Codes are released at https://github.com/mdswyz/TCVE . Yuanzhi Wang, Yong Li 0032, Xin Liu 0011, Anbo Dai, Antoni B. Chan, Zhen Cui 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | Deep Graph Structural InfomaxabstractIn the scene of self-supervised graph learning, Mutual Information (MI) was recently introduced for graph encoding to generate robust node embeddings. A successful representative is Deep Graph Infomax (DGI), which essentially operates on the space of node features but ignores topological structures, and just considers global graph summary. In this paper, we present an effective model called Deep Graph Structural Infomax (DGSI) to learn node representation. We explore to derive the structural mutual information from the perspective of Information Bottleneck (IB), which defines a trade-off between the sufficiency and minimality of representation on the condition of the topological structure preservation. Intuitively, the derived constraints formally maximize the structural mutual information both edge-wise and local neighborhood-wise. Besides, we develop a general framework that incorporates the global representational mutual information, local representational mutual information, and sufficient structural information into the node representation. Essentially, our DGSI extends DGI and could capture more fine-grained semantic information as well as beneficial structural information in a self-supervised manner, thereby improving node representation and further boosting the learning performance. Extensive experiments on different types of datasets demonstrate the effectiveness and superiority of the proposed method. Wenting Zhao 0001, Gongping Xu, Zhen Cui 0001, Siqiang Luo, Cheng Long 0001, Tong Zhang 0021 |
AAAI | 3 |
| 2023 | Exploratory Inference Learning for Scribble Supervised Semantic SegmentationabstractScribble supervised semantic segmentation has achieved great advances in pseudo label exploitation, yet suffers insufficient label exploration for the mass of unannotated regions. In this work, we propose a novel exploratory inference learning (EIL) framework, which facilitates efficient probing on unlabeled pixels and promotes selecting confident candidates for boosting the evolved segmentation. The exploration of unannotated regions is formulated as an iterative decision-making process, where a policy searcher learns to infer in the unknown space and the reward to the exploratory policy is based on a contrastive measurement of candidates. In particular, we devise the contrastive reward with the intra-class attraction and the inter-class repulsion in the feature space w.r.t the pseudo labels. The unlabeled exploration and the labeled exploitation are jointly balanced to improve the segmentation, and framed in a close-looping end-to-end network. Comprehensive evaluations on the benchmark datasets (PASCAL VOC 2012 and PASCAL Context) demonstrate the superiority of our proposed EIL when compared with other state-of-the-art methods for the scribble-supervised semantic segmentation problem. Chuanwei Zhou, Zhen Cui 0001, Chunyan Xu, Cao Han, Jian Yang 0003 |
AAAI | 2 |
| 2023 | Progressive Bayesian Inference for Scribble-Supervised Semantic SegmentationabstractThe scribble-supervised semantic segmentation is an important yet challenging task in the field of computer vision. To deal with the pixel-wise sparse annotation problem, we propose a Progressive Bayesian Inference (PBI) framework to boost the performance of the scribble-supervised semantic segmentation, which can effectively infer the semantic distribution of these unlabeled pixels to guide the optimization of the segmentation network. The PBI dynamically improves the model learning from two aspects: the Bayesian inference module (i.e., semantic distribution learning) and the pixel-wise segmenter (i.e., model updating). Specifically, we effectively infer the semantic probability distribution of these unlabeled pixels with our designed Bayesian inference module, where its guidance is estimated through the Bayesian expectation maximization under the situation of partially observed data. The segmenter can be progressively improved under the joint guidance of the original scribble information and the learned semantic distribution. The segmenter optimization and semantic distribution promotion are encapsulated into a unified architecture where they could improve each other with mutual evolution in a progressive fashion. Comprehensive evaluations of several benchmark datasets demonstrate the effectiveness and superiority of our proposed PBI when compared with other state-of-the-art methods applied to the scribble-supervised semantic segmentation task. Chuanwei Zhou, Chunyan Xu, Zhen Cui 0001 |
AAAI | 3 |
| 2023 | Decoupled Multimodal Distilling for Emotion RecognitionabstractHuman multimodal emotion recognition (MER) aims to perceive human emotions via language, visual and acoustic modalities. Despite the impressive performance of previous MER approaches, the inherent multimodal heterogeneities still haunt and the contribution of different modalities varies significantly. In this work, we mitigate this issue by proposing a decoupled multimodal distillation (DMD) approach that facilitates flexible and adaptive crossmodal knowledge distillation, aiming to enhance the discriminative features of each modality. Specially, the representation of each modality is decoupled into two parts, i.e., modality-irrelevant/-exclusive spaces, in a self-regression manner. DMD utilizes a graph distillation unit (GD-Unit) for each decoupled part so that each GD can be performed in a more specialized and effective manner. A GD-Unit consists of a dynamic graph where each vertice represents a modality and each edge indicates a dynamic knowledge distillation. Such GD paradigm provides a flexible knowledge transfer manner where the distillation weights can be automatically learned, thus enabling diverse crossmodal knowledge transfer patterns. Experimental results show DMD consistently obtains superior performance than state-of-the-art MER methods. Visualization results show the graph edges in DMD exhibit meaningful distributional patterns w.r.t. the modality-irrelevant/-exclusive feature spaces. Codes are re leased at https://github.com/mdswyz/DMD. Yuanzhi Wang, Zhen Cui 0001 |
CVPR | 3 |
| 2023 | Unbiased Multiple Instance Learning for Weakly Supervised Video Anomaly DetectionabstractWeakly Supervised Video Anomaly Detection (WSVAD) is challenging because the binary anomaly label is only given on the video level, but the output requires snippet-level predictions. So, Multiple Instance Learning (MIL) is prevailing in WSVAD. However, MIL is notoriously known to suffer from many false alarms because the snippet-level detector is easily biased towards the abnormal snippets with simple context, confused by the normality with the same bias, and missing the anomaly with a different pattern. To this end, we propose a new MIL framework: Unbiased MIL (UMIL), to learn unbiased anomaly features that improve WSVAD. At each MIL training iteration, we use the current detector to divide the samples into two groups with different context biases: the most confident abnormal/normal snippets and the rest ambiguous ones. Then, by seeking the invariant features across the two sample groups, we can remove the variant context biases. Extensive experiments on benchmarks UCF-Crime and TAD demonstrate the effectiveness of our UMIL. Our code is provided at https://github.com/ktr-hubrt/UMIL. Zhongqi Yue, Qianru Sun, Bin Luo 0008, Zhen Cui 0001, Hanwang Zhang |
CVPR | 5 |
| 2023 | Few-shot Continual Infomax LearningabstractFew-shot continual learning is the ability to continually train a neural network from a sequential stream of few-shot data. In this paper, we propose a Few-shot Continual Infomax Learning (FCIL) framework that makes a deep model to continually/incrementally learn new concepts from few labeled samples, relieving the catastrophic forgetting of past knowledge. Specifically, inspired by the theoretical definition of transfer entropy, we introduce a feature embedding infomax to effectively perform the few-shot learning, which can transfer the strong encoding capability of the base network to learn the feature embedding of these novel classes by maximizing the mutual information of different-level feature distributions. Further, considering that the learned knowledge in the human brain is a generalization of actual information and exists in a certain relational structure, we perform continual structure infomax learning to relieve the catastrophic forgetting problem in the continual learning process. The information structure of this learned knowledge can be preserved through maximizing the mutual information across these continual-changing relations of inter-classes. Comprehensive evaluations on CIFAR100, miniImageNet, and CUB200 datasets demonstrate the superiority of our FCIL when compared against state-of-the-art methods on the few-shot continual learning task. Ziqi Gu, Chunyan Xu, Jian Yang 0003, Zhen Cui 0001 |
ICCV | 4 |
| 2023 | Distribution-Consistent Modal Recovering for Incomplete Multimodal LearningabstractRecovering missing modality is popular in incomplete multimodal learning because it usually benefits downstream tasks. However, the existing methods often directly estimate missing modalities from the observed ones by deep neural networks, lacking consideration of the distribution gap between modalities, resulting in the inconsistency of distributions between the recovered and the true data. To mitigate this issue, in this work, we propose a novel recovery paradigm, Distribution-Consistent Modal Recovering (DiCMoR), to transfer the distributions from available modalities to missing modalities, which thus maintains the distribution consistency of recovered data. In particular, we design a class-specific flow based modality recovery method to transform cross-modal distributions on the condition of sample class, which could well predict a distribution-consistent space for missing modality by virtue of the invertibility and exact density estimation of normalizing flow. The generated data from the predicted distribution is integrated with available modalities for the task of classification. Experiments show that DiCMoR gains superior performances and is more robust than existing state-of-the-art methods under various missing patterns. Visualization results show that the distribution gaps between recovered modalities and missing modalities are mitigated. Codes are released at https://github.com/mdswyz/DiCMoR. Yuanzhi Wang, Zhen Cui 0001 |
ICCV | 2 |
| 2023 | Incomplete Multimodality-Diffused Emotion RecognitionabstractHuman multimodal emotion recognition (MER) aims to perceive and understand human emotions via various heterogeneous modalities, such as language, vision, and acoustic. Compared with unimodality, the complementary information in the multimodalities facilitates robust emotion understanding. Nevertheless, in real-world scenarios, the missing modalities hinder multimodal understanding and result in degraded MER performance. In this paper, we propose an Incomplete Multimodality-Diffused emotion recognition (IMDer) method to mitigate the challenge of MER under incomplete multimodalities. To recover the missing modalities, IMDer exploits the score-based diffusion model that maps the input Gaussian noise into the desired distribution space of the missing modalities and recovers missing data abided by their original distributions. Specially, to reduce semantic ambiguity between the missing and the recovered modalities, the available modalities are embedded as the condition to guide and refine the diffusion-based recovering process. In contrast to previous work, the diffusion-based modality recovery mechanism in IMDer allows to simultaneously reach both distribution consistency and semantic disambiguation. Feature visualization of the recovered modalities illustrates the consistent modality-specific distribution and semantic alignment. Besides, quantitative experimental results verify that IMDer obtains state-of-the-art MER accuracy under various missing modality patterns. Yuanzhi Wang, Zhen Cui 0001 |
NeurIPS | 3 |
| 2023 | Grassmann Graph Embedding for Few-Shot Class Incremental Learning
Ziqi Gu, Chunyan Xu, Zhen Cui 0001 |
PRCV (8) | 3 |
| 2023 | Global Variational Convolution Network for Semi-supervised Node Classification on Large-Scale Graphs
Yide Qiu, Tong Zhang 0021, Zhen Cui 0001 |
PRCV (8) | 4 |
| 2023 | Learning cross-modal interaction for RGB-T tracking
Chunyan Xu, Zhen Cui 0001, Chaoqun Wang 0012, Chuanwei Zhou, Jian Yang 0003 |
Sci. China Inf. Sci. | 2 |
| 2023 | Two-directional two-dimensional fractional-order embedding canonical correlation analysis for multi-view dimensionality reduction and set-based video recognition
Yinghui Sun, Xizhan Gao, Sijie Niu, Dong Wei 0007, Zhen Cui 0001 |
Expert Syst. Appl. | 5 |
| 2023 | Prototype-guided Instance matching for multiple pedestrian tracking
Qiang Wang 0023, Wankou Yang, Chunyan Xu, Zhen Cui 0001 |
Neurocomputing | 5 |
| 2023 | A graph-based interpretability method for deep neural networks
Xiangwei Zheng 0001, Zhen Cui 0001, Chunyan Xu |
Neurocomputing | 4 |
| 2023 | Quality-aware pattern diffusion for video object segmentation
Chuanwei Zhou, Chunyan Xu, Jun Li 0027, Zhen Cui 0001, Jian Yang 0003 |
Neurocomputing | 4 |
| 2023 | Visual Micro-Pattern PropagationabstractStatistic observations demonstrate that visual feature patterns or structure patterns recur high-frequently within/across homo/heterogeneous images. Motivated by the interdependencies of visual patterns, we propose visual micro-pattern propagation (VMPP) to facilitate universal visual pattern learning. Especially, we present a graph framework to unify the conventional micro-pattern propagations in spatial, temporal, cross-modal and cross-task domains. A general formulation of pattern propagation named cross-graph model is presented under this framework, and accordingly a factorized version is derived for more efficient computation as well as better understanding. To correlate homo/heterogeneous patterns, in cross-graph we introduce two types of pattern relations from feature-level and structure-level. The structure pattern relation defines second-order visual connections for heterogeneous patterns by measuring first-order visual relations of homogeneous feature patterns. In virtue of the constructed first-/second-order connections, we design feature pattern diffusion and structure pattern diffusion to prop up various pattern propagation cases. To fulfill different pattern diffusions involved, further, we deeply study two fundamental visual problems, multi-task pixel-level prediction and online dual-modal object tracking, and accordingly propose two end-to-end pattern propagation networks by encapsulating and integrating some necessary diffusion modules therein. We conduct extensive experiments by dissecting every diffusion component as well as comparing numerous advanced methods. The experiments validate the effectiveness of our proposed various pattern diffusion ways and meantime report the state-of-the-art results on the two representative visual problems. Zhen Cui 0001, Chaoqun Wang 0012, Chunyan Xu, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Variational Instance-Adaptive Graph for EEG Emotion RecognitionabstractThe individual differences and the dynamic uncertain relationships among different electroencephalogram (EEG) regions are essential factors that limit EEG emotion recognition. To address these issues, in this article, we propose a variational instance-adaptive graph method (V-IAG) that simultaneously captures the individual dependencies among different EEG electrodes and estimates the underlying uncertain information. Specifically, we employ two branches, i.e., instance-adaptive branch and variational branch, to construct the graph. Inspired by the attention mechanism, the instance-adaptive branch generates the graph based on the input so as to characterize the individual dependencies among EEG channels. The variational branch generates the probabilistic graph, which quantifies the uncertainties. We combine these two types of graphs to extract more discriminative features. To present more precise graph representation, we propose a new operation named the multi-level and multi-graph convolution operation, which aggregates the features of EEG channels from different frequencies with different graphs. Furthermore, we design the graph coarsening and employ the sparse constraint to obtain more robust features. We conduct extensive experiments on three widely-used EEG emotion recognition databases, i.e., SJTU emotion EEG dataset (SEED), multi-modal physiological emotion recognition dataset (MPED) and DREAMER. The results demonstrate that the proposed model achieves the-state-of-the-art performance. Tengfei Song, Suyuan Liu, Wenming Zheng, Yuan Zong, Zhen Cui 0001, Yang Li 0019 |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | FENP: A Database of Neonatal Facial Expression for Pain AnalysisabstractIn this article, we introduce a new neonatal facial expression database for pain analysis. This database, called facial expression of neonatal pain (FENP), contains 11,000 neonatal facial expression images associated with 106 Chinese neonates from two children's hospitals, i.e., the Children's Hospital Affiliated to Nanjing Medical University and Second Affiliated Hospital Affiliated to Nanjing Medical University in China. The facial expression images cover four categories of facial expressions, i.e., severe pain expression, mild pain expression, crying expression and calmness expression, where each category contains 2750 neonatal facial expression images. Based on this database, we also investigate the pain facial expression recognition problem using several state-of-the-art facial expression features and expression recognition methods, such as Gabor+SVM, LBP+SVM, HOG+SVM, LBP+HOG+SVM, and several Convolutional Neural Network (CNN) methods (including AlexNet, VGGNet, GoogLeNet, ResNet and DenseNet). The experimental results indicate that the proposed neonatal pain facial expression database is very suitable for the study of both neonatal pain and facial expression recognition. Moreover, the FENP database is publicly available after signing a license agreement (the users can contact Jingjie Yan ([email protected]), Guanming Lu ([email protected])) or Xiaonan Li ([email protected]). Jingjie Yan, Guanming Lu, Wenming Zheng, Chengwei Huang, Zhen Cui 0001, Yuan Zong, Mengying Chen, Jindu Zhu, Haibo Li 0001 |
IEEE Trans. Affect. Comput. | 6 |
| 2023 | Progressive Context-Dependent Inference for Object Detection in Remote Sensing ImageryabstractInspired by our observation that numerous objects of remote sensing imageries are extremely consistent in geometric characteristics (e.g., object sizes/angles/layouts), in this work, we propose a novel Progressive Context-dependent Inference (PCI) method to make full use of large-scope contextual cues for better localizing objects in remote sensing imagery. Especially, to represent candidate objects and their geometric distributions, we build all of them into candidate object graphs, and subsequently perform inference learning by diffusing contextual object information. To make the inference more credible, we progressively accumulate these historical learning experiences on both label prediction and location regression processes into the next stage of network evolution, where topology structures and attributes of candidate object graphs would be dynamically updated. The graph update and ground object detection are jointly encapsulated as a closed-looping learning process. Hereby the problem of multi-object localization is converted into a progressive construction of dynamic graphs. Extensive experiments on three public datasets demonstrate the superiority of our proposed method over other state-of-the-art methods for ground object detection in remote sensing imagery. Chunyan Xu, Zhen Cui 0001, Jian Yang 0003 |
IEEE Trans. Image Process. | 3 |
| 2023 | Spatial-Temporal Tensor Graph Convolutional Network for Traffic Speed PredictionabstractAccurate traffic speed prediction is crucial for the guidance and management of urban traffic, which at the same time requires a model with a satisfactory computational burden and memory space in applications. In this paper, we propose a factorized Spatial-Temporal Tensor Graph Convolutional Network for traffic speed prediction. Traffic networks are modeled and unified into a graph tensor that integrates spatial and temporal information simultaneously. We extend graph convolution into tensor space and propose a tensor graph convolution network to extract more discriminating features from spatial-temporal graph data. We further introduce Tucker decomposition and derive a factorized tensor convolution to reduce the computational burden, which performs separate filtering in small-scale space, time, and feature modes. Besides, we can benefit from noise suppression of traffic data when discarding those trivial components in the process of tensor decomposition. Extensive experiments on the three real-world datasets demonstrate that our method is more effective than traditional prediction methods, and achieves state-of-the-art performance. Xuran Xu, Tong Zhang 0021, Chunyan Xu, Zhen Cui 0001, Jian Yang 0003 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | 3D3M: 3D Modulated Morphable Model for Monocular Face Reconstructionabstract3D face reconstruction from a single image is a vital task in various multimedia applications. A key challenge for 3D face shape reconstruction is to build the correct dense face correspondence between the monocular input face and the deformable mesh. Most existing methods rely on shape labels fitted by traditional methods or strong priors such as multi-view geometry consistency. In contrast, we propose an innovative 3D Modulated Morphable Model (3D3M) to learn the dense shape correspondence from monocular images in a self-supervised manner. Specifically, given a batch of input faces, 3D3M encodes their 3DMM attributes (shape, texture, lighting, etc.) and then randomly shuffles the 3DMM attributes to generate the attribute-changed faces. The attribute-changed faces can be encoded and rendered back in a cycle-consistent manner, which enables us to utilize the self-supervised consistencies in dense mesh vertices and reconstructed pixels. The dense shape and pixel correspondence enable us to adopt a series of self-supervised constraints to fit the 3D face model accurately and learn the per-vertex correctives end-to-end. 3D3M builds excellent high-quality 3D face reconstruction results from monocular images. Both quantitative and qualitative experimental results have verified the superiority of 3D3M over prior arts on 3D face reconstruction and face alignment. Yong Li 0032, Jianguo Hu, Xinmiao Pan, Zechao Li, Zhen Cui 0001 |
IEEE Trans. Multim. | 6 |
| 2023 | OMGH: Online Manifold-Guided Hashing for Flexible Cross-Modal RetrievalabstractCross-modal hashing hasrecently gained an increasing attention for its efficiency and fast retrieval speed in indexing the multimedia data across different modalities. Nevertheless, the multimedia data points often emerge in a streaming manner, and existing online methods often lack of learning capacity to handle both labeled and unlabeled data.To alleviate these concerns, this paper proposes an Online Manifold-Guided Hashing (OMGH) framework, which can incrementally learn the compact hash code of streaming data while adaptively optimizing the hash function in a streaming manner. To be specific, OMGH first exploits a matrix tri-factorization framework to learn the discriminative hash codes for streaming multi-modal data. Then, an online anchor-based manifold structure is designed to sparsely represent the old data and adaptively guide the hash code learning process, which can wellreduce the complexity in preserving the semantic correlation between the old data and streaming data. Meanwhile, such anchor-based manifold embedding is adaptive to the unsupervised and supervised learning strategies in a flexible way. Besides, an online discrete optimization method is efficiently addressed to incrementally update the hash functions and optimize the hash codes on streaming data points. As a result, the derived hash codes are more semantically meaningful for various online cross-modal retrieval tasks. Extensive experiments verify the advantages of the proposed OMGH model, by achieving and improving the state-of-the-art cross-modal retrieval performances on three benchmark datasets. Xin Liu 0011, Jinhan Yi, Yiu-Ming Cheung, Xing Xu 0001, Zhen Cui 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Instance-Aware Deep Graph Learning for Multi-Label ClassificationabstractGraph convolutional neural network (GCN) has effectively boosted the multi-label image recognition task by modeling correlation among labels. In previous methods, label correlation is computed based on statistical information through label diffusion, and therefore the same for all samples. This, however, makes graph inference on labels insufficient to handle huge variations among numerous image instances. In this paper, we propose an instance-aware graph convolutional neural network (IA_GCN) framework for the multi-label classification. As a whole, two fused branches of sub-networks are involved in the framework: a global branch modeling the whole image and a local branch exploring dependencies among regions of interests (ROIs). For both the branches, an image-dependent label correlation matrix (ID_LCM), fusing both the statistical label correlation matrix (LCM) and an individual one of each image instance, is constructed to inject adaptive information of label-awareness into the learned features of the model through graph convolution. Specifically, the individual LCM of each image is obtained by mining the label dependencies based on the predicted label scores of those detected ROIs. In this process, considering the contribution differences of ROIs to multi-label classification, variational inference is introduced to learn adaptive scaling factors for those ROIs by considering their complex distribution. Finally, extensive experiments on MS-COCO and VOC datasets show that our proposed approach outperforms existing state-of-the-art methods. Yun Wang 0028, Tong Zhang 0021, Chuanwei Zhou, Zhen Cui 0001, Jian Yang 0003 |
IEEE Trans. Multim. | 4 |
| 2023 | Tube-Embedded Transformer for Pixel PredictionabstractMulti-task pixel-level learning, which aims to exploit the inter-task interactions to improve the learning of each task, is an important but challenging issue in visual perception and multimedia applications. Measuring the inter-task correlation and intra-task specificity, we propose a tube-embedded transformer (TET) framework for robust multi-task pixel prediction. To facilitate inter-task interactions, we aggregate and project all tasks into a shared tube pool to generate the latent multi-task representation during the coarse-to-fine decoding stages. The resulting task-tube interactions replace the two-by-two task-task interactions to reduce the model complexity significantly. In addition, we introduce the transformer mechanism to adaptively transfer tube features to the target task. Concretely, on the one hand, multi-task features aggregate in the tube to generate the shared feature representation bases; on the other hand, based on the task-tube association and complementarity, the tube outputs the query entry and the weighting coefficients of the target task. Experimentally, on the joint learning of semantic segmentation, depth estimation, and surface normal estimation, the comparison experiments show the superiority of the TET multi-task learning method over other state-of-the-art approaches, and the ablation experiments verify the effectiveness of the TET mechanism. Zhen Cui 0001, Zechao Li, Jin Xie 0001, Jian Yang 0003 |
IEEE Trans. Multim. | 3 |
| 2022 | CVNet: Contour Vibration Network for Building ExtractionabstractThe classic active contour model raises a great promising solution to polygon-based object extraction with the progress of deep learning recently. Inspired by the physical vibration theory, we propose a contour vibration network (CVNet) for automatic building boundary delineation. Different from the previous contour models, the CVNet originally roots in the force and motion principle of contour string. Through the infinitesimal analysis and Newton's second law, we derive the spatial-temporal contour vibration model of object shapes, which is mathematically reduced to second-order differential equation. To concretize the dynamic model, we transform the vibration model into the space of image features, and reparameterize the equation coefficients as the learnable state from feature domain. The contour changes are finally evolved in a progressive mode through the computation of contour vibration equation. Both the polygon contour evolution and the model optimization are modulated to form a close-looping end-to-end network. Comprehensive experiments on three datasets demonstrate the effectiveness and superiority of our CVNet over other baselines and state-of-the-art methods for the polygon-based building extraction. The code is available at https://github.com/xzq-njust/CVNet. Chunyan Xu, Zhen Cui 0001, Xiangwei Zheng 0001, Jian Yang 0003 |
CVPR | 3 |
| 2022 | Inconsistency Distillation For Consistency: Enhancing Multi-View Clustering via Mutual Contrastive Teacher-Student LeaningabstractMulti-view clustering has attracted more attention recently since many real-world data are comprised of different representations or views. Recent multi-view clustering works mainly exploit the instance consistency to obtain the shared representations across different views, and apply a single-view clustering method to perform data partitions. However, these existing methods often ignore the inconsistency of instance associations within the views, which may enlarge the intra-class diversity among the views and therefore degrade the clustering performance. To address this issue, this paper proposes an efficient mutual contrastive teacher-student leaning (MC-TSL) model to enhance the multi-view clustering, which is the first attempt to study the inconsistency distillation for consistency learning. First, the proposed MC-TSL approach exploits a view-specific encoder with two heads, an instance encoding head and a semantic distillation head, respectively, for capturing the consistent and discriminative feature representations. To be specific, the former head exploits a cross-view contrastive learning method to obtain a redundancy-free consistent representation at the instance level, while the latter head designs a mutual teacher-student learning module to capture the intra-view information at semantic level. By training these two heads in an end-to-end manner, the discriminative multi-view embeddings are efficiently obtained and refined by minimizing the weighted sum of the reconstruction loss, contrastive loss and contrast distillation loss. Extensive experiments verify the superiorities of the proposed MC-TSL framework and show its competitive clustering performances. Dunqiang Liu, Shu-Juan Peng, Xin Liu 0011, Lei Zhu 0002, Zhen Cui 0001, Taihao Li |
ICDM | 5 |
| 2022 | Temporal Discriminative Micro-Expression Recognition via Graph Contrastive LearningabstractMicro-Expressions (MEs) are involuntary and im-perceptible facial movements that reflect the underlying emotions and inner activities. Recently, ME recognition technology has been widely used in several fields such as medical treatment. Due to the subtle variations among the video sequence and the limited training data, the ME recognition task still remains a challenging problem. Existing methods tend to address the ME recognition problem from two aspects: (1) Data augmentation and (2) Expression signal amplification. Few works realize the importance of temporal variation hidden in the ME sequence. Based on the above observation, we propose a Graph Contrastive Learning (GCL) framework to effectively perceive subtle temporal variation for robust ME recognition. Specifically, the strong spatial feature representation is captured through the transformer-based ME feature encoder. Then, the proposed GCL builds the graph structure for the ME sequence and introduces the graph convolution to model the temporal relationship. To capture and highlight the temporal variation hidden in the ME sequence, a contrastive learning framework is designed to discriminately learn the differences between the normal and the abnormal ME samples. Both quantitative and qualitative experimental results show the effectiveness and superiority of our method compared with the prior state-of-the-arts. Lingjie Lao, Menglin Liu, Chunyan Xu, Zhen Cui 0001 |
ICPR | 5 |
| 2022 | Extending generalized unsupervised manifold alignment
Xiaoyi Yin, Zhen Cui 0001, Hong Chang 0001, Bingpeng Ma, Shiguang Shan |
Sci. China Inf. Sci. | 2 |
| 2022 | Context-dependent emotion recognition
Lingjie Lao, Yong Li 0044, Tong Zhang 0021, Zhen Cui 0001 |
J. Vis. Commun. Image Represent. | 6 |
| 2022 | Direction-induced convolution for point cloud analysis
Chunyan Xu, Chuanwei Zhou, Zhen Cui 0001, Chunlong Hu |
Multim. Syst. | 4 |
| 2022 | Searching part-specific neural fabrics for human pose estimation
Wankou Yang, Zhen Cui 0001 |
Pattern Recognit. | 3 |
| 2022 | From Regional to Global Brain: A Novel Hierarchical Spatial-Temporal Neural Network Model for EEG Emotion RecognitionabstractIn this paper, we propose a novel Electroencephalograph (EEG) emotion recognition method inspired by neuroscience with respect to the brain response to different emotions. The proposed method, denoted by R2G-STNN, consists of spatial and temporal neural network models with regional to global hierarchical feature learning process to learn discriminative spatial-temporal EEG features. To learn the spatial features, a bidirectional long short term memory (BiLSTM) network is adopted to capture the intrinsic spatial relationships of EEG electrodes within brain region and between brain regions, respectively. Considering that different brain regions play different roles in the EEG emotion recognition, a region-attention layer into the R2G-STNN model is also introduced to learn a set of weights to strengthen or weaken the contributions of brain regions. Based on the spatial feature sequences, BiLSTM is adopted to learn both regional and global spatial-temporal features and the features are fitted into a classifier layer for learning emotion-discriminative features, in which a domain discriminator working corporately with the classifier is used to decrease the domain shift between training and testing data. Finally, to evaluate the proposed method, we conduct both subject-dependent and subject-independent EEG emotion recognition experiments on SEED database, and the experimental results show that the proposed method achieves state-of-the-art performance. Yang Li 0019, Wenming Zheng, Lei Wang 0001, Yuan Zong, Zhen Cui 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2022 | LHNet: Laplacian Convolutional Block for Remote Sensing Image Scene ClassificationabstractRecently, many state-of-the-art results for remote sensing image scene classification have been achieved by convolutional neural networks (CNNs) due to their large learning capability. However, in the forward process of CNNs, the high-frequency/texture features are gradually blurred with hierarchical down-sampling and convolution operations. High-frequency features are important to capture the diversity within a class and the similarity between classes. For example, the line features are crucial to distinguish a tennis court from a basketball court. For tennis court in different scenes, the highlight of line features can effectively avoid the influence of diverse background. As a consequence, we propose a Laplacian high-frequency convolutional block (LHCB) based on CNN to extract useful high-frequency features by trainable Laplacian operator. To propagate high-frequency features, we embed LHCB into the existing CNN structures and obtain LHNet. In LHNet, there are two pathways. The original CNN architecture can be taken as the low-frequency pathway and we propose a high-frequency pathway based on LHCB that propagates the residual high-frequency features blurred in each low-frequency layer. Considering that the high-frequency features usually show large variance between images of the same class, we propose a new objective for high-frequency pathway to enhance the intra-class similarity of high-frequency features. The final objective function is obtained by combining the new objective and the baseline classification objective. Numerous experiments on three public available remote sensing image scene classification data sets NWPU-RESISC45, AID and UC Mercerd demonstrate the superior performance of the proposed method. Licheng Jiao, Fang Liu 0001, Jia Liu 0020, Zhen Cui 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Graph Jigsaw Learning for Cartoon Face RecognitionabstractCartoon face recognition is challenging as they typically have smooth color regions and emphasized edges, the key to recognizing cartoon faces is to precisely perceive their sparse and critical shape patterns. However, it is quite difficult to learn a shape-oriented representation for cartoon face recognition with convolutional neural networks (CNNs). To mitigate this issue, we propose the GraphJigsaw that constructs jigsaw puzzles at various stages in the classification network and solves the puzzles with the graph convolutional network (GCN) in a progressive manner. Solving the puzzles requires the model to spot the shape patterns of the cartoon faces as the texture information is quite limited. The key idea of GraphJigsaw is constructing a jigsaw puzzle by randomly shuffling the intermediate convolutional feature maps in the spatial dimension and exploiting the GCN to reason and recover the correct layout of the jigsaw fragments in a self-supervised manner. The proposed GraphJigsaw avoids training the classification model with the deconstructed images that would introduce noisy patterns and are harmful for the final classification. Specially, GraphJigsaw can be incorporated at various stages in a top-down manner within the classification model, which facilitates propagating the learned shape patterns gradually. GraphJigsaw does not rely on any extra manual annotation during the training process and incorporates no extra computation burden at inference time. Both quantitative and qualitative experimental results have verified the feasibility of our proposed GraphJigsaw, which consistently outperforms other face recognition or jigsaw-based methods on two popular cartoon face datasets with considerable improvements. Yong Li 0032, Lingjie Lao, Zhen Cui 0001, Shiguang Shan, Jian Yang 0003 |
IEEE Trans. Image Process. | 3 |
| 2022 | Cross-Database Micro-Expression Recognition: A BenchmarkabstractCross-database micro-expression recognition (CDMER) is one of recently emerging and interesting problem in micro-expression analysis. CDMER is more challenging than the conventional micro-expression recognition (MER), because the training and testing samples in CDMER come from different micro-expression databases, resulting in inconsistency of the feature distributions between the training and testing sets. In this paper, we contribute to this topic from three aspects. First, we establish a CDMER experimental evaluation protocol aiming to allow the researchers to conveniently work on this topic and evaluate their proposed methods under the same standard. Second, we conduct benchmark experiments by using NINE state-of-the-art domain adaptation (DA) methods and SIX popular spatiotemporal descriptors for investigating CDMER problem from two different perspectives. Third, we propose a novel DA method called region selective transfer regression (RSTR) to deal with the CDMER task. The overall superior performance of RSTR over the state-of-the-art DA methods demonstrates that taking into consideration the facial local region information used in RSTR contributes to developing effective DA methods for dealing with CDMER problem. Tong Zhang 0015, Yuan Zong, Wenming Zheng, C. L. Philip Chen, Xiaopeng Hong, Chuangao Tang, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2022 | Self-Teaching Video Object SegmentationabstractVideo object segmentation (VOS) is one of the most fundamental tasks for numerous sequent video applications. The crucial issue of online VOS is the drifting of segmenter when incrementally updated on continuous video frames under unconfident supervision constraints. In this work, we propose a self-teaching VOS (ST-VOS) method to make segmenter to learn online adaptation confidently as much as possible. In the segmenter learning at each time slice, the segment hypothesis and segmenter update are enclosed into a self-looping optimization circle such that they can be mutually improved for each other. To reduce error accumulation of the self-looping process, we specifically introduce a metalearning strategy to learn how to do this optimization within only a few iteration steps. To this end, the learning rates of segmenter are adaptively derived through metaoptimization in the channel space of convolutional kernels. Furthermore, to better launch the self-looping process, we calculate an initial mask map through part detectors and motion flow to well-establish a foundation for subsequent refinement, which could result in the robustness of the segmenter update. Extensive experiments demonstrate that this ST idea can boost the performance of baselines, and in the meantime, our ST-VOS achieves encouraging performance on the DAVIS16, Youtube-objects, DAVIS17, and SegTrackV2 data sets, where, in particular, the accuracy of 75.7% in J-mean metric is obtained on the multi-instance DAVIS17 data set. Chuanwei Zhou, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Graph Game EmbeddingabstractGraph embedding aims to encode nodes/edges into low-dimensional continuous features, and has become a crucial tool for graph analysis including graph/node classification, link prediction, etc. In this paper we propose a novel graph learning framework, named graph game embedding, to learn discriminative node representation as well as encode graph structures. Inspired by the spirit of game learning, node embedding is converted to the selection/searching process of player strategies, where each node corresponds to one player and each edge corresponds to the interaction of two players. Then, a utility function, which theoretically satisfies the Nash Equilibrium, is defined to measure the benefit/loss of players during graph evolution. Furthermore, a collaboration and competition mechanism is introduced to increase the discriminant learning ability. Under this graph game embedding framework, considering different interaction manners of nodes, we propose two specific models, named paired game embedding for paired nodes and group game embedding for group interaction. Comparing with existing graph embedding methods, our algorithm possesses two advantages: (1) the designed utility function ensures the stable graph evolution with theoretical convergence and Nash Equilibrium satisfaction; (2) the introduced collaboration and competition mechanism endows the graph game embedding framework with discriminative feature leaning ability by guiding each node to learn an optimal strategy distinguished from others. We test the proposed method on three public datasets about citation networks, and the experimental results verify the effectiveness of our method. Xiaobin Hong 0002, Tong Zhang 0021, Zhen Cui 0001, Yuge Huang, Pengcheng Shen, Shaoxin Li 0001, Jian Yang 0003 |
AAAI | 3 |
| 2021 | Deep Wasserstein Graph Discriminant Learning for Graph ClassificationabstractGraph topological structures are crucial to distinguish different-class graphs. In this work, we propose a deep Wasserstein graph discriminant learning (WGDL) framework to learn discriminative embeddings of graphs in Wasserstein-metric (W-metric) matching space. In order to bypass the calculation of W-metric class centers in discriminant analysis, as well as better support batch process learning, we introduce a reference set of graphs (aka graph dictionary) to express those representative graph samples (aka dictionary keys). On the bridge of graph dictionary, every input graph can be projected into the latent dictionary space through our proposed Wasserstein graph transformation (WGT). In WGT, we formulate inter-graph distance in W-metric space by virtue of the optimal transport (OT) principle, which effectively expresses the correlations of cross-graph structures. To make WGDL better representation ability, we dynamically update graph dictionary during training by maximizing the ratio of inter-class versus intra-class Wasserstein distance. To evaluate our WGDL method, comprehensive experiments are conducted on six graph classification datasets. Experimental results demonstrate the effectiveness of our WGDL, and state-of-the-art performance. Tong Zhang 0021, Yun Wang 0028, Zhen Cui 0001, Chuanwei Zhou, Baoliang Cui, Haikuan Huang, Jian Yang 0003 |
AAAI | 3 |
| 2021 | Learning Normal Dynamics in Videos With Meta Prototype NetworkabstractFrame reconstruction (current or future frame) based on Auto-Encoder (AE) is a popular method for video anomaly detection. With models trained on the normal data, the reconstruction errors of anomalous scenes are usually much larger than those of normal ones. Previous methods introduced the memory bank into AE, for encoding diverse normal patterns across the training videos. However, they are memory-consuming and cannot cope with unseen new scenarios in the testing data. In this work, we propose a dynamic prototype unit (DPU) to encode the normal dynamics as prototypes in real time, free from extra memory cost. In addition, we introduce meta-learning to our DPU to form a novel few-shot normalcy learner, namely Meta-Prototype Unit (MPU). It enables the fast adaption capability on new scenes by only consuming a few iterations of update. Extensive experiments are conducted on various benchmarks. The superior performance over the state-of-the-art demonstrates the effectiveness of our method. Our code is available at https://github.com/ktr-hubrt/MPN/. Chen Chen 0001, Zhen Cui 0001, Chunyan Xu, Yong Li 0044, Jian Yang 0003 |
CVPR | 3 |
| 2021 | Consistent Instance False Positive Improves Fairness in Face RecognitionabstractDemographic bias is a significant challenge in practical face recognition systems. Existing methods heavily rely on accurate demographic annotations. However, such annotations are usually unavailable in real scenarios. Moreover, these methods are typically designed for a specific demographic group and are not general enough. In this paper, we propose a false positive rate penalty loss, which mitigates face recognition bias by increasing the consistency of instance False Positive Rate (FPR). Specifically, we first define the instance FPR as the ratio between the number of the non-target similarities above a unified threshold and the total number of the non-target similarities. The unified threshold is estimated for a given total FPR. Then, an additional penalty term, which is in proportion to the ratio of instance FPR overall FPR, is introduced into the denominator of the softmax-based loss. The larger the instance FPR, the larger the penalty. By such unequal penalties, the instance FPRs are supposed to be consistent. Compared with the previous debiasing methods, our method requires no demographic annotations. Thus, it can mitigate the bias among demographic groups divided by various attributes, and these attributes are not needed to be previously predefined during training. Extensive experimental results on popular benchmarks demonstrate the superiority of our method over state-of-the-art competitors. Code and pre-trained models are available at https://github.com/xkx0430/FairnessFR. Xingkun Xu, Yuge Huang, Pengcheng Shen, Shaoxin Li 0001, Feiyue Huang, Yong Li 0044, Zhen Cui 0001 |
CVPR | 8 |
| 2021 | Wasserstein Coupled Graph Learning for Cross-Modal RetrievalabstractGraphs play an important role in cross-modal image-text understanding as they characterize the intrinsic structure which is robust and crucial for the measurement of crossmodal similarity. In this work, we propose a Wasserstein Coupled Graph Learning (WCGL) method to deal with the cross-modal retrieval task. First, graphs are constructed according to two input cross-modal samples separately, and passed through the corresponding graph encoders to extract robust features. Then, a Wasserstein coupled dictionary, containing multiple pairs of counterpart graph keys with each key corresponding to one modality, is constructed for further feature learning. Based on this dictionary, the input graphs can be transformed into the dictionary space to facilitate the similarity measurement through a Wasserstein Graph Embedding (WGE) process. The WGE could capture the graph correlation between the input and each corresponding key through optimal transport, and hence well characterize the inter-graph structural relationship. To further achieve discriminant graph learning, we specifically define a Wasserstein discriminant loss on the coupled graph keys to make the intra-class (counterpart) keys more compact and inter-class (non-counterpart) keys more dispersed, which further promotes the final cross-modal retrieval task. Experimental results demonstrate the effectiveness and state-of-the-art performance. Yun Wang 0028, Tong Zhang 0021, Xueya Zhang, Zhen Cui 0001, Yuge Huang, Pengcheng Shen, Shaoxin Li 0001, Jian Yang 0003 |
ICCV | 4 |
| 2021 | Scribble-Supervised Semantic Segmentation InferenceabstractIn this paper, we propose a progressive segmentation inference (PSI) framework to tackle with scribble-supervised semantic segmentation. In virtue of latent contextual dependency, we encapsulate two crucial cues, contextual pattern propagation and semantic label diffusion, to enhance and refine pixel-level segmentation results from partially known seeds. In contextual pattern propagation, different-granular contextual patterns are correlated and leveraged to properly diffuse pattern information based on graphical model, so as to increase the inference confidence of pixel label prediction. Further, depending on high-confidence scores of estimated pixels, the initial annotated seeds are progressively spread over the image through dynamically learning an adaptive decision strategy. The two cues are finally modularized to form a close-looping update process during pixel-wise label inference. Extensive experiments demonstrate that our proposed progressive segmentation inference can benefit from the combination of spatial and semantic context cues, and meantime achieve the state-of-the-art performance on two public scribble segmentation datasets. Jingshan Xu, Chuanwei Zhou, Zhen Cui 0001, Chunyan Xu, Yuge Huang, Pengcheng Shen, Shaoxin Li 0001, Jian Yang 0003 |
ICCV | 3 |
| 2021 | Graph Deformer NetworkabstractConvolution learning on graphs draws increasing attention recently due to its potential applications to a large amount of irregular data. Most graph convolution methods leverage the plain summation/average aggregation to avoid the discrepancy of responses from isomorphic graphs. However, such an extreme collapsing way would result in a structural loss and signal entanglement of nodes, which further cause the degradation of the learning ability. In this paper, we propose a simple yet effective Graph Deformer Network (GDN) to fulfill anisotropic convolution filtering on graphs, analogous to the standard convolution operation on images. Local neighborhood subgraphs (acting like receptive fields) with different structures are deformed into a unified virtual space, coordinated by several anchor nodes. In the deformation process, we transfer components of nodes therein into affinitive anchors by learning their correlations, and build a multi-granularity feature space calibrated with anchors. Anisotropic convolutional kernels can be further performed over the anchor-coordinated space to well encode local variations of receptive fields. By parameterizing anchors and stacking coarsening layers, we build a graph deformer network in an end-to-end fashion. Theoretical analysis indicates its connection to previous work and shows the promising property of graph isomorphism testing. Extensive experiments on widely-used datasets validate the effectiveness of GDN in graph and node classifications. Wenting Zhao 0001, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
IJCAI | 3 |
| 2021 | Transfer Vision Patterns for Multi-Task Pixel LearningabstractMulti-task pixel perception is one of the most important topics in the field of machine intelligence. Inspired by the observation of cross-task interdependencies of visual patterns, we propose a multi-task vision pattern transformation (VPT) method to adaptively correlate and transfer cross-task visual patterns by leveraging the powerful transformer mechanism. To better transfer visual patterns, specifically, we build two types of pattern transformation based on the statistic prior that the affinity relations across tasks are correlated. One aims to transfer feature patterns for the integration of different task features; the other aims to exchange structure patterns for mining and leveraging the latent interaction cues. These two types of transformations are encapsulated into two VPT units, which provide universal matching interfaces for multi-task learning, complement each other to guide the transmission of feature/structure patterns, and finally realize an adaptive selection of important patterns across tasks. Extensive experiments on the joint learning of semantic segmentation, depth prediction and surface normal estimation demonstrate that our proposed method is more effective than those baselines and achieve the state-of-that-art performance in three pixel-level visual tasks. Yong Li 0032, Zhen Cui 0001, Jin Xie 0001, Jian Yang 0003 |
ACM Multimedia | 4 |
| 2021 | Robust and label efficient bi-filtering graph convolutional networks for node classification
Shuaihui Wang, Jin Zhang 0024, Xingyu Zhou 0002, Zhen Cui 0001, Guyu Hu, Zhisong Pan 0003 |
Knowl. Based Syst. | 5 |
| 2021 | A Bi-Hemisphere Domain Adversarial Neural Network Model for EEG Emotion RecognitionabstractIn this paper, we propose a novel neural network model, called bi-hemisphere domain adversarial neural network (BiDANN) model, for electroencephalograph (EEG) emotion recognition. The BiDANN model is inspired by the neuroscience findings that the left and right hemispheres of human's brain are asymmetric to the emotional response. It contains a global and two local domain discriminators that work adversarially with a classifier to learn discriminative emotional features for each hemisphere. At the same time, it tries to reduce the possible domain differences in each hemisphere between the source and target domains so as to improve the generality of the recognition model. In addition, we also propose an improved version of BiDANN, denoted by BiDANN-S, for subject-independent EEG emotion recognition problem by lowering the influences of the personal information of subjects to the EEG emotion recognition. Extensive experiments on the SEED database are conducted to evaluate the performance of both BiDANN and BiDANN-S. The experimental results have shown that the proposed BiDANN and BiDANN models achieve state-of-the-art performance in the EEG emotion recognition. Yang Li 0019, Wenming Zheng, Yuan Zong, Zhen Cui 0001, Tong Zhang 0021 |
IEEE Trans. Affect. Comput. | 4 |
| 2021 | Localizing Anomalies From Weakly-Labeled VideosabstractVideo anomaly detection under video-level labels is currently a challenging task. Previous works have made progresses on discriminating whether a video sequence contains anomalies. However, most of them fail to accurately localize the anomalous events within videos in the temporal domain. In this paper, we propose a Weakly Supervised Anomaly Localization (WSAL) method focusing on temporally localizing anomalous segments within anomalous videos. Inspired by the appearance difference in anomalous videos, the evolution of adjacent temporal segments is evaluated for the localization of anomalous segments. To this end, a high-order context encoding model is proposed to not only extract semantic representations but also measure the dynamic variations so that the temporal context could be effectively utilized. In addition, in order to fully utilize the spatial context information, the immediate semantics are directly derived from the segment representations. The dynamic variations as well as the immediate semantics, are efficiently aggregated to obtain the final anomaly scores. An enhancement strategy is further proposed to deal with noise interference and the absence of localization guidance in anomaly detection. Moreover, to facilitate the diversity requirement for anomaly detection benchmarks, we also collect a new traffic anomaly (TAD) dataset which specifies in the traffic conditions, differing greatly from the current popular anomaly detection evaluation benchmarks. Thedataset and the benchmark test codes, as well as experimental results, are made public on http://vgg-ai.cn/pages/Resource/ and https://github.com/ktr-hubrt/WSAL. Extensive experiments are conducted to verify the effectiveness of different components, and our proposed method achieves new state-of-the-art performance on the UCF-Crime and TAD datasets. Chuanwei Zhou, Zhen Cui 0001, Chunyan Xu, Yong Li 0032, Jian Yang 0003 |
IEEE Trans. Image Process. | 3 |
| 2021 | Meta-VOS: Learning to Adapt Online Target-Specific SegmentationabstractThe task of video object segmentation is a fundamental but challenging problem in the field of computer vision. To deal with large variations in target objects and background clutter, we propose an online adaptive video object segmentation (VOS) framework, named Meta-VOS, that learns to adapt the target-specific segmentation. Meta-VOS builds an online adaptive learning process by exploiting cumulative expertise after searching for confidence patterns across different videos/frames, and then dynamically improves the model learning from two aspects: Meta-seg learner (i.e., module updating) and Meta-seg criterion (i.e., rule of expertise). As our goal is to rapidly determine which patterns best represent the essential characteristics of specific targets in a video, Meta-seg learner is introduced to adaptively learn to update the parameters and hyperparameters of segmentation network in very few gradient descent steps. Furthermore, a Meta-seg criterion of learned expertise, which is constructed to evaluate the Meta-seg learner for the online adaptation of the segmentation network, can confidently online update positive/negative patterns under the guidance of motion cues, object appearances and learned knowledge. Comprehensive evaluations on several benchmark datasets demonstrate the superiority of our proposed Meta-VOS when compared with other state-of-the-art methods applied to the VOS problem. Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
IEEE Trans. Image Process. | 3 |
| 2021 | Dual-Stream Structured Graph Convolution Network for Skeleton-Based Action RecognitionabstractIn this work, we propose a dual-stream structured graph convolution network ( DS-SGCN ) to solve the skeleton-based action recognition problem. The spatio-temporal coordinates and appearance contexts of the skeletal joints are jointly integrated into the graph convolution learning process on both the video and skeleton modalities. To effectively represent the skeletal graph of discrete joints, we create a structured graph convolution module specifically designed to encode partitioned body parts along with their dynamic interactions in the spatio-temporal sequence. In more detail, we build a set of structured intra-part graphs, each of which can be adopted to represent a distinctive body part (e.g., left arm, right leg, head). The inter-part graph is then constructed to model the dynamic interactions across different body parts; here each node corresponds to an intra-part graph built above, while an edge between two nodes is used to express these internal relationships of human movement. We implement the graph convolution learning on both intra- and inter-part graphs in order to obtain the inherent characteristics and dynamic interactions, respectively, of human action. After integrating the intra- and inter-levels of spatial context/coordinate cues, a convolution filtering process is conducted on time slices to capture these temporal dynamics of human motion. Finally, we fuse two streams of graph convolution responses in order to predict the category information of human action in an end-to-end fashion. Comprehensive experiments on five single/multi-modal benchmark datasets (including NTU RGB+D 60, NTU RGB+D 120, MSR-Daily 3D, N-UCLA, and HDM05) demonstrate that the proposed DS-SGCN framework achieves encouraging performance on the skeleton-based action recognition task. Chunyan Xu, Tong Zhang 0021, Zhen Cui 0001, Jian Yang 0003, Chunlong Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2020 | Instance-Adaptive Graph for EEG Emotion RecognitionabstractTo tackle the individual differences and characterize the dynamic relationships among different EEG regions for EEG emotion recognition, in this paper, we propose a novel instance-adaptive graph method (IAG), which employs a more flexible way to construct graphic connections so as to present different graphic representations determined by different input instances. To fit the different EEG pattern, we employ an additional branch to characterize the intrinsic dynamic relationships between different EEG channels. To give a more precise graphic representation, we design the multi-level and multi-graph convolutional operation and the graph coarsening. Furthermore, we present a type of sparse graphic representation to extract more discriminative features. Experiments on two widely-used EEG emotion recognition datasets are conducted to evaluate the proposed model and the experimental results show that our method achieves the state-of-the-art performance. Tengfei Song, Suyuan Liu, Wenming Zheng, Yuan Zong, Zhen Cui 0001 |
AAAI | 5 |
| 2020 | Variational Pathway Reasoning for EEG Emotion RecognitionabstractResearch on human emotion cognition revealed that connections and pathways exist between spatially-adjacent and functional-related areas during emotion expression (Adolphs 2002a; Bullmore and Sporns 2009). Deeply inspired by this mechanism, we propose a heuristic Variational Pathway Reasoning (VPR) method to deal with EEG-based emotion recognition. We introduce random walk to generate a large number of candidate pathways along electrodes. To encode each pathway, the dynamic sequence model is further used to learn between-electrode dependencies. The encoded pathways around each electrode are aggregated to produce a pseudo maximum-energy pathway, which consists of the most important pair-wise connections. To find those most salient connections, we propose a sparse variational scaling (SVS) module to learn scaling factors of pseudo pathways by using the Bayesian probabilistic process and sparsity constraint, where the former endows good generalization ability while the latter favors adaptive pathway selection. Finally, the salient pathways from those candidates are jointly decided by the pseudo pathways and scaling factors. Extensive experiments on EEG emotion recognition demonstrate that the proposed VPR is superior to those state-of-the-art methods, and could find some interesting pathways w.r.t. different emotions. Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu, Wenming Zheng, Jian Yang 0003 |
AAAI | 2 |
| 2020 | Cross-Modal Pattern-Propagation for RGB-T TrackingabstractMotivated by our observations on RGB-T data that pattern correlations are high-frequently recurred across modalities also along sequence frames, in this paper, we propose a cross-modal pattern-propagation (CMPP) tracking framework to diffuse instance patterns across RGB-T data on spatial domain as well as temporal domain. To bridge RGB-T modalities, the cross-modal correlations on intra-modal paired pattern-affinities are derived to reveal those latent cues between heterogenous modalities. Through the correlations, the useful patterns may be mutually propagated between RGB-T modalities so as to fulfill inter-modal pattern-propagation. Further, considering the temporal continuity of sequence frames, we adopt the spirit of pattern propagation to dynamic temporal domain, in which long-term historical contexts are adaptively correlated and propagated into the current frame for more effective information inheritance. Extensive experiments demonstrate that the effectiveness of our proposed CMPP, and the new state-of-the-art results are achieved with the significant improvements on two RGB-T object tracking benchmarks. Chaoqun Wang 0012, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
CVPR | 3 |
| 2020 | Pattern-Structure Diffusion for Multi-Task LearningabstractInspired by the observation that pattern structures high-frequently recur within intra-task also across tasks, we propose a pattern-structure diffusion (PSD) framework to mine and propagate task-specific and task-across pattern structures in the task-level space for joint depth estimation, segmentation and surface normal prediction. To represent local pattern structures, we model them as small-scale graphlets, and propagate them in two different ways, i.e., intra-task and inter-task PSD. For the former, to overcome the limit of the locality of pattern structures, we use the high-order recursive aggregation on neighbors to multiplicatively increase the spread scope, so that long-distance patterns are propagated in the intra-task space. In the inter-task PSD, we mutually transfer the counterpart structures corresponding to the same spatial position into the task itself based on the matching degree of paired pattern structures therein. Finally, the intra-task and inter-task pattern structures are jointly diffused among the task-level patterns, and encapsulated into an end-to-end PSD network to boost the performance of multi-task learning. Extensive experiments on two widely-used benchmarks demonstrate that our proposed PSD is more effective and also achieves the state-of-the-art or competitive results. Zhen Cui 0001, Chunyan Xu, Zhenyu Zhang 0005, Chaoqun Wang 0012, Tong Zhang 0021, Jian Yang 0003 |
CVPR | 2 |
| 2020 | Graph Wasserstein Correlation Analysis for Movie Retrieval
Xueya Zhang, Tong Zhang 0021, Xiaobin Hong 0002, Zhen Cui 0001, Jian Yang 0003 |
ECCV (25) | 4 |
| 2020 | Cross-Graph Convolution Learning for Large-Scale Text-Picture Shopping Guide in E-Commerce SearchabstractIn this work, a new e-commerce search service named text-picture shopping guide (TPSG) is investigated and deployed to one of the most popular shopping platforms called Taobao. Different from traditional services that only contain text options, the TPSG provides pairs of text terms and user-friendly pictures for shopping guide, named text-picture options (TPOs). Instead of manually labeling pictures, we aim to automatically recommend personalized pictures in TPOs. To this end, we build a large-scale graph model on a great amount of data about users, pictures, and terms. Accordingly, a cross-graph convolution learning (CGCL) method is proposed to facilitate the accurate and efficient inference on the constructed graph. To separate the cue of personalized preferences of users to commodities, we factorize the entire mixture-relation graph involving attributes/relations of users and commodities into the user graph, the commodity graph, and the cross user-commodity graph which just characterizes the preferences. Further, we introduce powerful graph convolution to learn more effective representation of these graphs. To reduce the computation burden, specifically, we generalize graph convolution and propose a tensor graph convolution method to learn representation on cross graphs. We conduct extensive offline and online experiments on the large-scale datasets. The results show that the proposed CGCL is very effective and the TPOs recommendation method outperforms manual/advanced selection methods. Tong Zhang 0021, Baoliang Cui, Zhen Cui 0001, Haikuan Huang, Jian Yang 0003, Hongbo Deng, Bo Zheng 0007 |
ICDE | 3 |
| 2020 | Graph inference learning for semi-supervised classification
Chunyan Xu, Zhen Cui 0001, Xiaobin Hong 0002, Tong Zhang 0021, Jian Yang 0003, Wei Liu 0005 |
ICLR | 2 |
| 2020 | Global Information Guided Video Anomaly DetectionabstractVideo anomaly detection (VAD) is currently a challenging task due to the complexity of "anomaly" as well as the lack of labor-intensive temporal annotations. In this paper, we propose an end-to-end Global Information Guided (GIG) anomaly detection framework for anomaly detection using the video-level annotations (i.e., weak labels). We propose to first mine the global pattern cues by leveraging the weak labels in a GIG module. Then we build a spatial reasoning module to measure the relevance between vectors in spatial domain with the global cue vectors, and select the most related feature vectors for temporal anomaly detection. The experimental results on the CityScene challenge demonstrate the effectiveness of our model. Chunyan Xu, Zhen Cui 0001 |
ACM Multimedia | 3 |
| 2020 | Fast Hyper-walk Gridded Convolution on Graph
Xiaobin Hong 0002, Tong Zhang 0021, Zhen Cui 0001, Chunyan Xu, Liangfang Zhang, Jian Yang 0003 |
PRCV (3) | 3 |
| 2020 | Joint Task-Recursive Learning for RGB-D Scene UnderstandingabstractRGB-D scene understanding under monocular camera is an emerging and challenging topic with many potential applications. In this paper, we propose a novel Task-Recursive Learning (TRL) framework to jointly and recurrently conduct three representative tasks therein containing depth estimation, surface normal prediction and semantic segmentation. TRL recursively refines the prediction results through a series of task-level interactions, where one-time cross-task interaction is abstracted as one network block of one time stage. In each stage, we serialize multiple tasks into a sequence and then recursively perform their interactions. To adaptively enhance counterpart patterns, we encapsulate interactions into a specific Task-Attentional Module (TAM) to mutually-boost the tasks from each other. Across stages, the historical experiences of previous states of tasks are selectively propagated into the next stages by using Feature-Selection unit (FS-Unit), which takes advantage of complementary information across tasks. The sequence of task-level interactions is also evolved along a coarse-to-fine scale space such that the required details may be refined progressively. Finally the task-abstracted sequence problem of multi-task prediction is framed into a recursive network. Extensive experiments on NYU-Depth v2 and SUN RGB-D datasets demonstrate that our method can recursively refines the results of the triple tasks and achieves state-of-the-art performance. Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Zequn Jie, Xiang Li 0041, Jian Yang 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | EEG Emotion Recognition Using Dynamical Graph Convolutional Neural NetworksabstractIn this paper, a multichannel EEG emotion recognition method based on a novel dynamical graph convolutional neural networks (DGCNN) is proposed. The basic idea of the proposed EEG emotion recognition method is to use a graph to model the multichannel EEG features and then perform EEG emotion classification based on this model. Different from the traditional graph convolutional neural networks (GCNN) methods, the proposed DGCNN method can dynamically learn the intrinsic relationship between different electroencephalogram (EEG) channels, represented by an adjacency matrix, via training a neural network so as to benefit for more discriminative EEG feature extraction. Then, the learned adjacency matrix is used to learn more discriminative features for improving the EEG emotion recognition. We conduct extensive experiments on the SJTU emotion EEG dataset (SEED) and DREAMER dataset. The experimental results demonstrate that the proposed method achieves better recognition performance than the state-of-the-art methods, in which the average recognition accuracy of 90.4 percent is achieved for subject dependent experiment while 79.95 percent for subject independent cross-validation one on the SEED database, and the average accuracies of 86.23, 84.54 and 85.02 percent are respectively obtained for valence, arousal and dominance classifications on the DREAMER database. Tengfei Song, Wenming Zheng, Peng Song 0002, Zhen Cui 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2020 | Toward Bridging Microexpressions From Different DomainsabstractRecently, microexpression recognition has attracted a lot of researchers' attention due to its challenges and valuable applications. However, it is noticed that currently most of the existing proposed methods are often evaluated and tested on the single database and, hence, this brings us a question whether these methods are still effective if the training and testing samples belong to different domains, for example, different microexpression databases. In this case, a large feature distribution difference may exist between training (source) and testing (target) samples and, hence, microexpression recognition tasks would become more difficult. To solve this challenging problem, that is, cross-domain microexpression recognition, in this paper, we propose an effective method consisting of an auxiliary set selection model (ASSM) and a transductive transfer regression model (TTRM). In our method, an ASSM is designed to automatically select an optimal set of samples from the target domain to serve as the auxiliary set, which is used for subsequent TTRM training. As for TTRM, it aims at bridging the feature distribution gap between the source and target domains by learning a joint regression model with the source domain samples and the auxiliary set selected from the target domain. We evaluate the proposed TTRM plus ASSM by extensive cross-domain microexpression recognition experiments on SMIC and CASME II databases. Compared with the recent state-of-the-art domain adaptation methods, our proposed method has a more satisfactory performance in dealing with the cross-domain microexpression recognition tasks. Yuan Zong, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001, Bin Hu 0001 |
IEEE Trans. Cybern. | 3 |
| 2020 | Hierarchical Semantic Propagation for Object Detection in Remote Sensing ImageryabstractObject detection in remote sensing imagery is a critical yet challenging task in the field of computer vision due to the bird's-eye-view perspective. Although existing object detection approaches in remote sensing imagery have achieved great advances through the utilization of deep features or rotation proposals, but they give insufficient consideration to multilevel semantic information and its propagation for guiding the learning process. Accordingly, in this article, we propose a hierarchical semantic propagation (HSP) framework to boost object detection performance in remote sensing imagery, which is better able to propagate hierarchical semantic information among different components in a unified network. Given a remote sensing image as input, the HSP framework can detect instances of semantic objects belonging to certain categories in an end-to-end way. First, the multiscale representation is captured by a basic feature pyramid network, which can hierarchically combine spatial attention details and the global semantic structure in order to learn more discriminative visual features. Second, the soft-segmentation prediction is used as an auxiliary objective in the intermediate layer of our HSP; its output instance-aware semantic information can be propagated to suppress noisy background information and thereby guide the proposal generation in the region proposal network. By further propagating this hierarchical semantic information into the region of interest module, we can then predict the object category information and the corresponding horizontal and oriented bounding boxes. Comprehensive evaluations on three benchmark data sets demonstrate the superiority of our HSP to the existing state-of-the-art methods for object detection in remote sensing imagery. Chunyan Xu, Chengzheng Li, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2020 | Deep Manifold-to-Manifold Transforming Network for Skeleton-Based Action RecognitionabstractIn this paper, we will investigate skeleton-based action recognition by employing high-order statistics feature and first-order statistics feature, where the high-order statistics feature is characterized by symmetric positive definite (SPD) matrices. Noting that SPD matrices are theoretically embedded on Riemannian manifolds, we propose an end-to-end deep manifold-to-manifold transforming network (DMT-Net), which can make SPD matrices flow from one Riemannian manifold to another one for facilitating the action recognition task. To learn discriminative SPD features from both spatial and temporal dependencies, we propose a neural network model with three novel layers on manifolds: i.e., (1) the local SPD convolutional layer, (2) the non-linear SPD activation layer, and (3) the Riemannian-preserved recursive layer. The SPD property is preserved through all layers without the singular value decomposition (SVD) operation, which has to be conducted in the existing methods with expensive computation cost. Furthermore, a diagonalizing SPD layer is designed to efficiently calculate the final metric for the classification task. Finally, DMT-Net is further fused with a first order layer to capture temporal evolution information. To evaluate our proposed method, we conduct extensive experiments on the task of action recognition, where the input signals are represented as SPD matrices. The experimental results demonstrate that the proposed method is competitive over state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Chaolong Li, Jian Yang 0003 |
IEEE Trans. Multim. | 3 |
| 2020 | Walk-Steered Convolution for Graph ClassificationabstractGraph classification is a fundamental but challenging issue for numerous real-world applications. Despite recent great progress in image/video classification, convolutional neural networks (CNNs) cannot yet cater to graphs well because of graphical non-Euclidean topology. In this article, we propose a walk-steered convolutional (WSC) network to assemble the essential success of standard CNNs, as well as the powerful representation ability of random walk. Instead of deterministic neighbor searching used in previous graphical CNNs, we construct multiscale walk fields (a.k.a. local receptive fields) with random walk paths to depict subgraph structures and advocate graph scalability. To express the internal variations of a walk field, Gaussian mixture models are introduced to encode the principal components of walk paths therein. As an analogy to a standard convolution kernel on image, Gaussian models implicitly coordinate those unordered vertices/nodes and edges in a local receptive field after projecting to the gradient space of Gaussian parameters. We further stack graph coarsening upon Gaussian encoding by using dynamic clustering, such that high-level semantics of graph can be well learned like the conventional pooling on image. The experimental results on several public data sets demonstrate the superiority of our proposed WSC method over many state of the arts for graph classification. Jiatao Jiang, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Wenming Zheng, Jian Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Gaussian-Induced Convolution for GraphsabstractLearning representation on graph plays a crucial role in numerous tasks of pattern recognition. Different from gridshaped images/videos, on which local convolution kernels can be lattices, however, graphs are fully coordinate-free on vertices and edges. In this work, we propose a Gaussianinduced convolution (GIC) framework to conduct local convolution filtering on irregular graphs. Specifically, an edgeinduced Gaussian mixture model is designed to encode variations of subgraph region by integrating edge information into weighted Gaussian models, each of which implicitly characterizes one component of subgraph variations. In order to coarsen a graph, we derive a vertex-induced Gaussian mixture model to cluster vertices dynamically according to the connection of edges, which is approximately equivalent to the weighted graph cut. We conduct our multi-layer graph convolution network on several public datasets of graph classification. The extensive experiments demonstrate that our GIC is effective and can achieve the state-of-the-art results. Jiatao Jiang, Zhen Cui 0001, Chunyan Xu, Jian Yang 0003 |
AAAI | 2 |
| 2019 | Hashing Graph Convolution for Node ClassificationabstractConvolution on graphs has aroused great interest in AI due to its potential applications to non-gridded data. To bypass the influence of ordering and different node degrees, the summation/average diffusion/aggregation is often imposed on local receptive field in most prior works. However, the collapsing into one node in this way tends to cause signal entanglements of nodes, which would result in a sub-optimal feature and decrease the discriminability of nodes. To address this problem, in this paper, we propose a simple but effective Hashing Graph Convolution (HGC) method by using global-hashing and local-projection on node aggregation for the task of node classification. In contrast to the conventional aggregation with a full collision, the hash-projection can greatly reduce the collision probability during gathering neighbor nodes. Another incidental effect of hash-projection is that the receptive field of each node is normalized into a common-size bucket space, which not only staves off the trouble of different-size neighbors and their order but also makes a graph convolution run like the standard shape-gridded convolution. Considering the few training samples, also, we introduce a prediction-consistent regularization term into HGC to constrain the score consistency of unlabeled nodes in the graph. HGC is evaluated on both transductive and inductive experimental settings and achieves new state-of-the-art results on all datasets for node classification task. The extensive experiments demonstrate the effectiveness of hash-projection. Wenting Zhao 0001, Zhen Cui 0001, Chunyan Xu, Chengzheng Li, Tong Zhang 0021, Jian Yang 0003 |
CIKM | 2 |
| 2019 | Pattern-Affinitive Propagation Across Depth, Surface Normal and Semantic SegmentationabstractIn this paper, we propose a novel Pattern-Affinitive Propagation (PAP) framework to jointly predict depth, surface normal and semantic segmentation. The motivation behind it comes from the statistic observation that pattern-affinitive pairs recur much frequently across different tasks as well as within a task. Thus, we can conduct two types of propagations, cross-task propagation and task-specific propagation, to adaptively diffuse those similar patterns. The former integrates cross-task affinity patterns to adapt to each task therein through the calculation on non-local relationships. Next the latter performs an iterative diffusion in the feature space so that the cross-task affinity patterns can be widely-spread within the task. Accordingly, the learning of each task can be regularized and boosted by the complementary task-level affinities. Extensive experiments demonstrate the effectiveness and the superiority of our method on the joint three tasks. Meanwhile, we achieve the state-of-the-art or competitive results on the three related datasets, NYUD-v2, SUN-RGBD and KITTI. Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Yan Yan 0002, Nicu Sebe, Jian Yang 0003 |
CVPR | 2 |
| 2019 | Feature-Attentioned Object Detection in Remote Sensing ImageryabstractIn this work, we introduce a novel feature-attentioned object detection framework to boost its performance in remote sensing imagery, which can focus on learning these intrinsic representations from different aspects in an end-to-end framework. Firstly, when fusing multi-scale visual features of backbone network, we adopt the channel-wise and pixel-wise attentions to enhance these object-related representations and weaken the background/noise information. Secondly, an adaptive multiple receptive fields attention mechanism is employed to generate horizontal region proposals under the special situation where objects in the remote sensing imagery are always with different aspect ratios. Finally, the proposal-level feature attention is proposed to better consider both multi-layer convolutional and apparent representations so that the region of interest network can better predict the object-wise category and its corresponding location information. Comprehensive evaluations on DOTA and UCAS-AOD datasets well demonstrate the effectiveness of our feature-attentioned network for object detection in remote sensing imagery. Chengzheng Li, Chunyan Xu, Zhen Cui 0001, Tong Zhang 0021, Jian Yang 0003 |
ICIP | 3 |
| 2019 | Si-GCN: Structure-induced Graph Convolution Network for Skeleton-based Action RecognitionabstractIn recent years, the graph-convolution networks have been used to solve the problem of skeleton-based action recognition. Previous works often adopted a structure-fixed graph to model the physical joints of human skeleton, but cannot well consider these interactions of different human parts (e.g., the right arm and the left leg) to some extent. To deal with this problem, we propose a novel structure-induced graph convolution network (Si-GCN) framework to boost the performance of the skeleton-based action recognition task. Given a video sequence of human skeletons, the Si-GCN can produce the sample-wise category in an end-to-end way. Specifically, according to the natural divisions of human body, we define a collection of intra-part graphs for each input human skeleton (i.e., each graph denotes a specific part/global of human skeleton), and then formulate an inter-graph to model the relationships of different intra-part graphs. The Si-GCN framework, which will then perform the spectral graph convolutions on these constructed intra/inter-part graphs, can not only capture the internal modalities of each human part/subgraph, but also consider the interactions/relationships between different human parts. A temporal convolution follows to model the temporal and spatial dynamics of the skeleton in combination with the characteristics of time and space. Comprehensive evaluations on two public datasets (including NTU RGB+D and HDM05) well demonstrate the superiority of our proposed Si-GCN when compared with existing skeleton-based action recognition approaches. Chunyan Xu, Tong Zhang 0021, Wenting Zhao 0001, Zhen Cui 0001, Jian Yang 0003 |
IJCNN | 5 |
| 2019 | Cross-Database Micro-Expression Recognition: A BenchmarkabstractCross-database micro-expression recognition (CDMER) is one of recently emerging and interesting problems in micro-expression analysis. CDMER is more challenging than the conventional micro-expression recognition (MER), because the training and testing samples in CDMER come from different micro-expression databases, resulting in inconsistency of the feature distributions between the training and testing sets. In this paper, we contribute to this topic from two aspects. First, we establish a CDMER experimental evaluation protocol and provide a standard platform for evaluating their proposed methods. Second, we conduct extensive benchmark experiments by using NINE state-of-the-art domain adaptation (DA) methods and SIX popular spatiotemporal descriptors for investigating the CDMER problem from two different perspectives and deeply analyze and discuss the experimental results. In addition, all the data and codes involving CDMER in this paper are released on our project website: http://aip.seu.edu.cn/cdmer. Yuan Zong, Wenming Zheng, Xiaopeng Hong, Chuangao Tang, Zhen Cui 0001, Guoying Zhao 0001 |
ICMR | 5 |
| 2019 | EEG Emotion Recognition Based on Graph Regularized Sparse Linear Regression
Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Sheng Ge |
Neural Process. Lett. | 3 |
| 2019 | Recurrent Shape RegressionabstractAn end-to-end network architecture, the Recurrent Shape Regression (RSR), is presented to deal with the task of facial shape detection, a crucial step in many computer vision problems. The RSR generalizes the conventional cascaded regression into a recurrent dynamic network through abstracting common latent models with stage-to-stage operations. Instead of invariant regression transformation, we construct shape-dependent dynamic regressors to attain the recurrence of regression action itself. The regressors can be stacked into a high-order regression network to represent more complex shape regression. By further integrating feature learning as well as global shape constraint, the RSR becomes more controllable in entire optimization of shape regression, where the gradient computation can be efficiently back-propagated through time. To handle the possible partial occlusions of shapes, we propose a mimic virtual occlusion strategy by randomly disturbing certain point cliques without the requirement of any annotations of occlusion information or even occluded training data. Extensive experiments on five face datasets demonstrate that the proposed RSR outperforms the recent state-of-the-art cascaded approaches. Zhen Cui 0001, Shengtao Xiao, Zhiheng Niu, Shuicheng Yan, Wenming Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2019 | Recurrent Face Aging with Hierarchical AutoRegressive MemoryabstractModeling the aging process of human faces is important for cross-age face verification and recognition. In this paper, we propose a Recurrent Face Aging (RFA) framework which takes as input a single image and automatically outputs a series of aged faces. The hidden units in the RFA are connected autoregressively allowing the framework to age the person by referring to the previous aged faces. Due to the lack of labeled face data of the same person captured in a long range of ages, traditional face aging models split the ages into discrete groups and learn a one-step face transformation for each pair of adjacent age groups. Since human face aging is a smooth progression, it is more appropriate to age the face by going through smooth transitional states. In this way, the intermediate aged faces between the age groups can be generated. Towards this target, we employ a recurrent neural network whose recurrent module is a hierarchical triple-layer gated recurrent unit which functions as an autoencoder. The bottom layer of the module encodes the input to a latent representation, and the top layer decodes the representation to a corresponding aged face. The experimental results demonstrate the effectiveness of our framework. Wei Wang 0108, Yan Yan 0002, Zhen Cui 0001, Jiashi Feng, Shuicheng Yan, Nicu Sebe |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | Spatial-Temporal Recurrent Neural Network for Emotion RecognitionabstractIn this paper, we propose a novel deep learning framework, called spatial-temporal recurrent neural network (STRNN), to integrate the feature learning from both spatial and temporal information of signal sources into a unified spatial-temporal dependency model. In STRNN, to capture those spatially co-occurrent variations of human emotions, a multidirectional recurrent neural network (RNN) layer is employed to capture long-range contextual cues by traversing the spatial regions of each temporal slice along different directions. Then a bi-directional temporal RNN layer is further used to learn the discriminative features characterizing the temporal dependencies of the sequences, where sequences are produced from the spatial RNN layer. To further select those salient regions with more discriminative ability for emotion recognition, we impose sparse projection onto those hidden states of spatial and temporal domains to improve the model discriminant ability. Consequently, the proposed two-layer RNN model provides an effective way to make use of both spatial and temporal dependencies of the input signals for emotion recognition. Experimental results on the public emotion datasets of electroencephalogram and facial expression demonstrate the proposed STRNN method is more competitive over those state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Yang Li 0019 |
IEEE Trans. Cybern. | 3 |
| 2019 | Spectral Filter TrackingabstractVisual object tracking is a challenging computer vision task with numerous real-world applications. In this paper, we propose a simple but efficient Spectral Filter Tracking (SFT) method from the view of graph, where each candidate image region is modeled as a pixelwise grid graph. Instead of the conventional graph matching, we formulate the tracking as a plain least square regression problem of learning spectral filters on graphs to predict an optimal vertex, which indicates the center of the target. To bypass computationally expensive eigenvalue decomposition on graph Laplacian L, we parameterize spectral graph filters as a polynomial of L to aggregate local graph features according to spectral graph theory, in which Lk exactly encodes a k-hop local neighborhood of each vertex. Thus, different from the holistic regression in those correlation filter based methods, SFT can operate on localized regions around a pixel (i.e., a vertex), which can effectively reduce the influence of local variations and cluttered backgrounds. Furthermore, we observe that the correlation filter tracking may be viewed as a specific case of our proposed spectral filtering method. The implementation of SFT can simply boil down to only a few line codes, but surprisingly it beats the correlation filter based model with the same feature input, and achieves the state-of-the-art performance on OTB-2015 and VOT2016 under the same feature extraction strategy. Zhen Cui 0001, Youyi Cai, Wenming Zheng, Chunyan Xu, Jian Yang 0003 |
IEEE Trans. Image Process. | 1 |
| 2019 | ℓ1-Norm Heteroscedastic Discriminant Analysis Under Mixture of Gaussian DistributionsabstractFisher’s criterion is one of the most popular discriminant criteria for feature extraction. It is defined as the generalized Rayleigh quotient of the between-class scatter distance to the within-class scatter distance. Consequently, Fisher’s criterion does not take advantage of the discriminant information in the class covariance differences, and hence, its discriminant ability largely depends on the class mean differences. If the class mean distances are relatively large compared with the within-class scatter distance, Fisher’s criterion-based discriminant analysis methods may achieve a good discriminant performance. Otherwise, it may not deliver good results. Moreover, we observe that the between-class distance of Fisher’s criterion is based on the$\ell _{2}$-norm, which would be disadvantageous to separate the classes with smaller class mean distances. To overcome the drawback of Fisher’s criterion, in this paper, we first derive a new discriminant criterion, expressed as amixture of absolute generalized Rayleigh quotients, based on a Bayes error upper bound estimation, where mixture of Gaussians is adopted to approximate the real distribution of data samples. Then, the criterion is further modified by replacing$\ell _{2}$-norm with$\ell _{1}$one to better describe the between-class scatter distance, such that it would be more effective to separate the different classes. Moreover, we propose a novel$\ell _{1}$-norm heteroscedastic discriminant analysis method based on the new discriminant analysis (L1-HDA/GM) for heteroscedastic feature extraction, in which the optimization problem of L1-HDA/GM can be efficiently solved by using the eigenvalue decomposition approach. Finally, we conduct extensive experiments on four real data sets and demonstrate that the proposed method achieves much competitive results compared with the state-of-the-art methods. Wenming Zheng, Cheng Lu 0005, Zhouchen Lin, Tong Zhang 0021, Zhen Cui 0001, Wankou Yang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2018 | Spatio-Temporal Graph Convolution for Skeleton Based Action RecognitionabstractVariations of human body skeletons may be considered as dynamic graphs, which are generic data representation for numerous real-world applications. In this paper, we propose a spatio-temporal graph convolution (STGC) approach for assembling the successes of local convolutional filtering and sequence learning ability of autoregressive moving average. To encode dynamic graphs, the constructed multi-scale local graph convolution filters, consisting of matrices of local receptive fields and signal mappings, are recursively performed on structured graph data of temporal and spatial domain. The proposed model is generic and principled as it can be generalized into other dynamic models. We theoretically prove the stability of STGC and provide an upper-bound of the signal transformation to be learnt. Further, the proposed recursive model can be stacked into a multi-layer architecture. To evaluate our model, we conduct extensive experiments on four benchmark skeleton-based action datasets, including the large-scale challenging NTU RGB+D. The experimental results demonstrate the effectiveness of our proposed model and the improvement over the state-of-the-art. Chaolong Li, Zhen Cui 0001, Wenming Zheng, Chunyan Xu, Jian Yang 0003 |
AAAI | 2 |
| 2018 | Joint Task-Recursive Learning for Semantic Segmentation and Depth Estimation
Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Zequn Jie, Xiang Li 0041, Jian Yang 0003 |
ECCV (10) | 2 |
| 2018 | Action Recognition with Spatial-Temporal Representation Analysis Across Grassmannian Manifold and Euclidean SpaceabstractAction recognition plays an important character for numerous tasks of video area. Although previous works often learn the appearance and motion information with Convolutional Neural Networks (CNNs), they ignore the corresponding space structures of video representation. In this work, we address action recognition task with a Spatial-Temporal representation analysis algorithm Across Grassmannian manifold and Euclidean space (ST-AGE), which considers the appearance and motion information of video samples in an unified framework. For each video sample, we extract temporal features with classical CNNs (e.g., ConvNet, VGG, ResNet) and motion representation with the trajectory tracking method. Both spatial and temporal information can be then analyzed by embedding them on the Grassmannian manifold and Euclidean space, and an appropriate multi-kernel SVM is further conducted. Comprehensive evaluations on HMDB-51 and UCF-101 datasets demonstrate the significant superiority of STAGE over other state-of-the-art for human action recognition. Xinshu Qiao, Chuanwei Zhou, Chunyan Xu, Zhen Cui 0001, Jian Yang 0003 |
ICIP | 4 |
| 2018 | Deep Manifold-to-Manifold Transforming NetworkabstractIn this paper, we propose an end-to-end deep manifold-to-manifold transforming network (DMT-Net), which makes SPD matrices flow from one Riemannian manifold to another more discriminative one. For discriminative feature learning, two specific layers on manifolds are developed: (i) the local SPD convolutional layer, (ii) the non-linear SPD activation layer, where positive definiteness is satisfied for both two layers. Further, to relieve computational burden of kernels on relative large-scale data, we design a batch-kernelized layer to favor batchwise kernel optimization of deep networks. Specifically, one reference set dynamically changing with the network training is introduced to break the limitation of memory size. We evaluate our proposed method on action recognition datasets, where input signals are popularly modeled as SPD matrices. The experimental results demonstrate that our DMT-Net is more competitive than state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Chaolong Li |
ICIP | 3 |
| 2018 | A Novel Neural Network Model based on Cerebral Hemispheric Asymmetry for EEG Emotion RecognitionabstractIn this paper, we propose a novel neural network model, called bi-hemispheres domain adversarial neural network (BiDANN), for EEG emotion recognition. BiDANN is motivated by the neuroscience findings, i.e., the emotional brain's asymmetries between left and right hemispheres. The basic idea of BiDANN is to map the EEG feature data of both left and right hemispheres into discriminative feature spaces separately, in which the data representations can be classified easily. For further precisely predicting the class labels of testing data, we narrow the distribution shift between training and testing data by using a global and two local domain discriminators, which work adversarially to the classifier to encourage domain-invariant data representations to emerge. After that, the learned classifier from labeled training data can be applied to unlabeled testing data naturally. We conduct two experiments to verify the performance of our BiDANN model on SEED database. The experimental results show that the proposed model achieves the state-of-the-art performance. Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Tong Zhang 0021, Yuan Zong |
IJCAI | 3 |
| 2018 | Context-Dependent Diffusion Network for Visual Relationship DetectionabstractVisual relationship detection can bridge the gap between computer vision and natural language for scene understanding of images. Different from pure object recognition tasks, the relation triplets of subject-predicate-object lie on an extreme diversity space, such asperson-behind-person andcar-behind-building, while suffering from the problem of combinatorial explosion. In this paper, we propose a context-dependent diffusion network (CDDN) framework to deal with visual relationship detection. To capture the interactions of different object instances, two types of graphs, word semantic graph and visual scene graph, are constructed to encode global context interdependency. The semantic graph is built through language priors to model semantic correlations across objects, whilst the visual scene graph defines the connections of scene objects so as to utilize the surrounding scene information. For the graph-structured data, we design a diffusion network to adaptively aggregate information from contexts, which can effectively learn latent representations of visual relationships and well cater to visual relationship detection in view of its isomorphic invariance to graphs. Experiments on two widely-used datasets demonstrate that our proposed method is more effective and achieves the state-of-the-art performance. Zhen Cui 0001, Chunyan Xu, Wenming Zheng, Jian Yang 0003 |
ACM Multimedia | 1 |
| 2018 | Fusing magnitude and phase features with multiple face models for robust face recognition
Yan Li 0014, Shiguang Shan, Ruiping Wang 0001, Zhen Cui 0001, Xilin Chen 0001 |
Frontiers Comput. Sci. | 4 |
| 2018 | Face recognition based on recurrent regression neural network
Yang Li 0019, Wenming Zheng, Zhen Cui 0001, Tong Zhang 0021 |
Neurocomputing | 3 |
| 2018 | Multi-cue fusion for emotion recognition in the wild
Jingwei Yan, Wenming Zheng, Zhen Cui 0001, Chuangao Tang, Tong Zhang 0021, Yuan Zong |
Neurocomputing | 3 |
| 2018 | Unsupervised facial expression recognition using domain adaptation based dictionary learning approach
Wenming Zheng, Zhen Cui 0001, Yuan Zong, Tong Zhang 0021, Chuangao Tang |
Neurocomputing | 3 |
| 2018 | Deep Recurrent Regression for Facial Landmark DetectionabstractWe propose a novel end-to-end deep architecture for face landmark detection, based on a deep convolutional and deconvolutional network followed by carefully designed recurrent network structures. The pipeline of this architecture consists of three parts. Through the first part, we encode an input face image to resolution-preserved deconvolutional feature maps via a deep network with stacked convolutional and deconvolutional layers. Then, in the second part, we estimate the initial coordinates of the facial key points by an additional convolutional layer on top of these deconvolutional feature maps. In the last part, by using the deconvolutional feature maps and the initial facial key points as input, we refine the coordinates of the facial key points by a recurrent network that consists of multiple long short-term memory components. Extensive evaluations on several benchmark data sets show that the proposed deep architecture has superior performance against the state-of-the-art methods. Hanjiang Lai, Shengtao Xiao, Yan Pan 0002, Zhen Cui 0001, Jiashi Feng, Chunyan Xu, Jian Yin 0001, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2018 | Action-Attending Graphic Neural NetworkabstractThe motion analysis of human skeletons is crucial for human action recognition, which is one of the most active topics in computer vision. In this paper, we propose a fully end-to-end action-attending graphic neural network (A2GNN) for skeleton-based action recognition, in which each irregular skeleton is structured as an undirected attribute graph. To extract high-level semantic representation from skeletons, we perform the local spectral graph filtering on the constructed attribute graphs like the standard image convolution operation. Considering not all joints are informative for action analysis, we design an actionattending layer to detect those salient action units (AUs) by adaptively weighting skeletal joints. Herein the filtering responses are parameterized into a weighting function irrelevant to the order of input nodes. To further encode continuous motion variations, the deep features learnt from skeletal graphs are gathered along consecutive temporal slices and then fed into a recurrent gated network. Finally, the spectral graph filtering, action-attending and recurrent temporal encoding are integrated together to jointly train for the sake of robust action recognition as well as the intelligibility of human actions. To evaluate our A2GNN, we conduct extensive experiments on four benchmark skeletonbased action datasets, including the large-scale challenging NTU RGB+D dataset. The experimental results demonstrate that our network achieves the state-of-the-art performances. Chaolong Li, Zhen Cui 0001, Wenming Zheng, Chunyan Xu, Rongrong Ji, Jian Yang 0003 |
IEEE Trans. Image Process. | 2 |
| 2018 | Progressive Hard-Mining Network for Monocular Depth EstimationabstractDepth estimation from the monocular RGB image is a challenging task for computer vision due to no reliable cues as the prior knowledge. Most existing monocular depth estimation works including various geometric or network learning methods lack of an effective mechanism to preserve the cross-border details of depth maps, which yet is very important for the performance promotion. In this paper, we propose a novel end-to-end progressive hard-mining network (PHN) framework to address this problem. Specifically, we construct the hard-mining objective function, the intra-scale and inter-scale refinement subnetworks to accurately localize and refine those hard-mining regions. The intra-scale refining block recursively recovers details of depth maps from different semantic features in the same receptive field while the inter-scale block favors a complementary interaction among multi-scale depth cues of different receptive fields. For further reducing the uncertainty of the network, we design a difficulty-ware refinement loss function to guide the depth learning process, which can adaptively focus on mining these hard-regions where accumulated errors easily occur. All three modules collaborate together to progressively reduce the error propagation in the depth learning process, and then, boost the performance of monocular depth estimation to some extent. We conduct comprehensive evaluations on several public benchmark data sets (including NYU Depth V2, KITTI, and Make3D). The experiment results well demonstrate the superiority of our proposed PHN framework over other state of the arts for monocular depth estimation task. Zhenyu Zhang 0005, Chunyan Xu, Jian Yang 0003, Junbin Gao, Zhen Cui 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Domain Regeneration for Cross-Database Micro-Expression RecognitionabstractRecently, micro-expression recognition has attracted lots of researchers' attention due to its potential value in many practical applications, e.g., lie detection. In this paper, we investigate an interesting and challenging problem in micro-expression recognition, i.e., cross-database micro-expression recognition, in which the training and testing samples come from different micro-expression databases. Under this problem setting, the consistent feature distribution between the training and testing samples originally existing in conventional micro-expression recognition would be seriously broken and hence the performance of most current well-performing micro-expression recognition methods may sharply drop. In order to overcome it, we propose a simple yet effective framework called Domain Regeneration (DR) in this paper. DR framework aims at learning a domain regenerator to regenerate the micro-expression samples from source and target databases respectively such that they can abide by the same or similar feature distributions. Thus, we are able to use the classifier learned based on the labeled source micro-expression samples to predict the label information of the unlabeled target micro-expression samples. To evaluate the proposed DR framework, we conduct extensive cross-database micro-expression recognition experiments designed based on SMIC and CASME II databases. Experimental results show that compared with recent state-of-the-art cross-database emotion recognition methods, the proposed DR framework has more promising performance. Yuan Zong, Wenming Zheng, Xiaohua Huang 0003, Jingang Shi, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Learning From Hierarchical Spatiotemporal Descriptors for Micro-Expression RecognitionabstractMicro-expression recognition aims to infer genuine emotions that people try to conceal from facial video clips. It is a very challenging task because micro-expressions have a very low intensity and short duration, which makes micro-expressions difficult to observe. Recently, researchers have designed various spatiotemporal descriptors to describe micro-expressions. It is notable that for better capturing the low-intensity facial muscle movement, a fixed spatial division grid, 8× 8 for example, is commonly used to partition the facial images into a few facial blocks before extracting descriptors. However, it is hard to choose an ideal division grid for different micro-expression samples because the division grids affect the discriminative ability of spatiotemporal descriptors to distinguish micro-expressions. To address this problem, in this paper, we design a hierarchical spatial division scheme for spatiotemporal descriptor extraction. By using the proposed scheme, it would not be a problem to determine which division grid is most suitable regarding different micro-expression samples. Furthermore, we propose a kernelized group sparse learning (KGSL) model to process hierarchical scheme based spatiotemporal descriptors such that they are more effective for micro-expression recognition tasks. To evaluate the performance of the proposed micro-expression recognition method consisting of the hierarchical scheme based spatiotemporal descriptors and KGSL, extensive experiments are conducted on two public micro-expression databases: CASME II and SMIC. Compared with many recent state-of-the-art approaches, our method achieves more promising recognition results. Yuan Zong, Xiaohua Huang 0003, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001 |
IEEE Trans. Multim. | 4 |
| 2017 | View-Independent Facial Action Unit DetectionabstractAutomatic Facial Action Unit (AU) detection has drawn more and more attention over the past years due to its significance to facial expression analysis. Frontal-view AU detection has been extensively evaluated, but cross-pose AU detection is a less-touched problem due to the scarcity of the related dataset. The challenge of Facial Expression Recognition and Analysis (FERA2017) just released a large-scale videobased AU detection dataset across different facial poses. To deal with this challenging task, we develop a simple and efficient deep learning based system to detect AU occurrence under nine different facial views. In this system, we first crop out facial images by using morphology operations including binary segmentation, connected components labeling and region boundaries extraction, then for each type of AU, we train a corresponding expert network by specifically fine-tuning the VGG-Face network on cross-view facial images, so as to extract more discriminative features for the subsequent binary classification. In the AU detection sub-challenge, our proposed method achieves the mean accuracy of 77.8% (vs. the baseline 56.1%), and promotes the F1 score to 57.4% (vs. the baseline 45.2%). Chuangao Tang, Wenming Zheng, Jingwei Yan, Qiang Li 0044, Yang Li 0019, Tong Zhang 0021, Zhen Cui 0001 |
FG | 7 |
| 2017 | Learning a Target Sample Re-Generator for Cross-Database Micro-Expression RecognitionabstractIn this paper, we investigate the cross-database micro-expression recognition problem, where the training and testing samples are from two different micro-expression databases. Under this setting, the training and testing samples would have different feature distributions and hence the performance of most existing micro-expression recognition methods may decrease greatly. To solve this problem, we propose a simple yet effective method called Target Sample Re-Generator (TSRG) in this paper. By using TSRG, we are able to re-generate the samples from target micro-expression database and the re-generated target samples would share same or similar feature distributions with the original source samples. For this reason, we can then use the classifier learned based on the labeled source samples to accurately predict the micro-expression categories of the unlabeled target samples. To evaluate the performance of the proposed TSRG method, extensive cross-database micro-expression recognition experiments designed based on SMIC and CASME II databases are conducted. Compared with recent state-of-the-art cross-database emotion recognition methods, the proposed TSRG achieves more promising results. Yuan Zong, Xiaohua Huang 0003, Wenming Zheng, Zhen Cui 0001, Guoying Zhao 0001 |
ACM Multimedia | 4 |
| 2017 | Layerwise Class-Aware Convolutional Neural NetworkabstractThe human vision system usually has a specifically activated area of neurons when recognizing a category of images. Inspired by this visual mechanism, we propose a layerwise class-aware convolutional neural network architecture to explicitly discover category-tailored neurons on intermediate hidden layers to improve the network learning ability. Instead of directly selecting activated neurons for different categories, we inversely suppress those neurons of intermediate layers irrelevant with the given target class to produce a class-specific subnetwork, which implicitly enhances the discriminability of hidden layer features due to the increase of the inter-class discrepancy on them. Together with the classifier of the top layer, we jointly learn this network by formulating the suppressor of hidden layers as a penalty term in the objective function. To address class-specific neuron suppression in each hidden layer, we also introduce a statistic method based on mutual information to dynamically and automatically update the suppressed neurons during the network training. Extensive experiments demonstrate that the proposed model is superior to the state-of-the-art models. Zhen Cui 0001, Zhiheng Niu, Luoqi Liu, Shuicheng Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2016 | Recurrently Target-Attending TrackingabstractRobust visual tracking is a challenging task in computer vision. Due to the accumulation and propagation of estimation error, model drifting often occurs and degrades the tracking performance. To mitigate this problem, in this paper we propose a novel tracking method called Recurrently Target-attending Tracking (RTT). RTT attempts to identify and exploit those reliable parts which are beneficial for the overall tracking process. To bypass occlusion and discover reliable components, multi-directional Recurrent Neural Networks (RNNs) are employed in RTT to capture long-range contextual cues by traversing a candidate spatial region from multiple directions. The produced confidence maps from the RNNs are employed to adaptively regularize the learning of discriminative correlation filters by suppressing clutter background noises while making full use of the information from reliable parts. To solve the weighted correlation filters, we especially derive an efficient closedform solution with a sharp reduction in computation complexity. Extensive experiments demonstrate that our proposed RTT is more competitive over those correlation filter based methods. Zhen Cui 0001, Shengtao Xiao, Jiashi Feng, Shuicheng Yan |
CVPR | 1 |
| 2016 | Recurrent Face AgingabstractModeling the aging process of human face is important for cross-age face verification and recognition. In this paper, we introduce a recurrent face aging (RFA) framework based on a recurrent neural network which can identify the ages of people from 0 to 80. Due to the lack of labeled face data of the same person captured in a long range of ages, traditional face aging models usually split the ages into discrete groups and learn a one-step face feature transformation for each pair of adjacent age groups. However, those methods neglect the in-between evolving states between the adjacent age groups and the synthesized faces often suffer from severe ghosting artifacts. Since human face aging is a smooth progression, it is more appropriate to age the face by going through smooth transition states. In this way, the ghosting artifacts can be effectively eliminated and the intermediate aged faces between two discrete age groups can also be obtained. Towards this target, we employ a twolayer gated recurrent unit as the basic recurrent module whose bottom layer encodes a young face to a latent representation and the top layer decodes the representation to a corresponding older face. The experimental results demonstrate our proposed RFA provides better aging faces over other state-of-the-art age progression methods. Wei Wang 0108, Zhen Cui 0001, Yan Yan 0002, Jiashi Feng, Shuicheng Yan, Xiangbo Shu, Nicu Sebe |
CVPR | 2 |
| 2016 | Multi-clue fusion for emotion recognition in the wildabstractIn the past three years, Emotion Recognition in the Wild (EmotiW) Grand Challenge has drawn more and more attention due to its huge potential applications. In the fourth challenge, aimed at the task of video based emotion recognition, we propose a multi-clue emotion fusion (MCEF) framework by modeling human emotion from three mutually complementary sources, facial appearance texture, facial action, and audio. To extract high-level emotion features from sequential face images, we employ a CNN-RNN architecture, where face image from each frame is first fed into the fine-tuned VGG-Face network to extract face feature, and then the features of all frames are sequentially traversed in a bidirectional RNN so as to capture dynamic changes of facial textures. To attain more accurate facial actions, a facial landmark trajectory model is proposed to explicitly learn emotion variations of facial components. Further, audio signals are also modeled in a CNN framework by extracting low-level energy features from segmented audio clips and then stacking them as an image-like map. Finally, we fuse the results generated from three clues to boost the performance of emotion recognition. Our proposed MCEF achieves an overall accuracy of 56.66% with a large improvement of 16.19% with respect to the baseline. Jingwei Yan, Wenming Zheng, Zhen Cui 0001, Chuangao Tang, Tong Zhang 0021, Yuan Zong, Ning Sun 0001 |
ICMI | 3 |
| 2016 | A Novel Graph Regularized Sparse Linear Discriminant Analysis Model for EEG Emotion Recognition
Yang Li 0019, Wenming Zheng, Zhen Cui 0001 |
ICONIP (4) | 3 |
| 2016 | Cross-Database Facial Expression Recognition via Unsupervised Domain Adaptive Dictionary Learning
Wenming Zheng, Zhen Cui 0001, Yuan Zong |
ICONIP (2) | 3 |
| 2016 | Higher order partial least squares for object tracking: A 4D-tracking method
Bineng Zhong 0001, Xiangnan Yang, Yingju Shen, Cheng Wang 0020, Tian Wang 0001, Zhen Cui 0001, Hongbo Zhang 0002, Xiaopeng Hong, Duansheng Chen |
Neurocomputing | 6 |
| 2016 | Spatial Pyramid Covariance-Based Compact Video Code for Robust Face Retrieval in TV-SeriesabstractWe address the problem of face video retrieval in TV-series, which searches video clips based on the presence of specific character, given one face track of his/her. This is tremendously challenging because on one hand, faces in TV-series are captured in largely uncontrolled conditions with complex appearance variations, and on the other hand, retrieval task typically needs efficient representation with low time and space complexity. To handle this problem, we propose a compact and discriminative representation for the huge body of video data, named compact video code (CVC). Our method first models the face track by its sample (i.e., frame) covariance matrix to capture the video data variations in a statistical manner. To incorporate discriminative information and obtain more compact video signature suitable for retrieval, the high-dimensional covariance representation is further encoded as a much lower dimensional binary vector, which finally yields the proposed CVC. Specifically, each bit of the code, i.e., each dimension of the binary vector, is produced via supervised learning in a max margin framework, which aims to make a balance between the discriminability and stability of the code. Besides, we further extend the descriptive granularity of covariance matrix from traditional pixel-level to more general patch-level, and proceed to propose a novel hierarchical video representation named spatial pyramid covariance along with a fast calculation method. Face retrieval experiments on two challenging TV-series video databases, i.e., the Big Bang Theory and Prison Break, demonstrate the competitiveness of the proposed CVC over the state-of-the-art retrieval methods. In addition, as a general video matching algorithm, CVC is also evaluated in traditional video face recognition task on a standard Internet database, i.e., YouTube Celebrities, showing its quite promising performance by using an extremely compact code with only 128 bits. Yan Li 0014, Ruiping Wang 0001, Zhen Cui 0001, Shiguang Shan, Xilin Chen 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | A Deep Neural Network-Driven Feature Learning Method for Multi-view Facial Expression RecognitionabstractIn this paper, a novel deep neural network (DNN)-driven feature learning method is proposed and applied to multi-view facial expression recognition (FER). In this method, scale invariant feature transform (SIFT) features corresponding to a set of landmark points are first extracted from each facial image. Then, a feature matrix consisting of the extracted SIFT feature vectors is used as input data and sent to a well-designed DNN model for learning optimal discriminative features for expression classification. The proposed DNN model employs several layers to characterize the corresponding relationship between the SIFT feature vectors and their corresponding high-level semantic information. By training the DNN model, we are able to learn a set of optimal features that are well suitable for classifying the facial expressions across different facial views. To evaluate the effectiveness of the proposed method, two nonfrontal facial expression databases, namely BU-3DFE and Multi-PIE, are respectively used to testify our method and the experimental results show that our algorithm outperforms the state-of-the-art methods. Tong Zhang 0021, Wenming Zheng, Zhen Cui 0001, Yuan Zong, Jingwei Yan |
IEEE Trans. Multim. | 3 |
| 2015 | Sparsely encoded local descriptor for face verification
Zhen Cui 0001, Shiguang Shan, Ruiping Wang 0001, Lei Zhang 0006, Xilin Chen 0001 |
Neurocomputing | 1 |
| 2015 | Online learning 3D context for robust visual tracking
Bineng Zhong 0001, Yingju Shen, Yan Chen 0017, Weibo Xie, Zhen Cui 0001, Hongbo Zhang 0002, Duansheng Chen, Tian Wang 0001, Xin Liu 0011, Shu-Juan Peng, Jin Gou, Jixiang Du, Jing Wang 0049, Wenming Zheng |
Neurocomputing | 5 |
| 2014 | Representation Learning with Smooth Autoencoder
Kongming Liang, Hong Chang 0001, Zhen Cui 0001, Shiguang Shan, Xilin Chen 0001 |
ACCV (2) | 3 |
| 2014 | Compact Video Code and Its Application to Robust Face Retrieval in TV-Series
Yan Li 0014, Ruiping Wang 0001, Zhen Cui 0001, Shiguang Shan, Xilin Chen 0001 |
BMVC | 3 |
| 2014 | Deep Network Cascade for Image Super-resolution
Zhen Cui 0001, Hong Chang 0001, Shiguang Shan, Bineng Zhong 0001, Xilin Chen 0001 |
ECCV (5) | 1 |
| 2014 | Generalized Unsupervised Manifold Alignment
Zhen Cui 0001, Hong Chang 0001, Shiguang Shan, Xilin Chen 0001 |
NIPS | 1 |
| 2014 | Joint sparse representation for video-based face recognition
Zhen Cui 0001, Hong Chang 0001, Shiguang Shan, Bingpeng Ma, Xilin Chen 0001 |
Neurocomputing | 1 |
| 2014 | Robust tracking via patch-based appearance model and local background estimation
Bineng Zhong 0001, Yan Chen 0017, Yingju Shen, Yewang Chen, Zhen Cui 0001, Rongrong Ji, Xiao-Tong Yuan, Duansheng Chen |
Neurocomputing | 5 |
| 2014 | Structured partial least squares for simultaneous object tracking and segmentation
Bineng Zhong 0001, Xiao-Tong Yuan, Rongrong Ji, Yan Yan 0001, Zhen Cui 0001, Xiaopeng Hong, Yan Chen 0017, Tian Wang 0001, Duansheng Chen |
Neurocomputing | 5 |
| 2014 | Automatic motion capture data denoising via filtered subspace clustering and low rank matrix approximation
Xin Liu 0011, Yiu-Ming Cheung, Shu-Juan Peng, Zhen Cui 0001, Bineng Zhong 0001, Jixiang Du |
Signal Process. | 4 |
| 2014 | Flowing on Riemannian Manifold: Domain Adaptation by Shifting CovarianceabstractDomain adaptation has shown promising results in computer vision applications. In this paper, we propose a new unsupervised domain adaptation method called domain adaptation by shifting covariance (DASC) for object recognition without requiring any labeled samples from the target domain. By characterizing samples from each domain as one covariance matrix, the source and target domain are represented into two distinct points residing on a Riemannian manifold. Along the geodesic constructed from the two points, we then interpolate some intermediate points (i.e., covariance matrices), which are used to bridge the two domains. By utilizing the principal components of each covariance matrix, samples from each domain are further projected into intermediate feature spaces, which finally leads to domain-invariant features after the concatenation of these features from intermediate points. In the multiple source domain adaptation task, we also need to effectively integrate different types of features between each pair of source and target domains. We additionally propose an SVM based method to simultaneously learn the optimal target classifier as well as the optimal weights for different source domains. Extensive experiments demonstrate the effectiveness of our method for both single source and multiple source domain adaptation tasks. Zhen Cui 0001, Wen Li 0001, Dong Xu 0001, Shiguang Shan, Xilin Chen 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2013 | Automatic Motion Capture Data Denoising via Filtered Local Subspace Affinity and Low Rank ApproximationabstractIn this paper, we formulate the Motion capture (MoCap) data denoising problem as the concatenation of piecewise motion matrix recovery problem, in which the moving trajectories of each piecewise motion always share the similar subspace representation. To this end, we present an automatic MoCap data denoising approach based on the filtered local subspace affinity (LSA) and low rank approximation. The proposed approach does not need any physical information about the underling structure of MoCap data or require auxiliary data sets for the training priors. The experiments have shown the promising results. Shu-Juan Peng, Xin Liu 0011, Zhen Cui 0001, Zhipeng Xie, Duansheng Chen |
CAD/Graphics | 3 |
| 2013 | Fusing Robust Face Region Descriptors via Multiple Metric Learning for Face Recognition in the WildabstractIn many real-world face recognition scenarios, face images can hardly be aligned accurately due to complex appearance variations or low-quality images. To address this issue, we propose a new approach to extract robust face region descriptors. Specifically, we divide each image (resp. video) into several spatial blocks (resp. spatial-temporal volumes) and then represent each block (resp. volume) by sum-pooling the nonnegative sparse codes of position-free patches sampled within the block (resp. volume). Whitened Principal Component Analysis (WPCA) is further utilized to reduce the feature dimension, which leads to our Spatial Face Region Descriptor (SFRD) (resp. Spatial-Temporal Face Region Descriptor, STFRD) for images (resp. videos). Moreover, we develop a new distance metric learning method for face verification called Pairwise-constrained Multiple Metric Learning (PMML) to effectively integrate the face region descriptors of all blocks (resp. volumes) from an image (resp. a video). Our work achieves the state-of-the-art performances on two real-world datasets LFW and YouTube Faces (YTF) according to the restricted protocol. Zhen Cui 0001, Wen Li 0001, Dong Xu 0001, Shiguang Shan, Xilin Chen 0001 |
CVPR | 1 |
| 2012 | Image sets alignment for Video-Based Face RecognitionabstractVideo-based Face Recognition (VFR) can be converted to the matching of two image sets containing face images captured from each video. For this purpose, we propose to bridge the two sets with a reference image set that is well-defined and pre-structured to a number of local models offline. In other words, given two image sets, as long as each of them is aligned to the reference set, they are mutually aligned and well structured. Therefore, the similarity between them can be computed by comparing only the corresponded local models rather than considering all the pairs. To align an image set with the reference set, we further formulate the problem as a quadratic programming. It integrates three constrains to guarantee robust alignment, including appearance matching cost term exploiting principal angles, geometric structure consistency using affine invariant reconstruction weights, smoothness constraint preserving local neighborhood relationship. Extensive experimental evaluations are performed on three databases: Honda, MoBo and YouTube. Compared with competing methods, our approach can consistently achieve better results. Zhen Cui 0001, Shiguang Shan, Haihong Zhang, Shihong Lao, Xilin Chen 0001 |
CVPR | 1 |
| 2012 | Structured Sparse Linear Discriminant AnalysisabstractLinear Discriminant Analysis (LDA) is an efficient image feature extraction technique by supervised dimensionality reduction. In this paper, we extend LDA to Structured Sparse LDA (SSLDA), where the projecting vectors are not only constrained to sparsity but also structured with a pre-specified set of shapes. While the sparse priors deal with small sample size problem, the proposed structure regularization can also encode higher-order information with better interpretability. We also propose a simple and efficient optimization algorithm to solve the proposed optimization problem. Experiments on face images show the benefits of the proposed structured sparse LDA on both classification accuracy and interpretability. Zhen Cui 0001, Shiguang Shan, Haihong Zhang, Shihong Lao, Xilin Chen 0001 |
ICIP | 1 |
| 2011 | Sparsely Encoded Local Descriptor for face recognitionabstractIn this paper, a novel Sparsely Encoded Local Descriptor (SELD) is proposed for face recognition. Compared with K-means or Random-projection tree based previous methods, sparsity constraint is introduced in our dictionary learning and sequent image encoding, which implies more stable and discriminative face representation. Sparse coding also leads to an image descriptor of summation of sparse coefficient vectors, which is quite different from existing code-words appearance frequency(/histogram)-based descriptors. Extensive experiments on both FERET and challenging LFW database show the effectiveness of the proposed SELD method. Especially on the LFW dataset, recognition accuracy comparable to the best known results is achieved. Zhen Cui 0001, Shiguang Shan, Xilin Chen 0001, Lei Zhang 0036 |
FG | 1 |