EDBT 2026 Demo / reviewers in the wild / expert
Haipeng Chen 0002
dblp:53/8073-2
· DBLP profile ↗
67ranked-venue papers
11as first author
54since 2021 · last 2026
0000-0002-9410-4120ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 8 first-author · 41 since 2021Artificial intelligence and machine learning · 23 · 7 first-author · 21 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Computer networks · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Causality-Aligned Semantic Recovery for Incomplete Cross-Modal RetrievalabstractIncomplete cross-modal retrieval (ICMR) requires models to recover missing modalities and robustly align heterogeneous ones for effective retrieval. Existing methods, however, fall short in both aspects. They often rely on limited semantic cues, such as single samples or coarse category prototypes, which compromises reconstruction quality. Moreover, these approaches are vulnerable to learning spurious cross-modal correlations, thereby impairing accurate alignment and hindering retrieval performance. To address these challenges, we propose Causality-Aligned Semantic Recovery (CASR), a novel method designed to both comprehensively restore missing modalities and mitigate spurious associations between vision and language. Our CASR involves two essential components: i) the Missing Modality Imagination (MMI) module, which combines category semantic priors with relevant contextual information to achieve high-quality semantic reconstruction; ii) the Explicit Causal Alignment (ECA) module, which explicitly learns environment-invariant attention, effectively eliminating the interference of spurious correlations and improving retrieval performance. Furthermore, we extend CASR to the challenging task of Partially Aligned Cross-Modal Retrieval, where we treat unlabeled unpaired data as a form of incomplete data. By leveraging MMI and ECA modules, we are able to learn robust representations in this setting. Extensive experiments on benchmark datasets under various missing rates demonstrate that CASR achieves superior robustness and retrieval performance. Haipeng Chen 0002, Yu Liu 0004, Xun Yang 0001, Yuheng Liang, Yingda Lyu |
AAAI | 1 |
| 2026 | Dual Coding Theory in Action: Language-Assisted Human Pose Estimation in VideosabstractVideo-based human pose estimation aims to localize keypoints across frames, enabling robust analysis of human motion in applications such as sports, surveillance, and healthcare. However, existing methods rely solely on visual cues, limiting their robustness in complex scenes involving occlusion, motion blur, or poor lighting. In contrast, dual coding theory from psychology suggests that human cognition is inherently multimodal: we learn by integrating visual perception with linguistic context to form structured, semantic understandings of the world. Visual input provides concrete spatiotemporal grounding, while language offers symbolic abstraction that enhances reasoning and generalization. Motivated by this cognitive principle, we present the first framework that explicitly incorporates language as an auxiliary modality to enhance video-based pose estimation. To address the lack of paired video-text datasets, we first employ a Multimodal Large Language Model (MLLM) to generate textual descriptions of human interactions from videos. We then propose a novel coarse-to-fine multimodal alignment pipeline: a cross-modal semantic interaction module establishes initial grounding between spatiotemporal visual features and textual embeddings, while an optimal transport-based feature matching mechanism enforces fine-grained, geometry-aware alignment. This cognitively inspired design enables more accurate and robust pose estimation, especially in visually challenging scenes like occlusion and motion blur. Extensive experiments on three benchmarks confirm that our method consistently outperforms state-of-the-art approaches. Sifan Wu 0001, Haipeng Chen 0002, Yingda Lyu, Shaojing Fan, Zhenguang Liu, Yingying Jiao |
AAAI | 2 |
| 2026 | Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose EstimationabstractVideo-based human pose estimation has vast applications such as action recognition, sports analytics, and crime detection. However, this task is challenging as it involves interpreting both spatial context and temporal dynamics to accurately localize human anatomical keypoints in video sequences. Current approaches, often based on attention mechanisms, perform well but struggle in challenging scenarios like rapid motion and pose occlusion. We attribute these failures to two fundamental limitations: spatial uniformity, where models indiscriminately assign attention to both joint-relevant features and background clutter, thereby introducing spatial noise; and temporal rigidity, an inability to adapt to large joint displacements, resulting in severe feature misalignment during rapid motion. To overcome these challenges, we introduce PSTPose, a novel progressive spatiotemporal refinement framework. Specifically, to address the spatial uniformity problem, we propose a Discriminative Feature Enhancement (DFE) module that emphasizes joint-relevant features and a Feature Cluster Grouping (FCG) module that forms compact, semantically meaningful regions. For the temporal rigidity problem, we introduce a Deformable Spatiotemporal Fusion (DSF) module that adaptively aligns features across consecutive frames via deformation-aware sampling. This design ensures robust keypoint localization, particularly in cluttered and dynamic scenes. Extensive experiments on three large-scale benchmarks, PoseTrack2017, PoseTrack2018, PoseTrack21, demonstrate that PSTPose establishes a new state-of-the-art. Sifan Wu 0001, Haipeng Chen 0002, Yingda Lyu, Shaojing Fan, Zhenguang Liu, Yingying Jiao |
AAAI | 2 |
| 2026 | VGD: Value-Guided Diffusion Toward High-Utility Medical Image SegmentationabstractProgress in medical image segmentation is fundamentally constrained by the scarcity of annotated data. While diffusion models offer a promising solution by generating high-fidelity image–mask pairs, their utility for downstream tasks remains underexplored. A key bottleneck lies in the misalignment between generation outputs and task-specific needs—samples are produced independently of their utility for downstream training. To this end, we propose Value-Guided Diffusion (VGD), a lightweight sampling framework that integrates downstream model feedback into the generative inference process. VGD estimates a value score for each sample based on its utility to downstream training, and leverages this signal to iteratively guide the denoising trajectory toward high-reward regions of the data manifold. Crucially, VGD can be seamlessly integrated into existing medical diffusion models without any additional training or architectural modifications. Extensive experiments across multiple diffusion backbones and segmentation benchmarks demonstrate that VGD significantly boosts downstream segmentation performance while maintaining visual fidelity. Our findings highlight a task-aware sampling principle with potential to underpin future synthetic segmentation pipelines. Haipeng Chen 0002, Chengxin Yang, Yingda Lyu |
AAAI | 2 |
| 2026 | Pedestrian-Centric Discriminative and Fine-grained Semantic Mining for Text-based Person RetrievalabstractText-based Person Retrieval (TPR) aims to retrieve specific pedestrian images from a gallery based on the given textual descriptions, serving as a fine-grained instance of cross-modal retrieval on the Web. Current mainstream approaches primarily leverage pre-trained models and attention mechanisms to enhance multi-modal representations. Despite notable progress, they still struggle with two major challenges: 1) Intra-instance semantic asymmetry, which mainly derives from the partial semantic relevance conveyed by each image-text pair; and 2) Inter-instance semantic ambiguity, which arises from the high similarity of image-text pairs with different identities. These issues result in suboptimal semantic alignment and degraded retrieval accuracy. To this end, we propose a novel Pedestrian-Centric Discriminative and Fine-grained Semantic Mining (DFSM) framework for TPR. Specifically, our DFSM method comprises two essential components: 1) Text-aware Visual Refinement (TVR), which mitigates visual redundancy by selecting semantically relevant patches under textual guidance, and refines them via adaptive clustering and merging; 2) Token-level Semantic Alignment (TSA), which formulates the matching relationship between image regions and text words as a conditional transport (CT) problem, effectively mining fine-grained semantic differences and enhancing instance discrimination. Extensive experiments on four benchmarks validate the advantages of DFSM in terms of retrieval accuracy and visual interpretability. Yuheng Liang, Haipeng Chen 0002, Yu Liu 0004, Yingda Lyu |
WWW | 2 |
| 2026 | Explicit token modeling and hierarchical feature reasoning for multimodal fake news detection
Yixin Jia, Haipeng Chen 0002, Zenan Shi, Xun Yang 0001 |
Inf. Process. Manag. | 2 |
| 2026 | Bayesian perturbation-driven consistency regularization for semi-supervised medical image segmentation
Haipeng Chen 0002, Yingda Lyu, Zenan Shi, Yongping Yang, Yu Wang 0112 |
Knowl. Based Syst. | 2 |
| 2026 | ADNet: Delving into generalizable deepfake detection via adaptive expert selection and discrepancy learning
Haipeng Chen 0002, Yixin Jia, Zenan Shi |
Pattern Recognit. | 1 |
| 2026 | Rethinking Skeleton-Based Action Recognition From Action-Class Prediction Distribution PerspectiveabstractAction recognition has long been a fundamental and compelling problem in the field of computer vision. However, one aspect that has been overlooked so far is that current action recognition approaches often produce an unfavourable multi-peaked distribution when identifying the action class of a given motion sequence, which is ambiguous and hard to learn for neural networks. Moreover, current methods heavily rely on neural networks to extract action features for differentiating actions, lacking theoretical constraints ensuring that action-specific features are selectively extracted and ambiguous features common to multiple actions are effectively reduced. These shortcomings culminate in inadequate action recognition accuracy. Motivated by this, in this paper we seek to tackle the problem from three aspects: 1) We try to eliminate ambiguity by enforcing a smooth single-peaked distribution instead of a multi-peaked one for action-class prediction. 2) We theoretically analyze the lower bound of the label prediction log-likelihood and derive a training objective, which focuses on the extraction of action-specific features and the reduction of ambiguous features. 3) We further advocate feeding the model with richer information, including positive information like body-part structures and negative information like masked inputs. Empirically, our approach sets the new state-of-the-art performance on five large-scale benchmarks. Our code is released at https://github.com/ActionR-Group/DPM to facilitate future research. Yingying Jiao, Haipeng Chen 0002, Yingda Lyu, Shuang Wu 0002, Zhenguang Liu |
IEEE Trans. Image Process. | 2 |
| 2026 | Dual-Supervised Asymmetric Co-Training for Semi-Supervised Medical Domain GeneralizationabstractSemi-supervised domain generalization (SSDG) in medical image segmentation offers a promising solution for generalizing to unseen domains during testing, addressing domain shift challenges and minimizing annotation costs. However, conventional SSDG methods assume labeled and unlabeled data are available for each source domain in the training set, a condition that is not always met in practice. The coexistence of limited annotation and domain shift in the training set is a prevalent issue. Thus, this paper explores a more practical and challenging scenario, cross-domain semi-supervised domain generalization (CD-SSDG), where domain shifts occur between labeled and unlabeled training data, in addition to shifts between training and testing sets. Existing SSDG methods exhibit sub-optimal performance under such domain shifts because of inaccurate pseudo-labels. To address this issue, we propose a novel dual-supervised asymmetric co-training (DAC) framework tailored for CD-SSDG. Building upon the co-training paradigm with two sub-models offering cross pseudo supervision, our DAC framework integrates extra feature-level supervision and asymmetric auxiliary tasks for each sub-model. This feature-level supervision serves to address inaccurate pseudo supervision caused by domain shifts between labeled and unlabeled data, utilizing complementary supervision from the rich feature space. Additionally, two distinct auxiliary self-supervised tasks are integrated into each sub-model to enhance domain-invariant discriminative feature learning and prevent model collapse. Extensive experiments on real-world medical image segmentation datasets,i.e., Fundus, Polyp, and SCGM, demonstrate the robust generalizability of the proposed DAC framework. Jincai Song, Haipeng Chen 0002, Na Zhao 0004 |
IEEE Trans. Multim. | 2 |
| 2026 | Discriminative Representation Learning for Remote Sensing Visual Question AnsweringabstractRecently, Remote Sensing Visual Question Answering (RSVQA) has attracted increasing attention from both academia and industry, which is the basis for understanding the underlying correspondence between remote sensing imagery and text descriptions. However, current methods are still insufficient in learning discriminative visual and textual representations for answer reasoning, mainly due to two reasons: (1) the remote sensing image environment is complex and changeable, and the target scales vary significantly, making it difficult to extract discriminative visual features; and (2) there is a lack of effective guidance from remote sensing domain knowledge to learn discriminative features. To this end, we propose a D iscriminative R epresentation L earning (DRL) method that includes two key strategies: visual feature enhancement and prior knowledge guidance. Specifically, we employ the Fourier transform to simulate the diverse visual environment and force the model to mine discriminative visual representations by imposing consistency constraints with the original features. In addition, we leverage the Remote Sensing Multimodal Large Language Model (RSMLLM) to generate captions rich in remote sensing domain-specific prior knowledge. These captions, derived from RSMLLM’s powerful knowledge integration and summarization capabilities, can then be compared and fused with visual representations to generate more discriminative representations, which are ultimately used for answer reasoning. Finally, recognizing that most existing RSVQA methods rely solely on static remote sensing images, we introduce RSVideoQA, a novel satellite video question answering dataset. This dataset is designed to facilitate the exploration of the rich spatio-temporal dynamics inherent in video sequences. Experimental results across three distinct datasets validate the effectiveness of our proposed method. Our dataset and code will be released at https://github.com/chill-han/DRL . Yingda Lyu, Yu Liu 0004, Haipeng Chen 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Causal-Inspired Multitask Learning for Video-Based Human Pose EstimationabstractVideo-based human pose estimation has long been a fundamental yet challenging problem in computer vision. Previous studies focus on spatio-temporal modeling through the enhancement of architecture design and optimization strategies. However, they overlook the causal relationships in the joints, leading to models that may be overly tailored and thus estimate poorly to challenging scenes. Therefore, adequate causal reasoning capability, coupled with good interpretability of model, are both indispensable and prerequisite for achieving reliable results. In this paper, we pioneer a causal perspective on pose estimation and introduce a causal-inspired multitask learning framework, consisting of two stages. In the first stage, we try to endow the model with causal spatio-temporal modeling ability by introducing two self-supervision auxiliary tasks. Specifically, these auxiliary tasks enable the network to infer challenging keypoints based on observed keypoint information, thereby imbuing causal reasoning capabilities into the model and making it robust to challenging scenes. In the second stage, we argue that not all feature tokens contribute equally to pose estimation. Prioritizing causal (keypoint-relevant) tokens is crucial to achieve reliable results, which could improve the interpretability of the model. To this end, we propose a Token Causal Importance Selection module to identify the causal tokens and non-causal tokens (e.g., background and objects). Additionally, non-causal tokens could provide potentially beneficial cues but may be redundant. We further introduce a non-causal tokens clustering module to merge the similar non-causal tokens. Extensive experiments show that our method outperforms state-of-the-art methods on three large-scale benchmark datasets. Haipeng Chen 0002, Sifan Wu 0001, Yifang Yin, Yingying Jiao, Yingda Lyu, Zhenguang Liu |
AAAI | 1 |
| 2025 | Skeleton-based Action Recognition with Non-linear Dependency Modeling and Hilbert-Schmidt Independence CriterionabstractHuman skeleton-based action recognition has long been an indispensable aspect of artificial intelligence. Current state-of-the-art methods tend to consider only the dependencies between connected skeletal joints, limiting their ability to capture non-linear dependencies between physically distant joints. Moreover, most existing approaches distinguish action classes by estimating the probability density of motion representations, yet the high-dimensional nature of human motions invokes inherent difficulties in accomplishing such measurements. In this paper, we seek to tackle these challenges from two directions: (1) We propose a novel dependency refinement approach that explicitly models dependencies between any pair of joints, effectively transcending the limitations imposed by joint distance. (2) We further propose a framework that utilizes the Hilbert-Schmidt Independence Criterion to differentiate action classes without being affected by data dimensionality, and mathematically derive learning objectives guaranteeing precise recognition. Empirically, our approach sets the state-of-the-art performance on NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA datasets. Haipeng Chen 0002, Yingda Lyu |
AAAI | 1 |
| 2025 | Auxiliary Tasks Benefit Skeleton-based Action RecognitionabstractSkeleton-based action recognition has long been a fundamental and intriguing problem in machine intelligence. This task is challenging due to pose occlusion and rapid motion, which typically results in incomplete or noisy skeleton data. State-of-the-art methods tend to learn human motion directly from these corrupted skeletons as if they were reliable. Unfortunately, this might lead to unsatisfactory results when key regions of the skeleton are occluded or disturbed. To tackle the problem, we propose a novel framework that integrates auxiliary tasks into a motion modeling network. These auxiliary tasks corrupt partial human skeletons with masking or noise and then force the network to recover the corrupted data, explicitly facilitating robust feature representation learning. We further propose supervising the auxiliary tasks with mutual information losses, mathematically ensuring feature consistency and spatial alignment between the recovered and original skeleton data. Empirically, our approach sets the new state-of-the-art performance on three benchmark datasets. Haipeng Chen 0002 |
ICASSP | 2 |
| 2025 | Enhancing Semantic Clarity: Discriminative and Fine-grained Information Mining for Remote Sensing Image-Text RetrievalabstractRemote sensing image-text retrieval is a fundamental task in remote sensing multimodal analysis, promoting the alignment of visual and language representations. The mainstream approaches commonly focus on capturing shared semantic representations between visual and textual modalities. However, the inherent characteristics of remote sensing image-text pairs lead to a semantic confusion problem, stemming from redundant visual representations and high inter-class similarity. To tackle this problem, we propose a novel Discriminative and Fine-grained Information Mining (DFIM) model, which aims to enhance semantic clarity by reducing visual redundancy and increasing the semantic gap between different classes. Specifically, the Dynamic Visual Enhancement (DVE) module adaptively enhances the visual discriminative features under the guidance of multimodal fusion information. Meanwhile, the Fine-grained Semantic Matching (FSM) module cleverly models the matching relationship between image regions and text words as an optimal transport problem, thereby refining intra-instance matching. Extensive experiments on two benchmark datasets justify the superiority of DFIM in terms of retrieval accuracy and visual interpretability over the leading methods. Yu Liu 0004, Haipeng Chen 0002, Yuheng Liang, Xun Yang 0001, Yingda Lyu |
IJCAI | 2 |
| 2025 | Dataset-level color augmentation and multi-scale exploration methods for polyp segmentation
Haipeng Chen 0002, Honghong Ju, Jincai Song, Yingda Lyu, Xianzhu Liu |
Expert Syst. Appl. | 1 |
| 2025 | Enhancing Human Pose Estimation in Internet of Things via Diffusion Generative ModelsabstractWith the ongoing development of public video surveillance technology, accurate human pose estimation is becoming increasingly important in urban administration and law enforcement. However, existing methods rely on large-scale dense annotations, which are labor-intensive and time-consuming. To tackle this, we propose SparsePose which leverages training videos with sparse annotations (labeled every k frames) to learn to propagate temporal poses that help to estimate the poses in unlabeled frames. Technically, we engage in a novel dual-branch architecture that combines 1) pose forecasting of the consecutive neighboring frames with 2) visual clues of the current frame and the nearest labeled frames. We theoretically derive the intrabranch and interbranch mutual information loss to supervise that maximized pose-relevant features are extracted from the current frame and different branches complement each other to approach precise pose estimation. Additionally, we propose a diffusion generative enhancement, which improves the robustness of the model to challenging scenes from the perspective of diversity. Empirical results show that our method significantly outperforms the state-of-the-art methods in sparsely labeled pose estimation on three benchmark datasets. Sifan Wu 0001, Hongzhe Zhang, Zhenguang Liu, Haipeng Chen 0002, Yingying Jiao |
IEEE Internet Things J. | 4 |
| 2025 | Multi-level semantics probability embedding for image-text matching
Anan Liu, Wenhui Li 0001, Weizhi Nie, Xianzhu Liu, Haipeng Chen 0002 |
Inf. Process. Manag. | 6 |
| 2025 | Rethinking Polyp Segmentation from the Perspectives of Matching Views and Seeking Camouflage
Zhengfang Jiang, Haipeng Chen 0002, Yongping Yang, Xianzhu Liu, Yingda Lyu |
Multim. Syst. | 2 |
| 2025 | Multi-modality boundary-guided network for generalizable image manipulation localization
Yanyan Jiang 0005, Haipeng Chen 0002, Yingda Lyu |
Multim. Syst. | 3 |
| 2025 | Dual Space Representation Learning for Skeleton-Based Action RecognitionabstractSkeleton-based action recognition is crucial for machine intelligence. Current methods generally learn from 3D articulated motion sequences in the straightforward Euclidean space. Yet, thevanillaEuclidean space may not be the optimal choice for modeling the intricate correlations among human body joints. This challenge arises from the non-Euclidean nature of human anatomy, where joint correlations often vary non-linearly during movement. To address this, we propose a dual space representation learning method. Specifically, we represent the motion sequences in Hyperbolic space, leveraging its intrinsic properties to capture the non-Euclidean latent anatomy of human motions. We then incorporate the motion features from both Hyperbolic and Euclidean spaces, allowing us to precisely model the non-linear joint correlations while effectively sketching human poses. The proposed method empirically achieves state-of-the-art performance on the NTU RGB+D 60, NTURGB+D 120, and NW-UCLA datasets. Haipeng Chen 0002, Zhenguang Liu, Sihao Hu, Yingying Jiao |
IEEE Signal Process. Lett. | 2 |
| 2025 | Causality-Inspired Unsupervised Domain Adaptation With Target Style Imitation for Medical Image SegmentationabstractDeep learning performance may decrease substantially with unseen heterogeneous data. While most unsupervised domain adaptation (UDA) methods seek to address this through image alignment, they often ignore uncertainty style fluctuations within the target domain. When testing image styles vary in both direction and intensity, such models may fail to adapt. Furthermore, existing UDA methods tend to over-reliance on domain-level entire feature alignment, resulting in potentially over-exploiting semantic content-independent cues (e.g., intensity) as shortcut features. To address these limitations, this paper introduces an innovative and model-agnostic Causality-inspired Representation Learning Based on Target Style Imitation method for UDA. Specifically, we propose a novel Target Style Imitation (TSI) data augmentation approach to diversify the training data and align training and unseen target testing image styles. TSI constructs a Gaussian distribution for the target domain style and simulates unseen testing style variations through random sampling. Additionally, inspired by the stable and generalizable causal mechanism, we propose Causality-inspired Representation Learning (CRL) based on TSI method to enforce feature representations to adhere to causal properties (i.e., Separation and Independence) essential for robust UDA, thus fostering the model to focus on the domain-invariant semantic features. Our method surpasses state-of-the-art methods on two cross-modality medical image segmentation datasets. Jincai Song, Haipeng Chen 0002, Yingda Lyu, Weizhi Nie, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Customized Transformer Adapter With Frequency Masking for Deepfake DetectionabstractThe evolution of advanced artificial intelligence generated content approaches has heightened concerns about deepfake, due to the sophisticated forgeries and concealed appearances they produce. To this end, the pre-trained Vision Transformer (ViT) model has become a de facto choice for deepfake detection, thanks to its powerful learning capability. Despite favorable results achieved by existing ViT-based methods, they have inherent limitations that could result in suboptimal performance in scenarios with continuously evolving forgery techniques, such as overfitting to single forgery patterns or placing excessive emphasis on dominant forgery regions. In this paper, we propose CUTA, a simple yet effective deepfake detection paradigm that utilizes ViT adapters as the medium and fully exploits the spatial- and frequency-domain features of given images to overcome the limitations of existing methods. Specifically, CUTA focuses onfrequency domain maskingwithin the input space, which obscures parts of the high-frequency image to intensify the training challenge while preserving subtle forgery cues in the frequency domain to facilitate comprehensive forgery representations. Furthermore, we propose two task-customized modules within the ViT model, i.e., thetexture enhancement moduleand themulti-scale perceptron module, to seamlessly integrate local texture and rich contextual features. These two modules ensure an organic interaction between the task-specific forgery patterns and general semantic features within the pre-trained ViT framework. The experimental results on several publicly available benchmark datasets demonstrate CUTA’s superiority in performance, particularly showcasing its significant advantages in both cross-dataset and cross-manipulation scenarios. Zenan Shi, Haipeng Chen 0002, Yixin Jia, Wei Lu 0001, Xun Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Bias Mitigation and Representation Optimization for Noise-Robust Cross-Modal RetrievalabstractThe remarkable progress in cross-modal retrieval relies on accurately annotated multimedia datasets. In practice, most existing datasets used for training cross-modal retrieval models are automatically collected from the Internet to reduce data collection costs. However, it inevitably contains mismatched pairs, i.e., noisy correspondences, thus degrading the model performance. Recent advances utilize the predicted similarity distribution of individual samples for noise validation and correction, which easily faces two challenging dilemmas: (1) confirmation bias and (2) unstable performance with increasing noise. In light of the above, we propose a generalized Bias Mitigation and Representation Optimization (BMRO) framework. Specifically, we propose a Bias Estimator (BE) to estimate the unbiased confidence factor of a sample by contrasting it against its nearest neighbors. Unbiased confidence factor can precisely adjust sample contribution and enhance accurate sample division. This facilitates the Adaptive Representation Optimizer (ARO) in providing tailored optimization strategies for clean and noisy samples. ARO performs contrastive learning between clean samples and generated hard samples, thus promoting the generalizability and robustness of the representation. Besides, it utilizes complementary learning to reduce incorrect guidance from noisy samples. Extensive experiments on five visual-text benchmarks verify that our BMRO can significantly improve the matching accuracy and performance stability against noisy correspondences. Yu Liu 0004, Haipeng Chen 0002, Guihe Qin, Jincai Song, Xun Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Face Reconstruction-Based Generalized Deepfake Detection Model with Residual Outlook AttentionabstractWith the continuous development of deep counterfeiting technology, the information security in our daily life is under serious threat. While existing face forgery detection methods exhibit impressive accuracy when applied to datasets such as FaceForensics++ and Celeb-DF, they falter significantly when confronted with out-of-domain scenarios. This causes specialization of learned representations to known forgery patterns presented in the training set, rendering it difficult to detect forgeries with unknown patterns. To address this challenge, we propose a novel end-to-end Face Reconstruction-Based Generalized Deepfake Detection (FRG2D) model with Residual Outlook Attention (ROA) , which emphasizes the robust visual representations of genuine faces and discerns the subtle differences between authentic and manipulated facial images. Our methodology entails reconstructing authentic face images using an encoder–decoder architecture based on U-net, facilitating a deeper understanding of disparities between genuine and manipulated facial images. Furthermore, we integrate the convolutional block attention module (CBAM) and channel attention block (CAB) to selectively focus the network’s attention on salient features within real face images. Furthermore, we employ ROA to guide the network’s focus towards precise features within manipulated facial images. Simultaneously, the computed reconstruction differences obtained through ROA serves as the ultimate representation fed into the classifier for face forgery detection. Both the reconstruction and classification learning processes are optimized end-to-end. Through extensive experimentation, our model demonstrated a substantial improvement in deepfake detection across unknown domains, while maintaining a high accuracy within the known domain. Zenan Shi, Wenyu Liu 0013, Haipeng Chen 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Causality-Inspired Invariant Representation Learning for Text-Based Person RetrievalabstractText-based Person Retrieval (TPR) aims to retrieve relevant images of specific pedestrians based on the given textual query. The mainstream approaches primarily leverage pretrained deep neural networks to learn the mapping of visual and textual modalities into a common latent space for cross-modality matching. Despite their remarkable achievements, existing efforts mainly focus on learning the statistical cross-modality correlation found in training data, other than the intrinsic causal correlation. As a result, they often struggle to retrieve accurately in the face of environmental changes such as illumination, pose, and occlusion, or when encountering images with similar attributes. In this regard, we pioneer the observation of TPR from a causal view. Specifically, we assume that each image is composed of a mixture of causal factors (which are semantically consistent with text descriptions) and non-causal factors (retrieval-irrelevant, e.g., background), and only the former can lead to reliable retrieval judgments. Our goal is to extract text-critical robust visual representation (i.e., causal factors) and establish domain invariant cross-modality correlations for accurate and reliable retrieval. However, causal/non-causal factors are unobserved, so we emphasize that ideal causal factors that can simulate causal scenes should satisfy two basic principles:1) Independence: being independent of non-causal factors, and 2)Sufficiency: being causally sufficient for TPR across different environments. Building on that, we propose an Invariant Representation Learning method for TPR (IRLT), that enforces the visual representations to satisfy the two aforementioned critical properties. Extensive experiments on three datasets clearly demonstrate the advantages of IRLT over leading baselines in terms of accuracy and generalization. Yu Liu 0004, Guihe Qin, Haipeng Chen 0002, Zhiyong Cheng 0001, Xun Yang 0001 |
AAAI | 3 |
| 2024 | Rethinking Human Motion Prediction with Symplectic IntegralabstractLong-term and accurate forecasting is the long-standing pursuit of the human motion prediction task. Existing methods typically suffer from dramatic degradation in prediction accuracy with increasing prediction horizon. It comes down to two reasons: 1) Insufficient numerical stability caused by unforeseen high noise and complex feature relationships in the data, and 2) Inadequate modeling stability caused by unreasonable step sizes and undesirable parameter updates in the prediction. In this paper, we design a novel and sym-plectic integral-inspired framework named symplectic integral neural network (SINN), which engages symplectic tra-jectories to optimize the pose representation and employs a stable symplectic operator to alternately model the dynamic context. Specifically, we design a Symplectic Repre-sentation Encoder that performs on enhanced human pose representation to obtain trajectories on the symplectic manifold, ensuring numerical stability based on Hamiltonian mechanics and symplectic spatial splitting algorithm. We further present the Symplectic Temporal Aggregation mod-ule, which splits the long-term prediction into multiple ac-curate short-term predictions generated by a symplectic operator to secure modeling stability. Moreover, our approach is model-agnostic and can be efficiently integrated with different physical dynamics models. The experimental results demonstrate that our method achieves the new state-of-the-art, outperforming existing methods by 20.1% on Human3.6M, 16.7% on CUM Mocap, and 10.2% on 3DPW. Haipeng Chen 0002, Kedi Lyu, Zhenguang Liu, Yifang Yin, Xun Yang 0001, Yingda Lyu |
CVPR | 1 |
| 2024 | Joint-Motion Mutual Learning for Pose Estimation in VideoabstractHuman pose estimation in videos has long been a compelling yet challenging task within the realm of computer vision. Nevertheless, this task remains difficult because of the complex video scenes, such as video defocus and self-occlusion. Recent methods strive to integrate multi-frame visual features generated by a backbone network for pose estimation. However, they often ignore the useful joint information encoded in the initial heatmap, which is a by-product of the backbone generation. Comparatively, methods that attempt to refine the initial heatmap fail to consider any spatio-temporal motion features. As a result, the performance of existing methods for pose estimation falls short due to the lack of ability to leverage both local joint (heatmap) information and global motion (feature) dynamics. Sifan Wu 0001, Haipeng Chen 0002, Yifang Yin, Sihao Hu, Runyang Feng, Yingying Jiao, Zhenguang Liu |
ACM Multimedia | 2 |
| 2024 | CRML-Net: Cross-Modal Reasoning and Multi-Task Learning Network for tooth image segmentation
Yingda Lyu, Zhehao Liu, Haipeng Chen 0002 |
Comput. Vis. Image Underst. | 4 |
| 2024 | Reinforced visual interaction fusion radiology report generation
Haipeng Chen 0002, Yu Liu 0004, Yingda Lyu |
Multim. Syst. | 2 |
| 2024 | View sequence prediction GAN: unsupervised representation learning for 3D shapes by decomposing view content and viewpoint variance
Heyu Zhou, Jiayu Li 0004, Xianzhu Liu, Yingda Lyu, Haipeng Chen 0002, Anan Liu |
Multim. Syst. | 5 |
| 2024 | Collaborative region-boundary interaction network for medical image segmentation
Na Ta 0009, Haipeng Chen 0002, Zenan Shi |
Multim. Tools Appl. | 2 |
| 2024 | Regular Constrained Multimodal Fusion for Image CaptioningabstractMore diverse and closer to human-like captions are of paramount importance in image captioning. Recent research has achieved significant advancements, with the majority adopting end-to-end encoder-decoder architectures that integrate specific feature-text processing. However, the homogeneity of their model structures, the simplicity or complexity of feature-text fusion, and the uniformity of training objectives have all to some extent affected the diversity and effectiveness of caption generation, thus limiting the potential applications of this task. Therefore, in this paper, we propose the Regular Constrained Multimodal Fusion (RCMF) method for image captioning to better integrate information across and within modalities, while also approaching human-like fine-grained semantic perception and relationship reasoning capabilities. Initially, our RCMF preprocesses images using a Swin-Transformer and then an extended encoder with a new intra-modal fusion module, utilizing window-focused linear attention to capture features and leveraging refined grid and global visual features. By combining text features, RCMF employs a cross-modal fusion module and decoder to deeply model the interaction between text and image. Additionally, RCMF first introduces a new additional regulatory modal fusion reasoning (MFR) branch, which surpasses the above architectures. Its MFR loss combined with cross-entropy loss forms a new training objective strategy, effectively mining fine-grained relationships between images and text, perceiving the semantic information of images and their corresponding captions, thereby regulating the generated captions to be more diverse and human-like. Experimental results based on the MS COCO 2014 dataset, particularly under the same experimental conditions, demonstrate the outstanding performance of our method, especially in terms of METEOR, ROUGE-L, CIDEr, and SPICE metrics. Visualization results further intuitively confirm the effectiveness of our RCMF method. Source code inhttps://github.com/200084/RCMF-for-image-caption. Haipeng Chen 0002, Yu Liu 0004, Yingda Lyu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Discrepancy-Guided Reconstruction Learning for Image Forgery DetectionabstractIn this paper, we propose a novel image forgery detection paradigm for boosting the model learning capacity on both forgery-sensitive and genuine compact visual patterns. Compared to the existing methods that only focus on the discrepant-specific patterns (\eg, noises, textures, and frequencies), our method has a greater generalization. Specifically, we first propose a Discrepancy-Guided Encoder (DisGE) to extract forgery-sensitive visual patterns. DisGE consists of two branches, where the mainstream backbone branch is used to extract general semantic features, and the accessorial discrepant external attention branch is used to extract explicit forgery cues. Besides, a Double-Head Reconstruction (DouHR) module is proposed to enhance genuine compact visual patterns in different granular spaces. Under DouHR, we further introduce a Discrepancy-Aggregation Detector (DisAD) to aggregate these genuine compact visual patterns, such that the forgery detection capability on unknown patterns can be improved. Extensive experimental results on four challenging datasets validate the effectiveness of our proposed method against state-of-the-art competitors. Zenan Shi, Haipeng Chen 0002, Long Chen 0016 |
IJCAI | 2 |
| 2023 | Action Recognition with Multi-stream Motion Modeling and Mutual Information MaximizationabstractAction recognition has long been a fundamental and intriguing problem in artificial intelligence. The task is challenging due to the high dimensionality nature of an action, as well as the subtle motion details to be considered. Current state-of-the-art approaches typically learn from articulated motion sequences in the straightforward 3D Euclidean space. However, the vanilla Euclidean space is not efficient for modeling important motion characteristics such as the joint-wise angular acceleration, which reveals the driving force behind the motion. Moreover, current methods typically attend to each channel equally and lack theoretical constrains on extracting task-relevant features from the input. In this paper, we seek to tackle these challenges from three aspects: (1) We propose to incorporate an acceleration representation, explicitly modeling the higher-order variations in motion. (2) We introduce a novel Stream-GCN network equipped with multi-stream components and channel attention, where different representations (i.e., streams) supplement each other towards a more precise action recognition while attention capitalizes on those important channels. (3) We explore feature-level supervision for maximizing the extraction of task-relevant information and formulate this into a mutual information loss. Empirically, our approach sets the new state-of-the-art performance on three benchmark datasets, NTU RGB+D, NTU RGB+D 120, and NW-UCLA. Haipeng Chen 0002, Zhenguang Liu, Yingda Lyu, Beibei Zhang 0007, Shuang Wu 0002, Zhibo Wang 0001, Kui Ren 0001 |
IJCAI | 2 |
| 2023 | A complementary and contrastive network for stimulus segmentation and generalization
Na Ta 0009, Haipeng Chen 0002, Yingda Lyu, Zenan Shi, Zhehao Liu |
Image Vis. Comput. | 2 |
| 2023 | FPF-Net: feature propagation and fusion based on attention mechanism for pancreas segmentation
Haipeng Chen 0002, Zenan Shi |
Multim. Syst. | 1 |
| 2023 | LET-Net: locally enhanced transformer network for medical image segmentationabstractAbstract Medical image segmentation has attracted increasing attention due to its practical clinical requirements. However, the prevalence of small targets still poses great challenges for accurate segmentation. In this paper, we propose a novel locally enhanced transformer network (LET-Net) that combines the strengths of transformer and convolution to address this issue. LET-Net utilizes a pyramid vision transformer as its encoder and is further equipped with two novel modules to learn more powerful feature representation. Specifically, we design a feature-aligned local enhancement module, which encourages discriminative local feature learning on the condition of adjacent-level feature alignment. Moreover, to effectively recover high-resolution spatial information, we apply a newly designed progressive local-induced decoder. This decoder contains three cascaded local reconstruction and refinement modules that dynamically guide the upsampling of high-level features by their adaptive reconstruction kernels and further enhance feature representation through a split-attention mechanism. Additionally, to address the severe pixel imbalance for small targets, we design a mutual information loss that maximizes task-relevant information while eliminating task-irrelevant noises. Experimental results demonstrate that our LET-Net provides more effective support for small target segmentation and achieves state-of-the-art performance in polyp and breast lesion segmentation tasks. Na Ta 0009, Haipeng Chen 0002, Xianzhu Liu, Nuo Jin |
Multim. Syst. | 2 |
| 2023 | BLE-Net: boundary learning and enhancement network for polyp segmentation
Na Ta 0009, Haipeng Chen 0002, Yingda Lyu, Taosuo Wu |
Multim. Syst. | 2 |
| 2023 | M-AResNet: a novel multi-scale attention residual network for melting curve image classification
Pengxiang Su, Xuanjing Shen, Haipeng Chen 0002, Di Gai, Yu Liu 0004 |
Multim. Tools Appl. | 3 |
| 2023 | PL-GNet: Pixel Level Global Network for detection and localization of image forgeries
Zenan Shi, Xuanjing Shen, Haipeng Chen 0002, Yingda Lyu |
Signal Process. Image Commun. | 3 |
| 2023 | Spatiotemporal Consistency Learning From Momentum Cues for Human Motion PredictionabstractExtrapolating future human motion based on the historical human pose sequence is the foundation of various intelligent applications. Numerous deep learning-based algorithms have been designed to address this task, achieving state-of-the-art performance on different human motion benchmark datasets. However, most existing methods employ three-dimensional coordinates of joints to demonstrate dynamic motion contexts implicitly. Unfortunately, it remains challenging in capturing motion information from the pose sequence. In this paper, we advocate explicitly describing dynamic contexts via the momentum of human motion mechanic space, as the momentum of a joint is explicit, temporal consistent, and can provide abundant information to the model. In addition, the single-stream methods play a dominant role in the field of human motion prediction. They usually capture motion information via the strategy of continuous or sparse sampling, which might obviate global or detailed local information. Therefore, we present a simple yet effective dual-stream method that can consider both the detailed and global temporal information through a combination of continuous and sparse sampling. The proposed dual-stream paradigm enables the improvement of computational efficiency and the short-term prediction accuracy concurrently. Furthermore, we present a novel temporal attention-based graph convolutional network (TA-GCN) to derive a spatiotemporally consistent motion representation, which can adequately consider the rationality of human body topology. Extensive experiments on two large motion prediction benchmark datasets (i.e., Human 3.6M and CMU Mocap) show that our algorithm achieves state-of-the-art performance both qualitatively and quantitatively. Haipeng Chen 0002, Wenyin Zhang, Pengxiang Su |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Transformer-Auxiliary Neural Networks for Image Manipulation Localization by Operator InductionsabstractImage manipulation localization (IML), which seeks to accurately segment tampered regions that are artfully fastened into a normal image, is a fundamental yet challenging computer vision task. Despite that impressive results have been achieved by some progressive deep learning methods, they usually fail in capturing the subtle manipulation artifacts at different object scales, which are not competent to generate a perfect segmentation mask with complete and fine object structures. Besides, the problem of coarse boundaries also occurs frequently. To this end, in this paper, we propose a Transformer-Auxiliary by operator-induced neural Network (TANet) to localize forged regions for IML. Specifically, a stacked multi-scale transformer (SMT) branch is first introduced as a compensation for feature representations of the mainstream convolutional neural network branch. SMT can detect structured abnormalities of the input image at multi-levels by operating on patches of different sizes. Then TANet explicitly exploits an operator induction module (OIM) to excavate valuable and manipulated region-related boundary semantics to guide the representative learning of the mainstream branch. The OIM encourages the network to generate features that highlight object structure, thereby promoting precise boundary localization of forged regions. We conduct extensive experiments on various datasets and settings to validate the effectiveness of TANet. Results show that TANet outperforms the state-of-the-art methods by a large margin under widely-used evaluation metrics. Zenan Shi, Haipeng Chen 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Multiscale Spatial and Temporal Learning for Human Motion Prediction
Pengxiang Su, Xuanjing Shen, Haipeng Chen 0002 |
ICANN (2) | 3 |
| 2022 | Spatial-Temporal Correlation Modeling for Motion PredictionabstractHuman motion prediction is fundamental for many applications in computer vision. Current methods typically handle motion prediction with seqential models, which ignore the fact that joint movement is driven by forces. In this paper, we provide a novel mechanical view to decompose force into magnitude and direction, which contributes to modeling the temporal evolution of joints. Moreover, existing graph convolution-based methods merely utilize the deep-level features, which is difficult to capture the complex spatial dependencies contexts. We introduce a novel spatial connections encoding model to capture the multi-level spatial dependencies between joints. Finally, to encode abundant temporal dependencies, we present a multi-head temporal encoding module. Comprehensive experiments show that our model sets the state-of-the-art performance on the largest human motion benchmark datasets. Yingying Jiao, Haipeng Chen 0002, Chang Yao 0001, Pengxiang Su, Chong Fu 0002, Xiang Wang 0010 |
ICME | 2 |
| 2022 | PR-NET: Progressively-refined neural network for image manipulation localizationabstractCurrent deep learning-based image manipulation localization methods achieve impressive performance when rich spatial features and information are fully utilized. However, most of them suffer from the irrelevance of semantic awareness when identifying various manipulation categories. This leads to false alarms on recognizing forged regions. In this paper, we propose a Progressively-Refined Neural Network (PR-Net), to localize the tampered regions progressively under a coarse-to-fine workflow. Specifically, PR-Net is composed of a Feature Extractor (FE) that captures feature intrinsic correlations and a Mask Generation Module (MGM) with three refining generators. The FE takes a CNN to extract the image features and introduces an attention mechanism Convolution Block Attention Module (CBAM) to suppress the image content and guide the extractor in exploring the inconsistencies between the manipulated and authentic regions. The MGM comprises three generators where the Coarse Mask RR-Generator generates a localization result roughly, the Candidate Mask RR-Generator generates a possible tampered region according to the rough localization measure, and the Fine Mask RR-Generator produces the final prediction of manipulated regions. We also utilize the Rotated Residual (RR) structure to suppress the image content during the generative process. The extensive experimental results on four benchmark data sets (NIST16, COVER, CASIA v1.0, and In-The-Wild) demonstrate the superior performance of PR-Net compared with the state-of-the-art methods in localizing the manipulated regions. Zenan Shi, Chaoqun Chang, Haipeng Chen 0002, Xiaoyu Du 0002, Hanwang Zhang |
Int. J. Intell. Syst. | 3 |
| 2022 | 3D human motion prediction: A survey
Kedi Lyu, Haipeng Chen 0002, Zhenguang Liu, Beiqi Zhang, Ruili Wang 0001 |
Neurocomputing | 2 |
| 2022 | Hybrid features and semantic reinforcement network for image forgery detection
Haipeng Chen 0002, Chaoqun Chang, Zenan Shi, Yingda Lyu |
Multim. Syst. | 1 |
| 2022 | Semantically guided projection for zero-shot 3D model classification and retrieval
Yuting Su 0001, Jiayu Li 0004, Wenhui Li 0001, Zan Gao 0002, Haipeng Chen 0002, Xuanya Li, Anan Liu |
Multim. Syst. | 5 |
| 2022 | Joint Local Correlation and Global Contextual Information for Unsupervised 3D Model Retrieval and ClassificationabstractUnsupervised 3D model analysis has attracted tremendous attentions with the increasing growth of 3D model data and the extensive human annotations. Many effective methods have been designed to address the 3D model analysis with labeled information, while rare methods devote to unsupervised deep learning due to the difficulty of mining reliable information. In this paper, we propose a novel unsupervised deep learning method named joint local correlation and global contextual information (LCGC) for 3D model retrieval and classification, which mines the reliable triplet set and uses triplet loss to optimize the deep neural network. Our method proposes two schemes: 1) Local self-correlation information learning, which adopts the intra and inter information to construct the view-level triplet set. 2) Global neighbor contextual information learning, which employs the neighbor contextual information to explore the reliable relations among 3D models and construct the model-level triplet set. The above schemes encourage that the selected triple set can been used to improve the discrimination of learned features. Extensive evaluations on two large-scale datasets, ModelNet40 and ShapeNet55, have demonstrated the effectiveness of our proposed method. Wenhui Li 0001, Zhenlan Zhao, Anan Liu, Zan Gao 0002, Chenggang Yan 0001, Zhendong Mao 0001, Haipeng Chen 0002, Weizhi Nie |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | GLPose: Global-Local Representation Learning for Human Pose EstimationabstractMulti-frame human pose estimation is at the core of many computer vision tasks. Although state-of-the-art approaches have demonstrated remarkable results for human pose estimation on static images, their performances inevitably come short when being applied to videos. A central issue lies in the visual degeneration of video frames induced by rapid motion and pose occlusion in dynamic environments. This problem, by nature, is insurmountable for a single frame. Therefore, incorporating complementary visual cues from other video frames becomes an intuitive paradigm. Current state-of-the-art methods usually leverage information from adjacent frames, which unfortunately place excessive focus on only the temporally nearby frames. In this paper, we argue that combining global semantically similar information and local temporal visual context will deliver more comprehensive and more robust representations for human pose estimation. Towards this end, we present an effective framework, namely global-local enhanced pose estimation ( GLPose ) network. Our framework consists of a feature processing module that conditionally incorporates global semantic information and local visual context to generate a robust human representation and a feature enhancement module that excavates complementary information from this aggregated representation to enhance keyframe features for precise estimation. We empirically find that the proposed GLpose outperforms existing methods by a large margin and achieves new state-of-the-art results on large benchmark datasets. Yingying Jiao, Haipeng Chen 0002, Runyang Feng, Haoming Chen, Sifan Wu 0001, Yifang Yin, Zhenguang Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | Aggregated Multi-GANs for Controlled 3D Human Motion PredictionabstractHuman motion prediction from historical pose sequence is at the core of many applications in machine intelligence. However, in current state-of-the-art methods, the predicted future motion is confined within the same activity. One can neither generate predictions that differ from the current activity, nor manipulate the body parts to explore various future possibilities. Undoubtedly, this greatly limits the usefulness and applicability of motion prediction. In this paper, we propose a generalization of the human motion prediction task in which control parameters can be readily incorporated to adjust the forecasted motion. Our method is compelling in that it enables manipulable motion prediction across activity types and allows customization of the human movement in a variety of fine-grained ways. To this aim, a simple yet effective composite GAN structure, consisting of local GANs for different body parts and aggregated via a global GAN is presented. The local GANs game in lower dimensions, while the global GAN adjusts in high dimensional space to avoid mode collapse. Extensive experiments show that our method outperforms state-of-the-art. The codes are available at https://github.com/herolvkd/AM-GAN. Zhenguang Liu, Kedi Lyu, Shuang Wu 0002, Haipeng Chen 0002, Yanbin Hao, Shouling Ji |
AAAI | 4 |
| 2021 | Motion Prediction using Trajectory CuesabstractPredicting human motion from a historical pose sequence is at the core of many applications in computer vision. Current state-of-the-art methods concentrate on learning motion contexts in the pose space, however, the high dimensionality and complex nature of human pose invoke inherent difficulties in extracting such contexts. In this paper, we instead advocate to model motion contexts in the joint trajectory space, as the trajectory of a joint is smooth, vectorial, and gives sufficient information to the model. Moreover, most existing methods consider only the dependencies between skeletal connected joints, disregarding prior knowledge and the hidden connections between geometrically separated joints. Motivated by this, we present a semi-constrained graph to explicitly encode skeletal connections and prior knowledge, while adaptively learn implicit dependencies between joints.We also explore the applications of our approach to a range of objects including human, fish, and mouse. Surprisingly, our method sets the new state-of-the-art performance on 4 different benchmark datasets, a remarkable highlight is that it achieves a 19.1% accuracy improvement over current state-of-the-art in average. To facilitate future research, we have released our code at https://github.com/Pose-Group/MPT. Zhenguang Liu, Pengxiang Su, Shuang Wu 0002, Xuanjing Shen, Haipeng Chen 0002, Yanbin Hao, Meng Wang 0001 |
ICCV | 5 |
| 2021 | Learning Human Motion Prediction via Stochastic Differential EquationsabstractHuman motion understanding and prediction is an integral aspect in our pursuit of machine intelligence and human-machine interaction systems. Current methods typically pursue a kinematics modeling approach, relying heavily upon prior anatomical knowledge and constraints. However, such an approach is hard to generalize to different skeletal model representations, and also tends to be inadequate in accounting for the dynamic range and complexity of motion, thus hindering predictive accuracy. In this work, we propose a novel approach in modeling the motion prediction problem based on stochastic differential equations and path integrals. The motion profile of each skeletal joint is formulated as a basic stochastic variable and modeled with the Langevin equation. We develop a strategy of employing GANs to simulate path integrals that amounts to optimizing over possible future paths. We conduct experiments in two large benchmark datasets, Human 3.6M and CMU MoCap. It is highlighted that our approach achieves a 12.48% accuracy improvement over current state-of-the-art methods in average. Kedi Lyu, Zhenguang Liu, Shuang Wu 0002, Haipeng Chen 0002, Xuhong Zhang 0002, Yuyu Yin |
ACM Multimedia | 4 |
| 2020 | Multi-focus noisy image fusion based on gradient regularized convolutional sparse representationeabstractThe method proposes a multi-focus noisy image fusion algorithm combining gradient regularized convolutional sparse representatione and spatial frequency. Firstly, the source image is decomposed into a base layer and a detail layer through two-scale image decomposition. The detail layer uses the Alternating Direction Method of Multipliers (ADMM) to solve the convolutional sparse coefficients with gradient penalties to complete the fusion of detail layer coefficients. Then, The base layer uses the spatial frequency to judge the focus area, the spatial frequency and the "choose-max" strategy are applied to achieved the multi-focus fusion result of base layer. Finally, the fused image is calculated as a superposition of the base layer and the detail layer. Experimental results show that compared with other algorithms, this algorithm provides excellent subjective visual perception and objective evaluation metrics. Xuanjing Shen, Haipeng Chen 0002, Di Gai |
MMAsia | 3 |
| 2020 | Medical image fusion using the PCNN based on IQPSO in NSST domainabstractIn this study, an improved quantum‐behaved particle swarm optimisation based pulse‐coupled neural network (IQPSO‐PCNN) is proposed in the non‐subsampled shearlet transform (NSST) domain for medical image fusion. First, NSST tool is used to decompose the source image into low‐frequency and high‐frequency subbands. Then, for low‐frequency subbands, the fusion rules of two different functions are presented, which simultaneously addresses two key issues of energy preservation and detail extraction. For high‐frequency subbands, unlike conventional PCNN‐based methods, parameters are manually set based on experience, and the decomposed high‐frequency subbands share a set of parameters. The IQPSO‐PCNN model can obtain the optimal parameters for each high‐frequency subband adaptively according to its own information. Finally, the fused low‐frequency subband and high‐frequency subbands are inversely transformed by NSST to acquire the final fused image. The proposed algorithm uses >90 pairs of images with four different modalities. In addition, fusion experiments are performed on different sequences of the three modes. The experimental results demonstrate that the proposed method is superior to existing state‐of‐art methods in subjective visual performance and objective evaluation. Di Gai, Xuanjing Shen, Haipeng Chen 0002, Zeyu Xie, Pengxiang Su |
IET Image Process. | 3 |
| 2020 | Multi-focus image fusion method based on two stage of convolutional neural network
Di Gai, Xuanjing Shen, Haipeng Chen 0002, Pengxiang Su |
Signal Process. | 3 |
| 2020 | Global Semantic Consistency Network for Image Manipulation DetectionabstractThis letter focuses on image manipulation detection which aims to recognize the manipulated regions under the contextual semantic information. Existing approaches usually overlook the semantic discrepancy between different levels of feature maps, and directly fuse (e.g., addition, or concatenation) them for detection. In this letter, we argue that the semantic gap is the main reason for the low effectiveness of feature fusion in manipulation predictions. To address this problem, we propose a Global Semantic Consistency Network (GSCNet) for image manipulation detection, which is based on an encoder-decoder structure. Specifically, to make GSCNet include more global texture information which has been empirically confirmed to be beneficial to manipulation detection, gram block is first deployed on each level of feature maps in the encoding stage. Based on that, bi-directional convolutional LSTM is further implemented on the decoding stage, such that feature maps of the same level have semantic consistency. Experimental results on NIST16, and CASIA v1.0 declare that GSCNet can accurately locate the manipulated regions. Furthermore, compared to the existing models, GSCNet can achieve new state-of-the-art results. Zenan Shi, Xuanjing Shen, Haipeng Chen 0002, Yingda Lyu |
IEEE Signal Process. Lett. | 3 |
| 2019 | A fusion algorithm for medical structural and functional images based on adaptive image decomposition
Xuanjing Shen, Haipeng Chen 0002, Yingda Lv, Xiaoli Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Splicing image forgery detection using textural features based on the grey level co-occurrence matricesabstractTo further improve the detection rate with relatively low dimension feature vector, a novel passive splicing detection method using textural features based on the grey level co‐occurrence matrices, namely TF‐GLCM, is proposed in this study. In the TF‐GLCM, the GLCM are calculated based on the difference block discrete cosine transform arrays to capture the textural information and the spatial relationship between image pixels sufficiently. The discriminable properties contained in the GLCM are described by six textural features, which include two new introduced ones and four independent ones. In addition, the statistical moments mean Me and standard deviation SD of textural features are used instead of themselves as elements in feature vector to reduce the dimensionality of feature vector and computational complexity. A support vector machine is employed for classification purpose. Experimental results show that the TF‐GLCM achieves the detection rates of 98% on CASIA v1.0, and 97% on CASIA v2.0 with 96‐D feature vector. The detection rates benefit from the two new textural features. Meanwhile, the TF‐GLCM is superior to some state‐of‐the‐art methods with lower dimension feature vector. Xuanjing Shen, Zenan Shi, Haipeng Chen 0002 |
IET Image Process. | 3 |
| 2017 | A novel automatic fuzzy clustering algorithm based on soft partition and membership information
Haipeng Chen 0002, Xuanjing Shen, Yingda Lv, Long Jian-Wu |
Neurocomputing | 1 |
| 2017 | Segmentation fusion based on neighboring information for MR brain images
Yuncong Feng, Xuanjing Shen, Haipeng Chen 0002, Xiaoli Zhang 0001 |
Multim. Tools Appl. | 3 |
| 2016 | A comparison of general-purpose distributed systems for data processingabstractGeneral-purpose distributed systems for data processing become popular in recent years due to the high demand from industry for big data analytics. However, there is a lack of comprehensive comparison among these systems and detailed analysis on their performance. In this paper, we conduct an extensive performance study on four state-of-the-art general-purpose distributed computing systems. Our results reveal useful insights on the design and implementation, which help the improvement of existing systems and the development of better new systems. James Cheng, Yunjian Zhao, Fan Yang 0091, Haipeng Chen 0002, Ruihao Zhao |
IEEE BigData | 6 |
| 2016 | Histogram-based colour image fuzzy clustering algorithm
Haipeng Chen 0002, Xuanjing Shen, Jianwu Long |
Multim. Tools Appl. | 1 |
| 2016 | Copy-move forgery detection based on scaled ORB
Xuanjing Shen, Haipeng Chen 0002 |
Multim. Tools Appl. | 3 |
| 2016 | A weighted-ROC graph based metric for image segmentation evaluation
Yuncong Feng, Xuanjing Shen, Haipeng Chen 0002, Xiaoli Zhang 0001 |
Signal Process. | 3 |
| 2011 | An improved image blind identification based on inconsistency in light source direction
Yingda Lv, Xuanjing Shen, Haipeng Chen 0002 |
J. Supercomput. | 3 |