Yingda Lyu

dblp:226/5277 · DBLP profile ↗
← Back
34ranked-venue papers
2as first author
31since 2021 · last 2026
0000-0002-2037-6692ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 1 first-author · 22 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Causality-Aligned Semantic Recovery for Incomplete Cross-Modal Retrieval
abstract
Incomplete cross-modal retrieval (ICMR) requires models to recover missing modalities and robustly align heterogeneous ones for effective retrieval. Existing methods, however, fall short in both aspects. They often rely on limited semantic cues, such as single samples or coarse category prototypes, which compromises reconstruction quality. Moreover, these approaches are vulnerable to learning spurious cross-modal correlations, thereby impairing accurate alignment and hindering retrieval performance. To address these challenges, we propose Causality-Aligned Semantic Recovery (CASR), a novel method designed to both comprehensively restore missing modalities and mitigate spurious associations between vision and language. Our CASR involves two essential components: i) the Missing Modality Imagination (MMI) module, which combines category semantic priors with relevant contextual information to achieve high-quality semantic reconstruction; ii) the Explicit Causal Alignment (ECA) module, which explicitly learns environment-invariant attention, effectively eliminating the interference of spurious correlations and improving retrieval performance. Furthermore, we extend CASR to the challenging task of Partially Aligned Cross-Modal Retrieval, where we treat unlabeled unpaired data as a form of incomplete data. By leveraging MMI and ECA modules, we are able to learn robust representations in this setting. Extensive experiments on benchmark datasets under various missing rates demonstrate that CASR achieves superior robustness and retrieval performance.
Haipeng Chen 0002, Yu Liu 0004, Xun Yang 0001, Yuheng Liang, Yingda Lyu
AAAI5
2026 Dual Coding Theory in Action: Language-Assisted Human Pose Estimation in Videos
abstract
Video-based human pose estimation aims to localize keypoints across frames, enabling robust analysis of human motion in applications such as sports, surveillance, and healthcare. However, existing methods rely solely on visual cues, limiting their robustness in complex scenes involving occlusion, motion blur, or poor lighting. In contrast, dual coding theory from psychology suggests that human cognition is inherently multimodal: we learn by integrating visual perception with linguistic context to form structured, semantic understandings of the world. Visual input provides concrete spatiotemporal grounding, while language offers symbolic abstraction that enhances reasoning and generalization. Motivated by this cognitive principle, we present the first framework that explicitly incorporates language as an auxiliary modality to enhance video-based pose estimation. To address the lack of paired video-text datasets, we first employ a Multimodal Large Language Model (MLLM) to generate textual descriptions of human interactions from videos. We then propose a novel coarse-to-fine multimodal alignment pipeline: a cross-modal semantic interaction module establishes initial grounding between spatiotemporal visual features and textual embeddings, while an optimal transport-based feature matching mechanism enforces fine-grained, geometry-aware alignment. This cognitively inspired design enables more accurate and robust pose estimation, especially in visually challenging scenes like occlusion and motion blur. Extensive experiments on three benchmarks confirm that our method consistently outperforms state-of-the-art approaches.
Sifan Wu 0001, Haipeng Chen 0002, Yingda Lyu, Shaojing Fan, Zhenguang Liu, Yingying Jiao
AAAI3
2026 Attentive Keypoint Identification: Progressive Spatiotemporal Refinement for Video-based Human Pose Estimation
abstract
Video-based human pose estimation has vast applications such as action recognition, sports analytics, and crime detection. However, this task is challenging as it involves interpreting both spatial context and temporal dynamics to accurately localize human anatomical keypoints in video sequences. Current approaches, often based on attention mechanisms, perform well but struggle in challenging scenarios like rapid motion and pose occlusion. We attribute these failures to two fundamental limitations: spatial uniformity, where models indiscriminately assign attention to both joint-relevant features and background clutter, thereby introducing spatial noise; and temporal rigidity, an inability to adapt to large joint displacements, resulting in severe feature misalignment during rapid motion. To overcome these challenges, we introduce PSTPose, a novel progressive spatiotemporal refinement framework. Specifically, to address the spatial uniformity problem, we propose a Discriminative Feature Enhancement (DFE) module that emphasizes joint-relevant features and a Feature Cluster Grouping (FCG) module that forms compact, semantically meaningful regions. For the temporal rigidity problem, we introduce a Deformable Spatiotemporal Fusion (DSF) module that adaptively aligns features across consecutive frames via deformation-aware sampling. This design ensures robust keypoint localization, particularly in cluttered and dynamic scenes. Extensive experiments on three large-scale benchmarks, PoseTrack2017, PoseTrack2018, PoseTrack21, demonstrate that PSTPose establishes a new state-of-the-art.
Sifan Wu 0001, Haipeng Chen 0002, Yingda Lyu, Shaojing Fan, Zhenguang Liu, Yingying Jiao
AAAI3
2026 VGD: Value-Guided Diffusion Toward High-Utility Medical Image Segmentation
abstract
Progress in medical image segmentation is fundamentally constrained by the scarcity of annotated data. While diffusion models offer a promising solution by generating high-fidelity image–mask pairs, their utility for downstream tasks remains underexplored. A key bottleneck lies in the misalignment between generation outputs and task-specific needs—samples are produced independently of their utility for downstream training. To this end, we propose Value-Guided Diffusion (VGD), a lightweight sampling framework that integrates downstream model feedback into the generative inference process. VGD estimates a value score for each sample based on its utility to downstream training, and leverages this signal to iteratively guide the denoising trajectory toward high-reward regions of the data manifold. Crucially, VGD can be seamlessly integrated into existing medical diffusion models without any additional training or architectural modifications. Extensive experiments across multiple diffusion backbones and segmentation benchmarks demonstrate that VGD significantly boosts downstream segmentation performance while maintaining visual fidelity. Our findings highlight a task-aware sampling principle with potential to underpin future synthetic segmentation pipelines.
Haipeng Chen 0002, Chengxin Yang, Yingda Lyu
AAAI4
2026 Pedestrian-Centric Discriminative and Fine-grained Semantic Mining for Text-based Person Retrieval
abstract
Text-based Person Retrieval (TPR) aims to retrieve specific pedestrian images from a gallery based on the given textual descriptions, serving as a fine-grained instance of cross-modal retrieval on the Web. Current mainstream approaches primarily leverage pre-trained models and attention mechanisms to enhance multi-modal representations. Despite notable progress, they still struggle with two major challenges: 1) Intra-instance semantic asymmetry, which mainly derives from the partial semantic relevance conveyed by each image-text pair; and 2) Inter-instance semantic ambiguity, which arises from the high similarity of image-text pairs with different identities. These issues result in suboptimal semantic alignment and degraded retrieval accuracy. To this end, we propose a novel Pedestrian-Centric Discriminative and Fine-grained Semantic Mining (DFSM) framework for TPR. Specifically, our DFSM method comprises two essential components: 1) Text-aware Visual Refinement (TVR), which mitigates visual redundancy by selecting semantically relevant patches under textual guidance, and refines them via adaptive clustering and merging; 2) Token-level Semantic Alignment (TSA), which formulates the matching relationship between image regions and text words as a conditional transport (CT) problem, effectively mining fine-grained semantic differences and enhancing instance discrimination. Extensive experiments on four benchmarks validate the advantages of DFSM in terms of retrieval accuracy and visual interpretability.
Yuheng Liang, Haipeng Chen 0002, Yu Liu 0004, Yingda Lyu
WWW4
2026 UML: uncertainty-aware and mutual learning for noise-robust cross-lingual cross-modal retrieval
Yingda Lyu
Sci. China Inf. Sci.4
2026 Bayesian perturbation-driven consistency regularization for semi-supervised medical image segmentation
Haipeng Chen 0002, Yingda Lyu, Zenan Shi, Yongping Yang, Yu Wang 0112
Knowl. Based Syst.3
2026 Rethinking Skeleton-Based Action Recognition From Action-Class Prediction Distribution Perspective
abstract
Action recognition has long been a fundamental and compelling problem in the field of computer vision. However, one aspect that has been overlooked so far is that current action recognition approaches often produce an unfavourable multi-peaked distribution when identifying the action class of a given motion sequence, which is ambiguous and hard to learn for neural networks. Moreover, current methods heavily rely on neural networks to extract action features for differentiating actions, lacking theoretical constraints ensuring that action-specific features are selectively extracted and ambiguous features common to multiple actions are effectively reduced. These shortcomings culminate in inadequate action recognition accuracy. Motivated by this, in this paper we seek to tackle the problem from three aspects: 1) We try to eliminate ambiguity by enforcing a smooth single-peaked distribution instead of a multi-peaked one for action-class prediction. 2) We theoretically analyze the lower bound of the label prediction log-likelihood and derive a training objective, which focuses on the extraction of action-specific features and the reduction of ambiguous features. 3) We further advocate feeding the model with richer information, including positive information like body-part structures and negative information like masked inputs. Empirically, our approach sets the new state-of-the-art performance on five large-scale benchmarks. Our code is released at https://github.com/ActionR-Group/DPM to facilitate future research.
Yingying Jiao, Haipeng Chen 0002, Yingda Lyu, Shuang Wu 0002, Zhenguang Liu
IEEE Trans. Image Process.3
2026 Discriminative Representation Learning for Remote Sensing Visual Question Answering
abstract
Recently, Remote Sensing Visual Question Answering (RSVQA) has attracted increasing attention from both academia and industry, which is the basis for understanding the underlying correspondence between remote sensing imagery and text descriptions. However, current methods are still insufficient in learning discriminative visual and textual representations for answer reasoning, mainly due to two reasons: (1) the remote sensing image environment is complex and changeable, and the target scales vary significantly, making it difficult to extract discriminative visual features; and (2) there is a lack of effective guidance from remote sensing domain knowledge to learn discriminative features. To this end, we propose a D iscriminative R epresentation L earning (DRL) method that includes two key strategies: visual feature enhancement and prior knowledge guidance. Specifically, we employ the Fourier transform to simulate the diverse visual environment and force the model to mine discriminative visual representations by imposing consistency constraints with the original features. In addition, we leverage the Remote Sensing Multimodal Large Language Model (RSMLLM) to generate captions rich in remote sensing domain-specific prior knowledge. These captions, derived from RSMLLM’s powerful knowledge integration and summarization capabilities, can then be compared and fused with visual representations to generate more discriminative representations, which are ultimately used for answer reasoning. Finally, recognizing that most existing RSVQA methods rely solely on static remote sensing images, we introduce RSVideoQA, a novel satellite video question answering dataset. This dataset is designed to facilitate the exploration of the rich spatio-temporal dynamics inherent in video sequences. Experimental results across three distinct datasets validate the effectiveness of our proposed method. Our dataset and code will be released at https://github.com/chill-han/DRL .
Yingda Lyu, Yu Liu 0004, Haipeng Chen 0002
ACM Trans. Multim. Comput. Commun. Appl.1
2025 Causal-Inspired Multitask Learning for Video-Based Human Pose Estimation
abstract
Video-based human pose estimation has long been a fundamental yet challenging problem in computer vision. Previous studies focus on spatio-temporal modeling through the enhancement of architecture design and optimization strategies. However, they overlook the causal relationships in the joints, leading to models that may be overly tailored and thus estimate poorly to challenging scenes. Therefore, adequate causal reasoning capability, coupled with good interpretability of model, are both indispensable and prerequisite for achieving reliable results. In this paper, we pioneer a causal perspective on pose estimation and introduce a causal-inspired multitask learning framework, consisting of two stages. In the first stage, we try to endow the model with causal spatio-temporal modeling ability by introducing two self-supervision auxiliary tasks. Specifically, these auxiliary tasks enable the network to infer challenging keypoints based on observed keypoint information, thereby imbuing causal reasoning capabilities into the model and making it robust to challenging scenes. In the second stage, we argue that not all feature tokens contribute equally to pose estimation. Prioritizing causal (keypoint-relevant) tokens is crucial to achieve reliable results, which could improve the interpretability of the model. To this end, we propose a Token Causal Importance Selection module to identify the causal tokens and non-causal tokens (e.g., background and objects). Additionally, non-causal tokens could provide potentially beneficial cues but may be redundant. We further introduce a non-causal tokens clustering module to merge the similar non-causal tokens. Extensive experiments show that our method outperforms state-of-the-art methods on three large-scale benchmark datasets.
Haipeng Chen 0002, Sifan Wu 0001, Yifang Yin, Yingying Jiao, Yingda Lyu, Zhenguang Liu
AAAI6
2025 Skeleton-based Action Recognition with Non-linear Dependency Modeling and Hilbert-Schmidt Independence Criterion
abstract
Human skeleton-based action recognition has long been an indispensable aspect of artificial intelligence. Current state-of-the-art methods tend to consider only the dependencies between connected skeletal joints, limiting their ability to capture non-linear dependencies between physically distant joints. Moreover, most existing approaches distinguish action classes by estimating the probability density of motion representations, yet the high-dimensional nature of human motions invokes inherent difficulties in accomplishing such measurements. In this paper, we seek to tackle these challenges from two directions: (1) We propose a novel dependency refinement approach that explicitly models dependencies between any pair of joints, effectively transcending the limitations imposed by joint distance. (2) We further propose a framework that utilizes the Hilbert-Schmidt Independence Criterion to differentiate action classes without being affected by data dimensionality, and mathematically derive learning objectives guaranteeing precise recognition. Empirically, our approach sets the state-of-the-art performance on NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA datasets.
Haipeng Chen 0002, Yingda Lyu
AAAI3
2025 Enhancing Semantic Clarity: Discriminative and Fine-grained Information Mining for Remote Sensing Image-Text Retrieval
abstract
Remote sensing image-text retrieval is a fundamental task in remote sensing multimodal analysis, promoting the alignment of visual and language representations. The mainstream approaches commonly focus on capturing shared semantic representations between visual and textual modalities. However, the inherent characteristics of remote sensing image-text pairs lead to a semantic confusion problem, stemming from redundant visual representations and high inter-class similarity. To tackle this problem, we propose a novel Discriminative and Fine-grained Information Mining (DFIM) model, which aims to enhance semantic clarity by reducing visual redundancy and increasing the semantic gap between different classes. Specifically, the Dynamic Visual Enhancement (DVE) module adaptively enhances the visual discriminative features under the guidance of multimodal fusion information. Meanwhile, the Fine-grained Semantic Matching (FSM) module cleverly models the matching relationship between image regions and text words as an optimal transport problem, thereby refining intra-instance matching. Extensive experiments on two benchmark datasets justify the superiority of DFIM in terms of retrieval accuracy and visual interpretability over the leading methods.
Yu Liu 0004, Haipeng Chen 0002, Yuheng Liang, Xun Yang 0001, Yingda Lyu
IJCAI6
2025 Dataset-level color augmentation and multi-scale exploration methods for polyp segmentation
Haipeng Chen 0002, Honghong Ju, Jincai Song, Yingda Lyu, Xianzhu Liu
Expert Syst. Appl.5
2025 Rethinking Polyp Segmentation from the Perspectives of Matching Views and Seeking Camouflage
Zhengfang Jiang, Haipeng Chen 0002, Yongping Yang, Xianzhu Liu, Yingda Lyu
Multim. Syst.5
2025 Multi-modality boundary-guided network for generalizable image manipulation localization
Yanyan Jiang 0005, Haipeng Chen 0002, Yingda Lyu
Multim. Syst.4
2025 Causality-Inspired Unsupervised Domain Adaptation With Target Style Imitation for Medical Image Segmentation
abstract
Deep learning performance may decrease substantially with unseen heterogeneous data. While most unsupervised domain adaptation (UDA) methods seek to address this through image alignment, they often ignore uncertainty style fluctuations within the target domain. When testing image styles vary in both direction and intensity, such models may fail to adapt. Furthermore, existing UDA methods tend to over-reliance on domain-level entire feature alignment, resulting in potentially over-exploiting semantic content-independent cues (e.g., intensity) as shortcut features. To address these limitations, this paper introduces an innovative and model-agnostic Causality-inspired Representation Learning Based on Target Style Imitation method for UDA. Specifically, we propose a novel Target Style Imitation (TSI) data augmentation approach to diversify the training data and align training and unseen target testing image styles. TSI constructs a Gaussian distribution for the target domain style and simulates unseen testing style variations through random sampling. Additionally, inspired by the stable and generalizable causal mechanism, we propose Causality-inspired Representation Learning (CRL) based on TSI method to enforce feature representations to adhere to causal properties (i.e., Separation and Independence) essential for robust UDA, thus fostering the model to focus on the domain-invariant semantic features. Our method surpasses state-of-the-art methods on two cross-modality medical image segmentation datasets.
Jincai Song, Haipeng Chen 0002, Yingda Lyu, Weizhi Nie, Anan Liu
IEEE Trans. Circuits Syst. Video Technol.3
2025 User Invariant Preference Learning for Multi-Behavior Recommendation
abstract
In multi-behavior recommendation scenarios, analyzing users’ diverse behaviors, such as click , purchase , and rating , enables a more comprehensive understanding of their interests, facilitating personalized and accurate recommendations. A fundamental assumption of multi-behavior recommendation methods is the existence of shared user preferences across behaviors, representing users’ intrinsic interests. Based on this assumption, existing approaches aim to integrate information from various behaviors to enrich user representations. However, they often overlook the presence of both commonalities and individualities in users’ multi-behavior preferences. These individualities reflect distinct aspects of preferences captured by different behaviors, where certain auxiliary behaviors may introduce noise, hindering the prediction of the target behavior. To address this issue, we propose a user invariant preference learning (UIPL) for multi-behavior recommendation, aiming to capture users’ intrinsic interests (referred to as invariant preferences) from multi-behavior interactions to mitigate the introduction of noise. Specifically, UIPL leverages the paradigm of invariant risk minimization to learn invariant preferences. To implement this, we employ a variational autoencoder (VAE) to extract users’ invariant preferences, replacing the standard reconstruction loss with an invariant risk minimization constraint. Additionally, we construct distinct environments by combining multi-behavior data to enhance robustness in learning these preferences. Finally, the learned invariant preferences are used to provide recommendations for the target behavior. Extensive experiments on four real-world datasets demonstrate that UIPL significantly outperforms current state-of-the-art methods.
Mingshi Yan, Zhiyong Cheng 0001, Fan Liu 0008, Yingda Lyu, Yahong Han
ACM Trans. Inf. Syst.4
2024 Rethinking Human Motion Prediction with Symplectic Integral
abstract
Long-term and accurate forecasting is the long-standing pursuit of the human motion prediction task. Existing methods typically suffer from dramatic degradation in prediction accuracy with increasing prediction horizon. It comes down to two reasons: 1) Insufficient numerical stability caused by unforeseen high noise and complex feature relationships in the data, and 2) Inadequate modeling stability caused by unreasonable step sizes and undesirable parameter updates in the prediction. In this paper, we design a novel and sym-plectic integral-inspired framework named symplectic integral neural network (SINN), which engages symplectic tra-jectories to optimize the pose representation and employs a stable symplectic operator to alternately model the dynamic context. Specifically, we design a Symplectic Repre-sentation Encoder that performs on enhanced human pose representation to obtain trajectories on the symplectic manifold, ensuring numerical stability based on Hamiltonian mechanics and symplectic spatial splitting algorithm. We further present the Symplectic Temporal Aggregation mod-ule, which splits the long-term prediction into multiple ac-curate short-term predictions generated by a symplectic operator to secure modeling stability. Moreover, our approach is model-agnostic and can be efficiently integrated with different physical dynamics models. The experimental results demonstrate that our method achieves the new state-of-the-art, outperforming existing methods by 20.1% on Human3.6M, 16.7% on CUM Mocap, and 10.2% on 3DPW.
Haipeng Chen 0002, Kedi Lyu, Zhenguang Liu, Yifang Yin, Xun Yang 0001, Yingda Lyu
CVPR6
2024 CRML-Net: Cross-Modal Reasoning and Multi-Task Learning Network for tooth image segmentation
Yingda Lyu, Zhehao Liu, Haipeng Chen 0002
Comput. Vis. Image Underst.1
2024 Reinforced visual interaction fusion radiology report generation
Haipeng Chen 0002, Yu Liu 0004, Yingda Lyu
Multim. Syst.4
2024 View sequence prediction GAN: unsupervised representation learning for 3D shapes by decomposing view content and viewpoint variance
Heyu Zhou, Jiayu Li 0004, Xianzhu Liu, Yingda Lyu, Haipeng Chen 0002, Anan Liu
Multim. Syst.4
2024 MCA-Net: multi-cascade attention network for polyp segmentation
Xuanjing Shen, Yingda Lyu
Multim. Tools Appl.3
2024 Regular Constrained Multimodal Fusion for Image Captioning
abstract
More diverse and closer to human-like captions are of paramount importance in image captioning. Recent research has achieved significant advancements, with the majority adopting end-to-end encoder-decoder architectures that integrate specific feature-text processing. However, the homogeneity of their model structures, the simplicity or complexity of feature-text fusion, and the uniformity of training objectives have all to some extent affected the diversity and effectiveness of caption generation, thus limiting the potential applications of this task. Therefore, in this paper, we propose the Regular Constrained Multimodal Fusion (RCMF) method for image captioning to better integrate information across and within modalities, while also approaching human-like fine-grained semantic perception and relationship reasoning capabilities. Initially, our RCMF preprocesses images using a Swin-Transformer and then an extended encoder with a new intra-modal fusion module, utilizing window-focused linear attention to capture features and leveraging refined grid and global visual features. By combining text features, RCMF employs a cross-modal fusion module and decoder to deeply model the interaction between text and image. Additionally, RCMF first introduces a new additional regulatory modal fusion reasoning (MFR) branch, which surpasses the above architectures. Its MFR loss combined with cross-entropy loss forms a new training objective strategy, effectively mining fine-grained relationships between images and text, perceiving the semantic information of images and their corresponding captions, thereby regulating the generated captions to be more diverse and human-like. Experimental results based on the MS COCO 2014 dataset, particularly under the same experimental conditions, demonstrate the outstanding performance of our method, especially in terms of METEOR, ROUGE-L, CIDEr, and SPICE metrics. Visualization results further intuitively confirm the effectiveness of our RCMF method. Source code inhttps://github.com/200084/RCMF-for-image-caption.
Haipeng Chen 0002, Yu Liu 0004, Yingda Lyu
IEEE Trans. Circuits Syst. Video Technol.4
2023 Action Recognition with Multi-stream Motion Modeling and Mutual Information Maximization
abstract
Action recognition has long been a fundamental and intriguing problem in artificial intelligence. The task is challenging due to the high dimensionality nature of an action, as well as the subtle motion details to be considered. Current state-of-the-art approaches typically learn from articulated motion sequences in the straightforward 3D Euclidean space. However, the vanilla Euclidean space is not efficient for modeling important motion characteristics such as the joint-wise angular acceleration, which reveals the driving force behind the motion. Moreover, current methods typically attend to each channel equally and lack theoretical constrains on extracting task-relevant features from the input. In this paper, we seek to tackle these challenges from three aspects: (1) We propose to incorporate an acceleration representation, explicitly modeling the higher-order variations in motion. (2) We introduce a novel Stream-GCN network equipped with multi-stream components and channel attention, where different representations (i.e., streams) supplement each other towards a more precise action recognition while attention capitalizes on those important channels. (3) We explore feature-level supervision for maximizing the extraction of task-relevant information and formulate this into a mutual information loss. Empirically, our approach sets the new state-of-the-art performance on three benchmark datasets, NTU RGB+D, NTU RGB+D 120, and NW-UCLA.
Haipeng Chen 0002, Zhenguang Liu, Yingda Lyu, Beibei Zhang 0007, Shuang Wu 0002, Zhibo Wang 0001, Kui Ren 0001
IJCAI4
2023 A complementary and contrastive network for stimulus segmentation and generalization
Na Ta 0009, Haipeng Chen 0002, Yingda Lyu, Zenan Shi, Zhehao Liu
Image Vis. Comput.3
2023 BLE-Net: boundary learning and enhancement network for polyp segmentation
Na Ta 0009, Haipeng Chen 0002, Yingda Lyu, Taosuo Wu
Multim. Syst.3
2023 PL-GNet: Pixel Level Global Network for detection and localization of image forgeries
Zenan Shi, Xuanjing Shen, Haipeng Chen 0002, Yingda Lyu
Signal Process. Image Commun.4
2023 UP-Net: Uncertainty-Supervised Parallel Network for Image Manipulation Localization
abstract
Image manipulation localization remains a hot topic due to its inherent semantic-independent nature and realistic needs. Virtually all localization studies are devoted to solving arbitrary tampering using multi-branch networks based on deep features or skip-connection structures based on full features, which may induce the loss of manipulation details or noisy interference from image semantics. This poses a challenge for existing localization methods to fully capture invisible manipulations, especially in post-processing settings and across dataset scenarios. To address the above issues, we propose an uncertainty-supervised parallel network (UP-Net) for image tampering localization that preserves more manipulation details while avoiding semantic noise. UP-Net cascades the frequency and RGB domains of the manipulated image as dual-domain embedding, instead of dual-domain parallel learning as in previous work. To learn semantic-independent manipulation features, two structurally identical parallel branches are designed to learn tampering inconsistencies from intermediate and deep coding features for gradually obtaining the initial and final localization predictions. Where attention-guided partial decoder (AGPD) integrates more precise manipulation edges and manipulation semantics without introducing additional noise by focusing on channel correlation and spatial dependence, making a significant contribution to performance. Moreover, the new concept of uncertainty-constrained loss supervision is introduced to guide UP-Net to continuously improve confidence in locating difficult pixels, which are easily misclassified due to post-processing operations. Experiments on three public manipulation datasets and two real challenge datasets show that our end-to-end UP-Net achieves significant performance in manipulation localization, generalization across datasets, and robustness compared to state-of-the-art methods.
Dengyun Xu, Xuanjing Shen, Yingda Lyu
IEEE Trans. Circuits Syst. Video Technol.3
2022 MC-Net: Learning mutually-complementary features for image manipulation localization
abstract
Deep learning has become an emerging technical for image manipulation localization, which can automatically recognize abnormal traces caused by manipulation. However, as manipulations mainly happens in the foreground regions, these methods largely focus on the foreground contents and neglect the background, which contain complementary signal for fully understanding the image and are meaningful for manipulation localization. We propose a Mutually-Complementary Network (MC-Net), which is a two-branch network to operate the foreground and background features, respectively. To distill complementary signals from the features, we propose a mutual attentive module composed of self-feature attentive, and cross-feature attentive components to advance the communication across the foreground and background branches. Extensive qualitative and quantitative experiments demonstrate that our proposed MC-Net distinctly improves the prediction of foreground and background, obtains consistent performance increments on four benchmark data sets, and significantly outperforms the state-of-the-art methods.
Dengyun Xu, Xuanjing Shen, Yingda Lyu, Xiaoyu Du 0002, Fuli Feng
Int. J. Intell. Syst.3
2022 Hybrid features and semantic reinforcement network for image forgery detection
Haipeng Chen 0002, Chaoqun Chang, Zenan Shi, Yingda Lyu
Multim. Syst.4
2021 CF Model: A Coarse-to-Fine Model Based on Two-Level Local Search for Image Copy-Move Forgery Detection
abstract
Copy-move forgery is the most predominant forgery technique in the field of digital image forgery. Block-based and interest-based are currently the two mainstream categories for copy-move forgery detection methods. However, block-based algorithm lacks the ability to resist affine transformation attacks, and interest point-based algorithm is limited to accurately locate the tampered region. To tackle these challenges, a coarse-to-fine model (CFM) is proposed. By extracting features, affine transformation matrix and detecting forgery regions, the localization of tampered areas from sparse to precise is realized. Specifically, in order to further exactly extract the forged regions and improve performance of the model, a two-level local search algorithm is designed in the refinement stage. In the first level, the image blocks are used as search units for feature matching, and the second level is to refine the edge of the region at pixel level. The method maintains a good balance between the complexity and effectiveness of forgery detection, and the experimental results show that it has a better detection effect than the traditional interest-based copy and move forgery detection method. In addition, CFM method has high robustness on postprocessing operations, such as scaling, rotation, noise, and JPEG compression.
Fang Mei, Tianchang Gao, Yingda Lyu
Secur. Commun. Networks3
2020 Global Semantic Consistency Network for Image Manipulation Detection
abstract
This letter focuses on image manipulation detection which aims to recognize the manipulated regions under the contextual semantic information. Existing approaches usually overlook the semantic discrepancy between different levels of feature maps, and directly fuse (e.g., addition, or concatenation) them for detection. In this letter, we argue that the semantic gap is the main reason for the low effectiveness of feature fusion in manipulation predictions. To address this problem, we propose a Global Semantic Consistency Network (GSCNet) for image manipulation detection, which is based on an encoder-decoder structure. Specifically, to make GSCNet include more global texture information which has been empirically confirmed to be beneficial to manipulation detection, gram block is first deployed on each level of feature maps in the encoding stage. Based on that, bi-directional convolutional LSTM is further implemented on the decoding stage, such that feature maps of the same level have semantic consistency. Experimental results on NIST16, and CASIA v1.0 declare that GSCNet can accurately locate the manipulated regions. Furthermore, compared to the existing models, GSCNet can achieve new state-of-the-art results.
Zenan Shi, Xuanjing Shen, Haipeng Chen 0002, Yingda Lyu
IEEE Signal Process. Lett.4
2019 Automatic Segmentation of Brain Tumor Image Based on Region Growing with Co-constraint
Siming Cui, Xuanjing Shen, Yingda Lyu
MMM (1)3
2018 Image splicing detection based on Markov features in discrete octonion cosine transform domain
abstract
To improve the poor robustness and low accuracy of the existing algorithms of image splicing detection, a novel passive image forgery detection method is proposed in this study, which is based on DOCT (discrete octonion cosine transform) and Markov. By introducing the octonion and DOCT, the colour information of six image channels (the RGB model and the HSI model) can be exhaustively extracted, which enhances the robustness of the algorithm. On the issue of improving the detection accuracy, the standard deviation is used to characterise the relationship of the colour information between the parts of DOCT coefficient matrix, and the K ‐fold cross‐validation is introduced to improve the identification performance of the classifier. The steps of the algorithm are as follows: Firstly, the 8 × 8 block DOCT transform is used to the original image to obtain parts of block DOCT coefficient. Secondly, the standard deviation is used to process the corresponding parts of all blocks of the image. Finally, the Markov feature vector of the DOCT coefficient is extracted and feds to the LIBSVM (a library for support vector machines). When using LIBSVM for classification, K ‐fold cross‐validation is executed to select the best parameter pairs. The experiment results demonstrate that the algorithm is superior to the other state‐of‐the‐art splicing detection methods.
Hongda Sheng, Xuanjing Shen, Yingda Lyu, Zenan Shi, Shuyang Ma
IET Image Process.3