EDBT 2026 Demo / reviewers in the wild / expert
Gang Yan 0001
dblp:87/7043-1
· DBLP profile ↗
16ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-7494-9353ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Context-aware token recovery and identity-guided global augmentation for occluded person re-identification
Gang Yan 0001, Shuze Geng |
Multim. Syst. | 1 |
| 2026 | FAPS-MER: Facial action position and semantic based interactive fusion for micro-expression recognition
Gang Yan 0001, Yubo He, Zian Liu, Yuqiang Guo, Shixin Cen |
Signal Process. Image Commun. | 1 |
| 2026 | ALL-IN-ONE: Divide-and-Conquer Strategy for Multi-Manipulation Image Classification and LocalizationabstractIn the advertising and media industries, image editing often involves multiple manipulation techniques to meet creative and technical requirements. Detecting tampered regions is crucial in scenarios like legal disputes or media integrity assessments. However, existing forensic methods often target single manipulation types or treat all manipulations as one, and many deep learning approaches lack flexibility in frequency and edge extraction, limiting their effectiveness. To address these challenges, this paper proposes an ALL-IN-ONE framework for comprehensive image forensic analysis, which adopts a divide-and-conquer strategy for multi-manipulation image classification and localization. Specifically, we introduce a Multi-Frequency Band Extraction Module (MBEM) to capture richer artifact information in the frequency domain. This is complemented by an Attention Window-based Fusion Module, which fuses same-frequency features across different scales and enhances the discriminative features more effectively. To improve the localization of copy-move manipulation, we design a Copy-Move Accurate Detection Module (CADM), which leverages the visual consistency between source and target regions. Furthermore, we propose a Precise Edge Generator (PEG) as part of the Edge-Guided Progressive Fine-Tune Module (EPFM), which can generate more accurate edge to enhance edge localization. To address the issue of insufficient labeled data, we construct a publicly available dataset, the Multi-Manipulation Image Dataset (MMID), consisting of 2,000 multi-manipulation images, each containing at least two types of forgeries. Extensive experiments are conducted, comparing our method with state-of-the-art approaches on MMID, as well as on single-manipulation datasets such as CASIA, CoMoFoD, and NIST. The results demonstrate that MMID is effective for training discriminative models and validate that our proposed method significantly outperforms existing approaches in terms of accuracy and robustness for simultaneous forgery localization and manipulation classification. Chang Ti, Gang Yan 0001, Yingchun Guo, Bin Li 0011 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Token recombination based shallow-deep feature fusion for occluded person re-identification
Shuze Geng, Gang Yan 0001, Haowei Wang 0001, Wenjie Xia |
Multim. Syst. | 3 |
| 2025 | TransStyle: Transformer-based StyleGAN for image inversion and editing
Yingchun Guo, Xueqi Lv, Gang Yan 0001, Shi Di |
Pattern Recognit. Lett. | 3 |
| 2025 | Pose-Skeleton Guided Cross-Attention Representation Fusion for Occluded Pedestrian Re-IdentificationabstractMost methods address occluded pedestrian Re-Identification (Re-ID) by employing external auxiliary models in the feature output stage of the backbone network to locate visible appearance areas. Nevertheless, these approaches suffer from issues such as occlusion information diffusion and imprecise masks generated by external models, indicating the need for further exploration in the decoupling of pedestrian features from occlusion information. In light of these challenges, we propose an innovative algorithm called Pose-Skeleton guided Cross-attention Representation fusion (PSCR) method. Firstly, we introduce the Visible Appearance Region Attention (VARA) model designed to leverage pose information for guiding the backbone network in effectively distinguishing between occlusion information and pedestrian features at the intermediate layer. By employing a suppression strategy, the model is able to effectively suppress occlusion interference and alleviate the diffusion of occlusion information. Next, to achieve precise localization of pedestrian-specific semantic regions, a groundbreaking Skeletal Area Modeling (SAM) is proposed. Leveraging the principles of mathematical modeling and capitalizing on the efficacy of human keypoint confidence, this module generates finely-grained masks for local skeleton regions and extracts an exhaustive set of local features. Lastly, under the constraints imposed by spatial attention masks, a cross-attention mechanism is employed to fuse the features acquired from the previous two steps with local features. This fusion process results in the generation of enhanced local features that seamlessly integrate aligning high-level semantic information. Extensive experimentation demonstrates that the proposed algorithm exhibits notable performance advancements when compared to existing methodologies. Shuze Geng, Zijin Wang, Gang Yan 0001, Yang Yu 0022, Yingchun Guo |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Continuous sign language recognition based on hierarchical memory sequence networkabstractAbstract With the goal of solving the problem of feature extractors lacking strong supervision training and insufficient time information concerning single‐sequence model learning, a hierarchical sequence memory network with a multi‐level iterative optimisation strategy is proposed for continuous sign language recognition. This method uses the spatial‐temporal fusion convolution network (STFC‐Net) to extract the spatial‐temporal information of RGB and Optical flow video frames to obtain the multi‐modal visual features of a sign language video. Then, in order to enhance the temporal relationships of visual feature maps, the hierarchical memory sequence network is used to capture local utterance features and global context dependencies across time dimensions to obtain sequence features. Finally, the decoder decodes the final sentence sequence. In order to enhance the feature extractor, the authors adopted a multi‐level iterative optimisation strategy to fine‐tune STFC‐Net and the utterance feature extractor. The experimental results on the RWTH‐Phoenix‐Weather multi‐signer 2014 dataset and the Chinese sign language dataset show the effectiveness and superiority of this method. Cui-Hong Xue, Jingli Jia, Ming Yu 0006, Gang Yan 0001, Yingchun Guo, Yuehao Liu |
IET Comput. Vis. | 4 |
| 2024 | A review of single image super-resolution reconstruction based on deep learning
Ming Yu 0006, Jiecong Shi, Cui-Hong Xue, Xiaoke Hao, Gang Yan 0001 |
Multim. Tools Appl. | 5 |
| 2024 | Age transformation based on deep learning: a survey
Yingchun Guo, Gang Yan 0001, Xueqi Lv |
Neural Comput. Appl. | 3 |
| 2024 | IRNet-RS: image retargeting network via relative saliency
Yingchun Guo, Xiaoke Hao, Gang Yan 0001 |
Neural Comput. Appl. | 4 |
| 2024 | Identity-Preserving Face Aging With Multi-Attribution Fusion and Multi-Scale AttentionabstractDespite the recent advances of face age transformation in the age accuracy of face synthesis or the integrity of identity information, only age used as the condition of aging is insufficient to depict the real aging process. In this work, we propose a novel GAN-based identity-preserving aging method with Multi-scale attention and Multi-attribution fusion (named M2aging) to address this problem. In particular, we design a High-fidelity Multi-scale Attention(HiMA) module to capture minute and local features at different scales, improving the model's identity-preserving ability. In addition, a Multi-attribution Fusion Modulator (MFM) is proposed by fusing age, gender, and race attributions to control the aging process, which enhances attribution preservation without introducing any additional classifiers and training losses. Experimental results show the proposed M2aging outperforms other SOTA techniques in terms of identity preservation and accurate age conversion. The ablation experiments manifest that HiMA can effectively heighten the fidelity of the transformed images, and MFM can effectively control the aging process according to age, gender, and race. Yingchun Guo, Mingyuan Li 0005, Gang Yan 0001 |
IEEE Signal Process. Lett. | 3 |
| 2023 | Continuous sign language recognition based on iterative alignment network and attention mechanism
Cui-Hong Xue, Ming Yu 0006, Gang Yan 0001, Yuehao Liu |
Multim. Tools Appl. | 3 |
| 2023 | Multi-operator Image Retargeting based on Saliency Object Ranking and Similarity Evaluation Metric
Yingchun Guo, Gang Yan 0001 |
Signal Process. Image Commun. | 4 |
| 2023 | Part-Based Representation Enhancement for Occluded Person Re-IdentificationabstractRetrieving an occluded pedestrian remains a challenging problem in person re-identification (re-id). Most existing methods utilize external detectors to disentangle the visible body parts. However, these methods are unstable due to domain bias and consume numerous computing resources. In this paper, we propose a novel and lightweight Part-based Representation Enhancement (PRE) network for occluded re-id that takes full advantages of the local correlations to aggregate distinctive information for local features without relying on auxiliary detectors. First, according to the information qualities of different body parts, we design a reasonable partition strategy to obtain the local features. Next, a Partial Relationship Aggregation (PRA) module is developed to self-mine the visibility of the body and construct a correlation matrix for collecting the information related to pre-defined classes. Following this, we propose an Inter-part Omnibearing Fusion (IOF) module that leverages the occlusion-suppressed class features to enhance the distinctiveness of the local features via feature completion and reverse fusion strategies. During the testing phase, the global and reconstructed local features are concatenated together for re-id without a complex visible region matching algorithm. Extensive experiments on occluded, partial, and holistic re-id benchmarks demonstrate the superiority of PRE over state-of-the-art methods in terms of accuracy and model complexity. Gang Yan 0001, Zijin Wang, Shuze Geng, Yang Yu 0022, Yingchun Guo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Multi-task Facial Activity Patterns Learning for micro-expression recognition using Joint Temporal Local Cube Binary Pattern
Shixin Cen, Yang Yu 0022, Gang Yan 0001, Ming Yu 0006, Yuqiang Guo |
Signal Process. Image Commun. | 3 |
| 2020 | AR-Net: Adaptive Attention and Residual Refinement Network for Copy-Move Forgery DetectionabstractIn copy-move forgery, the illumination and contrast of tampered and genuine regions are highly consistent, which poses a greater challenge in copy-move forgery detection. In this article, an end-to-end neural network is proposed based on adaptive attention and residual refinement network (AR-Net). Specifically, position and channel attention features are fused by the adaptive attention mechanism to fully capture context information and enrich the representation of features. Second, deep matching is adopted to compute the self-correlation between feature maps, and atrous spatial pyramid pooling fuses the scaled correlation maps to generate the coarse mask. Finally, the coarse mask is optimized through the residual refinement module, which retains the structure of object boundaries. Extensive experiments, evaluated on CASIAII, COVERAGE, and CoMoFoD datasets, demonstrate that the AR-Net has superior performance than state-of-the-art algorithms and can locate tampered and corresponding genuine regions at the pixel level. In addition, AR-Net has high robustness on postprocessing operations, such as noise, blur, and JPEG recompression. Gang Yan 0001, Yingchun Guo, Yongfeng Dong |
IEEE Trans. Ind. Informatics | 3 |