EDBT 2026 Demo / reviewers in the wild / expert
Yingchun Guo
dblp:137/0027 · also Ying-Chun Guo
· DBLP profile ↗
23ranked-venue papers
10as first author
19since 2021 · last 2026
0000-0002-3239-3086ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 7 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sign language translation via cross-modal alignment and graph convolution
Cui-Hong Xue, Yingchun Guo |
Neurocomputing | 4 |
| 2026 | ALL-IN-ONE: Divide-and-Conquer Strategy for Multi-Manipulation Image Classification and LocalizationabstractIn the advertising and media industries, image editing often involves multiple manipulation techniques to meet creative and technical requirements. Detecting tampered regions is crucial in scenarios like legal disputes or media integrity assessments. However, existing forensic methods often target single manipulation types or treat all manipulations as one, and many deep learning approaches lack flexibility in frequency and edge extraction, limiting their effectiveness. To address these challenges, this paper proposes an ALL-IN-ONE framework for comprehensive image forensic analysis, which adopts a divide-and-conquer strategy for multi-manipulation image classification and localization. Specifically, we introduce a Multi-Frequency Band Extraction Module (MBEM) to capture richer artifact information in the frequency domain. This is complemented by an Attention Window-based Fusion Module, which fuses same-frequency features across different scales and enhances the discriminative features more effectively. To improve the localization of copy-move manipulation, we design a Copy-Move Accurate Detection Module (CADM), which leverages the visual consistency between source and target regions. Furthermore, we propose a Precise Edge Generator (PEG) as part of the Edge-Guided Progressive Fine-Tune Module (EPFM), which can generate more accurate edge to enhance edge localization. To address the issue of insufficient labeled data, we construct a publicly available dataset, the Multi-Manipulation Image Dataset (MMID), consisting of 2,000 multi-manipulation images, each containing at least two types of forgeries. Extensive experiments are conducted, comparing our method with state-of-the-art approaches on MMID, as well as on single-manipulation datasets such as CASIA, CoMoFoD, and NIST. The results demonstrate that MMID is effective for training discriminative models and validate that our proposed method significantly outperforms existing approaches in terms of accuracy and robustness for simultaneous forgery localization and manipulation classification. Chang Ti, Gang Yan 0001, Yingchun Guo, Bin Li 0011 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Style-xLSTM for Facial Age EditingabstractFace age editing typically utilizes GAN or StyleGAN as the underlying framework, employing CNN or Transformer to engineer sophisticated modules for the subsequent processing of facial latent codes. Notwithstanding the laudable endeavors of researchers in this domain, these methods still face challenges, including facial attribute entanglement and suboptimal age modification. For instance, the process of editing may result in the reversal of gender, the presence of discernible artifacts, and the inability to achieve the intended effect regarding age. To address the aforementioned limitations, inspired by the newly proposed Extended Long Short-Term Memory (xLSTM), we explore the application of xLSTM in the facial age editing task for the first time. With the well-designed Style-mLSTM module, we achieve accurate retention of facial attributes while seamlessly transitioning between successive age modifications. In addition, considering the negative impact of the imbalanced age distribution in the existing dataset on the editing effect of different age groups, we propose a Weighted Age Distribution Calibration (WADC) mechanism to effectively mitigate this problem. Qualitative and quantitative experiments demonstrate that our approach outperforms existing techniques on several evaluation metrics and effectively demonstrates the great potential of xLSTM in facial attribute editing tasks. Mingyuan Li 0005, Songcheng Xu, Renda Han, Yingchun Guo |
IJCNN | 5 |
| 2025 | Instance style-aware transformer for domain generalizable person re-identification
Yingchun Guo, Shi Di, Xueqi Lv |
Neurocomputing | 1 |
| 2025 | DSM-CLIP: A framework designed for hard negative sampling in generalizable person re-identification
Yingchun Guo |
Neurocomputing | 1 |
| 2025 | TransStyle: Transformer-based StyleGAN for image inversion and editing
Yingchun Guo, Xueqi Lv, Gang Yan 0001, Shi Di |
Pattern Recognit. Lett. | 1 |
| 2025 | Pose-Skeleton Guided Cross-Attention Representation Fusion for Occluded Pedestrian Re-IdentificationabstractMost methods address occluded pedestrian Re-Identification (Re-ID) by employing external auxiliary models in the feature output stage of the backbone network to locate visible appearance areas. Nevertheless, these approaches suffer from issues such as occlusion information diffusion and imprecise masks generated by external models, indicating the need for further exploration in the decoupling of pedestrian features from occlusion information. In light of these challenges, we propose an innovative algorithm called Pose-Skeleton guided Cross-attention Representation fusion (PSCR) method. Firstly, we introduce the Visible Appearance Region Attention (VARA) model designed to leverage pose information for guiding the backbone network in effectively distinguishing between occlusion information and pedestrian features at the intermediate layer. By employing a suppression strategy, the model is able to effectively suppress occlusion interference and alleviate the diffusion of occlusion information. Next, to achieve precise localization of pedestrian-specific semantic regions, a groundbreaking Skeletal Area Modeling (SAM) is proposed. Leveraging the principles of mathematical modeling and capitalizing on the efficacy of human keypoint confidence, this module generates finely-grained masks for local skeleton regions and extracts an exhaustive set of local features. Lastly, under the constraints imposed by spatial attention masks, a cross-attention mechanism is employed to fuse the features acquired from the previous two steps with local features. This fusion process results in the generation of enhanced local features that seamlessly integrate aligning high-level semantic information. Extensive experimentation demonstrates that the proposed algorithm exhibits notable performance advancements when compared to existing methodologies. Shuze Geng, Zijin Wang, Gang Yan 0001, Yang Yu 0022, Yingchun Guo |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Disentangled Lifespan Synthesis via Transformer-Based Nonlinear RegressionabstractAbstract Lifespan face age transformation aims to generate facial images that accurately depict an individual's appearance at different age stages. This task is highly challenging due to the need for reasonable changes in facial features while preserving identity characteristics. Existing methods tend to synthesize unsatisfactory results, such as entangled facial attributes and low identity preservation, especially when dealing with large age gaps. Furthermore, over‐manipulating the style vector may deviate it from the latent space and damage image quality. To address these issues, this paper introduces a novel nonlinear regression model‐Disentangled Lifespan face Aging (DL‐Aging) to achieve high‐quality age transformation images. Specifically, we propose an age modulation encoder to extract age‐related multi‐scale facial features as key and value, and use the reconstructed style vector of the image as the query. The multi‐head cross‐attention in the W+ space is utilized to update the query for aging image reconstruction iteratively. This nonlinear transformation enables the model to learn a more disentangled mode of transformation, which is crucial for alleviating facial attribute entanglement. Additionally, we introduce a W+ space age regularization term to prevent excessive manipulation of the style vector and ensure it remains within the W+ space during transformation, thereby improving generation quality and aging accuracy. Extensive qualitative and quantitative experiments demonstrate that the proposed DL‐Aging outperforms state‐of‐the‐art methods regarding aging accuracy, image quality, attribute disentanglement, and identity preservation, especially for large age gaps. Mingyuan Li 0005, Yingchun Guo |
Comput. Graph. Forum | 2 |
| 2024 | Continuous sign language recognition based on hierarchical memory sequence networkabstractAbstract With the goal of solving the problem of feature extractors lacking strong supervision training and insufficient time information concerning single‐sequence model learning, a hierarchical sequence memory network with a multi‐level iterative optimisation strategy is proposed for continuous sign language recognition. This method uses the spatial‐temporal fusion convolution network (STFC‐Net) to extract the spatial‐temporal information of RGB and Optical flow video frames to obtain the multi‐modal visual features of a sign language video. Then, in order to enhance the temporal relationships of visual feature maps, the hierarchical memory sequence network is used to capture local utterance features and global context dependencies across time dimensions to obtain sequence features. Finally, the decoder decodes the final sentence sequence. In order to enhance the feature extractor, the authors adopted a multi‐level iterative optimisation strategy to fine‐tune STFC‐Net and the utterance feature extractor. The experimental results on the RWTH‐Phoenix‐Weather multi‐signer 2014 dataset and the Chinese sign language dataset show the effectiveness and superiority of this method. Cui-Hong Xue, Jingli Jia, Ming Yu 0006, Gang Yan 0001, Yingchun Guo, Yuehao Liu |
IET Comput. Vis. | 5 |
| 2024 | Age transformation based on deep learning: a survey
Yingchun Guo, Gang Yan 0001, Xueqi Lv |
Neural Comput. Appl. | 1 |
| 2024 | IRNet-RS: image retargeting network via relative saliency
Yingchun Guo, Xiaoke Hao, Gang Yan 0001 |
Neural Comput. Appl. | 1 |
| 2024 | Identity-Preserving Face Aging With Multi-Attribution Fusion and Multi-Scale AttentionabstractDespite the recent advances of face age transformation in the age accuracy of face synthesis or the integrity of identity information, only age used as the condition of aging is insufficient to depict the real aging process. In this work, we propose a novel GAN-based identity-preserving aging method with Multi-scale attention and Multi-attribution fusion (named M2aging) to address this problem. In particular, we design a High-fidelity Multi-scale Attention(HiMA) module to capture minute and local features at different scales, improving the model's identity-preserving ability. In addition, a Multi-attribution Fusion Modulator (MFM) is proposed by fusing age, gender, and race attributions to control the aging process, which enhances attribution preservation without introducing any additional classifiers and training losses. Experimental results show the proposed M2aging outperforms other SOTA techniques in terms of identity preservation and accurate age conversion. The ablation experiments manifest that HiMA can effectively heighten the fidelity of the transformed images, and MFM can effectively control the aging process according to age, gender, and race. Yingchun Guo, Mingyuan Li 0005, Gang Yan 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Progressive Mask Transformer With Edge Enhancement for Image Manipulation LocalizationabstractRecent developments in image editing techniques have given rise to serious challenges to the credibility of multimedia data. Although some deep learning methods have achieved impressive results, they often fail to detect subtle edge artefacts, and current mainstream methods focus mainly on the foreground content and ignore the background content, which also contains abundant information related to manipulation. To address this issue, this letter proposes a progressive mask transformer with an edge enhancement network for image manipulation localization. Specifically, an edge enhancement flow is introduced to detect subtle manipulated edge artefacts and guide the localization of manipulated regions. Then, the manipulated, genuine and global features are progressively refined using a progressive mask transformer module. We perform extensive experiments on NIST16, Coverage, CASIA and IMD20 datasets to verify the effectiveness of our method, and the results demonstrate that the proposed method outperforms state-of-the-art methods by a wide margin based on on commonly used evaluation metrics. Yang Yu 0022, Yingchun Guo, Xiaoke Hao |
IEEE Signal Process. Lett. | 4 |
| 2023 | Multi-operator Image Retargeting based on Saliency Object Ranking and Similarity Evaluation Metric
Yingchun Guo, Gang Yan 0001 |
Signal Process. Image Commun. | 1 |
| 2023 | Part-Based Representation Enhancement for Occluded Person Re-IdentificationabstractRetrieving an occluded pedestrian remains a challenging problem in person re-identification (re-id). Most existing methods utilize external detectors to disentangle the visible body parts. However, these methods are unstable due to domain bias and consume numerous computing resources. In this paper, we propose a novel and lightweight Part-based Representation Enhancement (PRE) network for occluded re-id that takes full advantages of the local correlations to aggregate distinctive information for local features without relying on auxiliary detectors. First, according to the information qualities of different body parts, we design a reasonable partition strategy to obtain the local features. Next, a Partial Relationship Aggregation (PRA) module is developed to self-mine the visibility of the body and construct a correlation matrix for collecting the information related to pre-defined classes. Following this, we propose an Inter-part Omnibearing Fusion (IOF) module that leverages the occlusion-suppressed class features to enhance the distinctiveness of the local features via feature completion and reverse fusion strategies. During the testing phase, the global and reconstructed local features are concatenated together for re-id without a complex visible region matching algorithm. Extensive experiments on occluded, partial, and holistic re-id benchmarks demonstrate the superiority of PRE over state-of-the-art methods in terms of accuracy and model complexity. Gang Yan 0001, Zijin Wang, Shuze Geng, Yang Yu 0022, Yingchun Guo |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | JAC-Net: Joint learning with adaptive exploration and concise attention for unsupervised domain adaptive person re-identification
Yingchun Guo, Xiaoke Hao, Xi Chen 0044 |
Neurocomputing | 1 |
| 2022 | Mining semantic information from intra-image and cross-image for few-shot segmentation
Yingchun Guo, Ming Yu 0006 |
Multim. Tools Appl. | 2 |
| 2021 | SEINet: Semantic-Edge Interaction Network for Image Manipulation Localization
Na Qi, Yingchun Guo, Bin Li 0011 |
PRCV (2) | 3 |
| 2021 | Hypergraph Neural Network for Skeleton-Based Action RecognitionabstractRecently, skeleton-based human action recognition has attracted a lot of research attention in the field of computer vision. Graph convolutional networks (GCNs), which model the human body skeletons as spatial-temporal graphs, have shown excellent results. However, the existing methods only focus on the local physical connection between the joints, and ignore the non-physical dependencies among joints. To address this issue, we propose a hypergraph neural network (Hyper-GNN) to capture both spatial-temporal information and high-order dependencies for skeleton-based action recognition. In particular, to overcome the influence of noise caused by unrelated joints, we design the Hyper-GNN to extract the local and global structure information via the hyperedge (i.e., non-physical connection) constructions. In addition, the hypergraph attention mechanism and improved residual module are induced to further obtain the discriminative feature representations. Finally, a three-stream Hyper-GNN fusion architecture is adopted in the whole framework for action recognition. The experimental results performed on two benchmark datasets demonstrate that our proposed method can achieve the best performance when compared with the state-of-the-art skeleton-based methods. Xiaoke Hao, Yingchun Guo, Ming Yu 0006 |
IEEE Trans. Image Process. | 3 |
| 2020 | Multi-modal neuroimaging feature selection with consistent metric constraint for diagnosis of Alzheimer's disease
Xiaoke Hao, Yongjin Bao, Yingchun Guo, Ming Yu 0006, Daoqiang Zhang, Shannon L. Risacher, Andrew J. Saykin, Xiaohui Yao, Li Shen 0001 |
Medical Image Anal. | 3 |
| 2020 | AR-Net: Adaptive Attention and Residual Refinement Network for Copy-Move Forgery DetectionabstractIn copy-move forgery, the illumination and contrast of tampered and genuine regions are highly consistent, which poses a greater challenge in copy-move forgery detection. In this article, an end-to-end neural network is proposed based on adaptive attention and residual refinement network (AR-Net). Specifically, position and channel attention features are fused by the adaptive attention mechanism to fully capture context information and enrich the representation of features. Second, deep matching is adopted to compute the self-correlation between feature maps, and atrous spatial pyramid pooling fuses the scaled correlation maps to generate the coarse mask. Finally, the coarse mask is optimized through the residual refinement module, which retains the structure of object boundaries. Extensive experiments, evaluated on CASIAII, COVERAGE, and CoMoFoD datasets, demonstrate that the AR-Net has superior performance than state-of-the-art algorithms and can locate tampered and corresponding genuine regions at the pixel level. In addition, AR-Net has high robustness on postprocessing operations, such as noise, blur, and JPEG recompression. Gang Yan 0001, Yingchun Guo, Yongfeng Dong |
IEEE Trans. Ind. Informatics | 4 |
| 2018 | Image retargeting quality assessment based on content deformation measurement
Yingchun Guo, Yuting Hao, Ming Yu 0006 |
Signal Process. Image Commun. | 1 |
| 2009 | Short-Term Load Forecasting Using Support Vector Regression Based on Pattern-BaseabstractA new idea is proposed that preprocessing is the key to improving the precision of short-term load forecasting (STLF). This paper presents a new model of STLF which is using support vector regression (SVR) based on pattern-base. Our model can be described as follows: firstly, it recognizes the different patterns of daily load according such features as weather and date type by means of data mining technology of classification and regression tree (CART); secondly, it sets up pattern-bases which are composed of daily load data sequence with highly similar features; thirdly, it establishes SVR forecasting model based on the pattern-base which matches to the forecasting day. Since the patterns of daily load are treated beforehand, the rule of the historical data sequence is more obvious. The model has many advantages: first, since the training data has similar pattern to the forecasting day, the model reflects the rule of daily load accurately and improves forecasting precision accordingly; second, as the pattern variables need not to be input into model, the mapping of the categorical variables is solved; third, as inputs are reduced, the model is simplified and the runtime is lessened. The simulation indicates that the new method is feasible and the forecasting precision is greatly improved. Yingchun Guo, Dongxiao Niu |
ACIIDS | 1 |