Ge Jiao

dblp:236/9468 · DBLP profile ↗
← Back
19ranked-venue papers
1as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HairEdit: Hierarchical latent fusion network for Text- and Image-Guided Hair Editing
Xiaodong Qian, Ge Jiao, Chen Li 0048
Eng. Appl. Artif. Intell.3
2026 A Tubercle Bacilli detection network combining frequency guidance and adaptive multi path fusion
Ge Jiao, Shao Jian Wu, Jia Yu Huang
Eng. Appl. Artif. Intell.2
2026 Degradation-aware pyramid heterogeneous mixture-of-experts network for underwater image restoration
Ge Jiao
Neurocomputing3
2026 DEER: Diffusion-empowered efficient restoration for underwater images
Wanhui Gao, Ge Jiao
Pattern Recognit.5
2026 SAM2-DEGNet: dual-stage edge guidance network for camouflaged object detection using SAM2
Zhenpeng Zhong, Ge Jiao, Guangcheng Li, Tanghui Liu
Vis. Comput.2
2025 Semi-supervised Iterative Learning Network for Camouflaged Object Detection
abstract
Current camouflaged object detection (COD) methods rely heavily on large-scale datasets with pixel-level annotations. We propose a semi-supervised iterative learning network (SILNet) to address the reliance on large-scale pixel-level annotations in COD. SILNet employs a co-training strategy with convolutional networks and Transformers as encoders, followed by a binary gated decoder (BGD) for feature fusion. To optimize the use of labeled data, we introduce an optimal representative election mechanism (OREM) to identify key sequences of unlabeled images, guiding iterative learning and pseudo-label generation. To reduce noise in pseudo-labels, we incorporate a long-range representation module (LRM) leveraging Mamba’s background modeling. Experiments show that SILNet trained with only 10% of the labeled data outperforms state-of-theart unsupervised and weakly supervised methods, achieving performance competitive with fully supervised models.
Guowen Yue, Ge Jiao, Jiahao Xiang
ICASSP2
2025 MSFENet: Multi-Scale Filter-Enhanced Network architecture for digital image forgery trace localization
Min Mao, Ge Jiao, Wanhui Gao, Jixun Ye
Comput. Vis. Image Underst.2
2025 More observation leads to more clarity: Multi-view collaboration network for camouflaged object detection
Fangyan Wang, Ge Jiao, Guowen Yue
Neurocomputing2
2025 SPNet: Seam carving detection via spatial-phase learning
Jiyou Chen, Zhi Lv, Ge Jiao, Gaobo Yang
J. Inf. Secur. Appl.3
2025 InDReCT: Intra-domain dual reconstruction for cross-domain transfer in camouflaged object detection
Guowen Yue, Ge Jiao, Fangyan Wang
Knowl. Based Syst.2
2025 Compact structural feature enhancement for unsupervised anomaly detection in chest radiographs
Jixun Ye, Wanhui Gao, Ge Jiao
Vis. Comput.4
2025 When CNN meet with ViT: decision-level feature fusion for camouflaged object detection
Guowen Yue, Ge Jiao, Chen Li 0048, Jiahao Xiang
Vis. Comput.2
2024 Dual-Color Granularity Alignment for Text-Based Person Search
abstract
Text-based Person Search (TBPS) aims to retrieve the person images based on the given text descriptions. Due to the heterogeneity between modalities and the fine granularity of the person, it is challenging to address the task. Existing methods often overlook granularity consistency across different color channels, which means there’s much potential to enhance retrieval performance. In this paper, we propose a Dual-Color Granularity Alignment (DCGA) method for Text-Based Person Search. DCGA harnesses both color and grayscale information to address issues of color reliance and granularity consistency. Moreover, by employing an improved CR Loss with grayscale information used as an additional weak supervision, DCGA addresses intra-class variance and dataset scarcity. Extensive experiments have demonstrated that our proposed DCGA method achieves state-of-the-art results on all three public datasets.
Yuxing Lu, Ge Jiao
ICASSP3
2024 Concentrated Reasoning and Unified Reconstruction for Multi-Modal Media Manipulation
abstract
Detecting and Grounding Multi-Modal Media Manipulation (DGM4) is an emerging task that aims to identify and locate manipulated elements in both textual and visual media. Given the complexity of this task, the model requires more sophisticated reasoning capabilities to align multi-modal features and capture forgery traces. To this end, we propose a Concentrated reasoning and Unified reconstruction framework (CrUr) for DGM4. Instead of adhering to traditional hierarchical reasoning paradigms, we directly carry out all inference tasks using integrated multi-modal features. Specifically, we extract and align features at a finer granularity, capturing subtle differences that may indicate manipulation by leveraging advanced mask signal modeling. Moreover, to adapt to fine-grained reasoning tasks, we design a transformer-based Reconstruction Harmonizer to facilitate more complex interactions among the reconstructed features, ultimately obtaining integrated features. Experimental results on the DGM4datasets show that our method achieves state-of-the-art performances.
Yuxing Lu, Ge Jiao
ICASSP3
2024 Multi-scale pooling learning for camouflaged instance segmentation
Chen Li 0048, Ge Jiao, Guowen Yue
Appl. Intell.2
2024 ETBHD-HMF: A Hierarchical Multimodal Fusion Architecture for Enhanced Text-Based Hair Design
abstract
Abstract Text‐based hair design (TBHD) represents an innovative approach that utilizes text instructions for crafting hairstyle and colour, renowned for its flexibility and scalability. However, enhancing TBHD algorithms to improve generation quality and editing accuracy remains a current research difficulty. One important reason is that existing models fall short in alignment and fusion designs. Therefore, we propose a new layered multimodal fusion network called ETBHD‐HMF, which decouples the input image and hair text information into layered hair colour and hairstyle representations. Within this network, the channel enhancement separation (CES) module is proposed to enhance important signals and suppress noise for text representation obtained from CLIP, thus improving generation quality. Based on this, we develop the weighted mapping fusion (WMF) sub‐networks for hair colour and hairstyle. This sub‐network applies the mapper operations to input image and text representations, acquiring joint information. The WMF then selectively merges image representation and joint information from various style layers using weighted operations, ultimately achieving fine‐grained hairstyle designs. Additionally, to enhance editing accuracy and quality, we design a modality alignment loss to refine and optimize the information transmission and integration of the network. The experimental results of applying the network to the CelebA‐HQ dataset demonstrate that our proposed model exhibits superior overall performance in terms of generation quality, visual realism, and editing accuracy. ETBHD‐HMF (27.8 PSNR, 0.864 IDS) outperformed HairCLIP (26.9 PSNR, 0.828 IDS), with a 3% higher PSNR and a 4% higher IDS.
Ge Jiao, Chen Li 0048
Comput. Graph. Forum2
2024 Camouflaged Instance Segmentation From Global Capture to Local Refinement
abstract
Camouflaged instance segmentation (CIS) aims to segment instances that are seamlessly embedded in their surroundings. Existing CIS methods often focus on utilizing global information but neglect local information, resulting in incomplete feature representation and reduced accuracy. To address this, we propose a global-to-local network (GLNet) for CIS, leveraging both global and local information for enhanced feature representation and segmentation. Specifically, GLNet consists of two main components: global capture and local refinement. In global capture, we introduce a novel dual-branch convolutional feedforward network (Dual-FFN), which aims to more effectively capture camouflaged instances in complex scenes. In local refinement, we design a U-shape feature fusion module (UFFM) and an edge-guide fusion module (EFM). These modules facilitate the fusion of multi-scale features by cascading. As a result, the network gains an enhanced ability to discern the intricate details of camouflaged instances. Experimental results demonstrate that our GLNet outperforms existing methods, with a 49.3% average precision (AP) on the COD10K-Test.
Chen Li 0048, Ge Jiao
IEEE Signal Process. Lett.2
2024 Declined Tactile Angle Discrimination in Young Patients With Migraine Without Aura or Tension-Type Headache
abstract
Many headache patients often report cognitive disturbances, but tactile cognitive data are limited. Applying computing-aided strategies to reveal the association between migraine without aura (MOA) or tension-type headache (TTH) and tactile cognition is one of the research highlights. The aim of this study was to investigate whether MOA or TTH patients had a decline in tactile discrimination by utilizing a tactile angle discrimination tester. A cross-sectional study was performed between 1 January 2021, and 1 January 2022. A total of 301 participants were enrolled, with 107 in control, 90 in MOA, and 104 in TTH groups. A tactile cognition tester was used to objectively examine tactile discrimination in all participants. Tactile angle discrimination thresholds were measured to compare tactile cognitive functions among three groups. There were no statistically significant differences in their demographic characteristics. Compared to the normal control group, the MOA and TTH groups exhibited significantly higher tactile angle discrimination thresholds (showing decline in tactile discrimination), whereas no significant differences were found between the MOA and TTH groups. Differences in tactile angle discrimination thresholds were observed between young (≤ 44 years old) and middle-aged/elderly (≥ 45 years old) participants in the normal control group but not in the MOA and TTH groups. Moreover, the tactile deficits shown in the MOA or TTH groups were evident only in young participants. This study first demonstrated that patients with MOA or TTH, especially those patients younger than 44 years old, had decreased tactile angle discrimination ability, suggesting decline of tactile cognition.
Ge Jiao, Jian Zhang 0119, Junru Zhu, Qunxi Dong, Aihua Wang, Shengyuan Yu
IEEE Trans. Comput. Soc. Syst.1
2022 Feature Passing Learning for Image Steganalysis
abstract
Image steganalysis aims to detect whether secret information is hidden in an image. This technique has critical applications in the field of information security. Most existing methods combine popular computer vision components for design without profoundly exploring the key factors applicable to image steganalysis. This letter reveals the limitations of existing feature passing and downsampling methods for image steganalysis tasks. We found that existing methods that pass shallow features through residual connections cannot cope with the problem of feature disappearance during network forward. In addition, the information-reducing downsampling methods used by these methods suppress the expression of steganographic features. To address these issues, we propose a feature enhancement passing module (FEPM) to help pass shallow features to deep layers and an attention downsampling module (ADM) to perform attention learning on full-resolution features. Combining these two structures, we designed an ultra-lightweight and highprecision image steganalysis network, FPNet, which contains only 0.16M parameters. The results of several experiments in the same environment show that our method outperforms current state-of-the-art methods in several aspects, including detection accuracy and computational effort. The code is available athttps://github.com/henryccl/FPNet.
Ge Jiao, Xiyu Sun
IEEE Signal Process. Lett.2