Mingyuan Ge

dblp:313/0443 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2026
0000-0002-4094-0078ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2026 DynBrush: Structure-Aware and Style-Adaptive Transfer for Traditional Chinese Landscape Rendering
abstract
Artistic style transfer synthesizes a stylized image that preserves content geometry while substituting texture, color, and stroke statistics from a target artwork. Most neural style transfer models are designed and evaluated under the visual priors of Western paintings, where sharp edges, high-contrast textures, and explicit contours broadly define the style distribution. Directly applying them to traditional Chinese landscape painting often exposes a core conflict: maintaining globally coherent scene structure (e.g., mountain ridges, river trajectories, spatial depth) while producing region-aware abstraction with meaningful brushwork. This gap commonly leads to incomplete stylization, inconsistent perspective, or structural distortion, particularly in areas that require long-range continuity. To address this challenge, we present DynBrush, a quality-efficient style-transfer framework specialized for Chinese landscape rendering. DynBrush adopts sequential spatial modeling with a decoupled dual-channel encoder that disentangles structural and stylistic representations. We introduce Structural State Evolution (SSE) to reinforce global layout and geometric stability progressively. Stylistic State Refinement(SSR) complements this process by enhancing chromatic embedding and ink-tone transitions to match ink-and-wash aesthetics better. For decoding, Structure-Guided State(SGS) anchors reconstruction to globally consistent structural cues, while Cross-domain Style Attention(CSA) improves regional style correspondence across appearance discrepancies. We further incorporate Blank-Preserving Multi-scale Context Enhancer(BMCE) to organize brushstroke hierarchies across scales, yielding coherent stroke granularity, stable ink-tone gradation, and reserved-white fidelity. Experiments demonstrate that DynBrush delivers stable structure preservation and style consistency across diverse landscapes, with a favorable quality–efficiency trade-off and strong adaptability to traditional Chinese landscape characteristics.
Hangshen Nong, Mingyuan Ge, Mingyong Li
ICMR2
2026 ITAdapter: Image-Tag adapter framework with retrieval knowledge enhancer for radiology report generation
Shuaipeng Ding, Jianan Shui, Mingyuan Ge, Mengnan Fan, Xin Li 0242, Mingyong Li
Expert Syst. Appl.3
2025 Text-Driven Fashion Image Editing with Compositional Concept Learning and Counterfactual Abduction
abstract
Fashion image editing is a valuable tool for designers to convey their creative ideas by visualizing design concepts. With the recent advances in text editing methods, significant progress has been made in fashion image editing. However, they face two key challenges: spurious correlations in training data often induce changes in other areas when editing an area representing the intended editing concept, and these models typically lack the ability to edit multiple concepts simultaneously. To address the above challenges, we propose a novel Text-driven Fashion Image ediTing framework called T-FIT to mitigate the impact of spurious correlation by integrating counterfactual reasoning with compositional concept learning to precisely ensure compositional multi-concept fashion image editing relying solely on text descriptions. Specifically, T-FIT includes three key components. (i) Counterfactual abduction module, which learns an exogenous variable of the source image by a denoising U-Net model. (ii) Concept learning module, which identifies concepts in fashion image editing—such as clothing types and colors and projects a target concept into the space spanned from a series of textual prompts. (iii) Concept composition module, which enables simultaneous adjustments of multiple concepts by aggregating each concept’s direction vector obtained from the concept learning module. Extensive experiments show that our method can achieve state-of-the-art performance on various fashion image editing tasks, including single-concept editing (e.g., sleeve length, clothing type) and multi-concept editing (e.g., color & sleeve length).
Shanshan Huang 0004, Haoxuan Li 0001, Chunyuan Zheng 0001, Mingyuan Ge, Lei Wang 0197, Li Liu 0001
CVPR4
2025 Spatially-Aware Entity Relation Exploration for Remote Sensing Image-Text Retrieval
abstract
In recent years, remarkable progress has been made in remote sensing image-text retrieval (RSITR), which has transitioned from relying on compact global features to more fine-grained local features representing salient objects in images. However, existing methods typically concentrate only on significant entity information in remote sensing images as local features, overlooking the correlations between entities, resulting in isolated entity information. Moreover, there is ample room for exploring text local features.To address these issues, this paper presents an Entity Spatial Relation enhancement Network (ESRN), leveraging global and local entity features in remote sensing images and texts. For local feature processing, Graph Convolutional Network (GCN) is employed to aggregate the correlation between entity features and entity information in remote sensing images, enhancing entity information learning. The self-attention mechanism is used to model the remote dependency of entity keywords and spatial orientation semantic relations in text, strengthening the text representation ability. A strategy of proportionally adding different levels of features is proposed to enhance the representation of salient features and reduce noise interference.The approach was evaluated on two renowned remote sensing datasets, RSICD and RSITMD, validating the model's capacity to perceive the semantics of remote sensing images and text entities. Performance comparison, ablation experiments, and visualization analysis convincingly demonstrate the state-of-the-art performance of the ESRN method in the RSITR task.
Jianan Shui, Shuaipeng Ding, Mingyuan Ge, Mingyong Li
ICMR3
2025 Enhanced-Similarity Attention Fusion for Unsupervised Cross-Modal Hashing Retrieval
abstract
Abstract Although the fact that current methods have some effects, unsupervised cross-modal hashing methods still face several common challenges. First of all, the text features that have been collected from text data are not comprehensive enough to provide sufficient guidance for building textual modal similarity matrices. Secondly, the fusion of similarity matrices from different modalities lacks adaptability, leading to a less accurate final similarity matrix. This work suggests Enhanced Similarity Attention Fusion Hashing (ESAFH) as a remedy for these problems. Firstly, we construct a text encoder to enrich text features, an adjacency matrix is built to represent the association relationship between pairs of samples. Additionally, it is thought that features can be extracted from the sample and its semantic neighbor samples to enhance text features. Furthermore, we enhance the original similarity matrix by incorporating related information. This step aims to improve the accuracy of similarity estimation by considering the enriched text features obtained in the previous step. Finally, we introduce an enhanced attention fusion mechanism. This mechanism adaptively fuses the similarity matrices from different modalities, creating a unified inter-modal similarity matrix. This fused matrix guides the learning of hash functions by preserving the most relevant information from each modality. Through comprehensive experiments on the three popular datasets, the suggested ESAFH method is thoroughly assessed. The findings show that on these datasets, ESAFH performs satisfactorily in cross-modal retrieval tasks. In conclusion, by boosting text features, improving the similarity matrix, and utilizing an attention fusion mechanism, ESAFH solves the shortcomings of current methods.
Mingyong Li, Mingyuan Ge
Data Sci. Eng.2
2025 Joint multi-grained similarity contrastive learning for video-text retrieval
Mingyong Li, Mingyuan Ge, Wenfeng Zhang
Neurocomputing2
2025 Ambiguity-Aware and High-order Relation learning for multi-grained image-text matching
Yihua Gao, Mingyuan Ge, Mingyong Li
Knowl. Based Syst.3
2024 JM-CLIP: A Joint Modal Similarity Contrastive Learning Model for Video-Text Retrieval
abstract
In recent years, the work on video-text retrieval has been well-developed due to the emergence of large-scale pre-training methods. However, these works focus solely on inter-modal interactions and contrasts, neglecting the contrasts of multigrained features within modalities, which makes the similarity measurement less accurate. Worse still, many of these works only contrast features of the same grain size, ignoring the important information implied by features of different grain sizes. For these reasons, we propose a joint modal similarity contrastive learning model, named JM-CLIP. Firstly, the method employs a multidimensional contrastive strategy of two modal features, including inter-modal and intra-modal multi-grained feature contrasts. Secondly, to fuse the various contrastive similarities well, we also design a joint modal attention module to fuse the various similarities into a final joint multigranularity similarity score. The significant performance improvement achieved on three popular video-text retrieval datasets demonstrates the effectiveness and superiority of our proposed method. The code is available at https://github.com/DannielGe/JM-CLIP.
Mingyuan Ge, Yewen Li, Honghao Wu, Mingyong Li
ICASSP1
2024 Pseudo-label Based Unsupervised Momentum Representation Learning for Multi-domain Image Retrieval
Mingyuan Ge, Jianan Shui, Mingyong Li
MMM (3)1
2023 Deep Enhanced-Similarity Attention Cross-modal Hashing Learning
abstract
Despite the great success of existing cross-modal retrieval methods, existing unsupervised cross-modal hashing methods still suffer from common problems. First, the features extracted from the text are too sparse. Second, the similarity matrices of each different modality cannot be fused adaptively. In this paper, we propose Deep Enhanced-Similarity Attention Hashing (DESAH) to alleviate the above problems. Firstly, we construct a text encoder expanding graph convolutional neural network to simultaneously extract features of samples and their semantic neighbors to enrich text features. Secondly, we propose an enhanced attention fusion mechanism. The mechanism is used to adaptively fuse the similarity matrices within different modalities to form a unified inter-modal similarity matrix to guide the learning of hash functions. Extensive experiments have demonstrated that DESAH provides significant improvements in cross-modal retrieval tasks compared to baseline methods.
Mingyuan Ge, Yewen Li, Mingyong Li
ICMR1
2023 CCAH: A CLIP-Based Cycle Alignment Hashing Method for Unsupervised Vision-Text Retrieval
abstract
Due to the advantages of low storage cost and fast retrieval efficiency, deep hashing methods are widely used in cross‐modal retrieval. Images are usually accompanied by corresponding text descriptions rather than labels. Therefore, unsupervised methods have been widely concerned. However, due to the modal divide and semantic differences, existing unsupervised methods cannot adequately bridge the modal differences, leading to suboptimal retrieval results. In this paper, we propose CLIP‐based cycle alignment hashing for unsupervised vision‐text retrieval (CCAH), which aims to exploit the semantic link between the original features of modalities and the reconstructed features. Firstly, we design a modal cyclic interaction method that aligns semantically within intramodality, where one modal feature reconstructs another modal feature, thus taking full account of the semantic similarity between intramodal and intermodal relationships. Secondly, introducing GAT into cross‐modal retrieval tasks. We consider the influence of text neighbour nodes and add attention mechanisms to capture the global features of text modalities. Thirdly, Fine‐grained extraction of image features using the CLIP visual coder. Finally, hash encoding is learned through hash functions. The experiments demonstrate on three widely used datasets that our proposed CCAH achieves satisfactory results in total retrieval accuracy. Our code can be found at: https://github.com/CQYIO/CCAH.git .
Mingyong Li, Yewen Li, Mingyuan Ge
Int. J. Intell. Syst.4