VLDB 2026 Research / reviewers in the wild / expert
Ming Jin 0007
dblp:34/3870-7
· DBLP profile ↗
8ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-5097-8053ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Entropy-aware Mutual Student Co-training for Source-free Domain Adaptation in Medical Image SegmentationabstractUnsupervised Domain Adaptation (UDA) has shown remarkable success in medical image segmentation but its practical deployment is often constrained by strict privacy regulations that prohibit access to source-domain data. Source-Free Domain Adaptation (SFDA) addresses this limitation by enabling knowledge transfer from pre-trained source models without requiring private source data. Existing SFDA methods for medical image segmentation mainly rely on self-training with pseudo-labels generated by the source model. However, severe domain discrepancies introduce substantial pseudo-label noise, while non-uniform domain shifts across target samples lead to heterogeneous data distributions, rendering uniform adaptation strategies suboptimal. To address these challenges, we propose Entropy-aware Mutual Student Co-training (EMSC), a novel SFDA framework comprising one teacher model and two student models. An Entropy-Guided Difficulty Assessment (EGDA) module partitions target samples into source-similar and source-dissimilar subsets based on predictive uncertainty. Each subset is handled by an independent student branch equipped with a Subset-specific Adaptive Injection (SAI) module, which injects semantic anchors for source-similar samples and structural noise for source-dissimilar samples to enable differentiated adaptation and suppress pseudo-label noise. Extensive experiments on two widely used medical image segmentation benchmarks demonstrate that EMSC consistently outperforms state-of-the-art SFDA methods across multiple evaluation metrics, validating its robustness under heterogeneous domain shifts. Guangrun Chen, Litan Sun, Ming Jin 0007, Xiaofeng Qu, Sijie Niu |
ICMR | 4 |
| 2026 | Video and text semantic center alignment for text-video cross-modal retrieval
Ming Jin 0007, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031 |
Signal Process. Image Commun. | 1 |
| 2026 | MDA-MAA: A Collaborative Augmentation Approach for Generalizing Cross-Domain RetrievalabstractIn video-text cross-domain retrieval tasks, the generalization ability of the retrieval models is key to improving their performance and is crucial for enhancing their practical applicability. However, existing retrieval models exhibit significant deficiencies in cross-domain generalization. On one hand, models tend to overfit specific training domain data, resulting in poor cross-domain matching and significantly reduced retrieval accuracy when dealing with data from different, new, or mixed domains. On the other hand, although data augmentation is a vital strategy for enhancing model generalization, most existing methods focus on unimodal augmentation and fail to fully exploit the multimodal correlations between video and text. As a result, the augmented data lack semantic diversity, which further limits the model's ability to understand and perform in complex cross-domain scenarios. To address these challenges, this paper proposes an innovative collaborative augmentation approach named MDA-MAA, which includes two core modules: the Masked Attention Augmentation (MAA) module and the Multimodal Diffusion Augmentation (MDA) module. The MAA module applies masking to the original video frame features and uses an attention mechanism to predict the masked features, effectively reducing overfitting to training data and enhancing model generalization. The MDA module generates subtitles from video frames and uses the LLaMA model to infer comprehensive video captions. These captions, combined with the original video frames, are integrated into a diffusion model for joint learning, ultimately generating semantically enriched augmented video frames. This process leverages the multimodal relationship between video and text to increase the diversity of the training data distribution. Experimental results demonstrate that this collaborative augmentation method significantly improves the performance of video-text cross-domain retrieval models, validating its effectiveness in enhancing model generalization. Ming Jin 0007, Richang Hong |
IEEE Trans. Image Process. | 1 |
| 2025 | Revealing Security Flaws in Cross-Modal Retrieval Models Through Video PoisoningabstractVideo-text cross-modal retrieval is widely studied to improve retrieval accuracy. However, the security of video-text cross-modal retrieval models receives little attention. If attackers exploit the security vulnerabilities in these models, it poses a significant threat to the retrieval models. Thus, identifying security flaws in video-text cross-modal retrieval models becomes the focus of our research. We are the first to design a video poisoning model to uncover security vulnerabilities in retrieval models. Existing poisoning models have certain limitations when it comes to exploiting vulnerabilities in retrieval models. These include failing to comprehensively embed malicious information into the original video and being unable to maintain visual consistency between the original and poisoned videos. These limitations can result in unsuccessful attacks on retrieval models and an inability to effectively identify security flaws within them. To address these shortcomings, we design an efficient poisoning model that embeds malicious information thoroughly into the original clean data to attack video-text cross-modal retrieval models. We are the first to use a poisoning model to attack retrieval models, thereby uncovering their security vulnerabilities. Second, we introduce a bi-level poisoning module to ensure that malicious information is thoroughly embedded into the original video, thereby enhancing the attack capability of the poisoning model. Finally, we design an adversarial module to improve visual consistency between the original and poisoned videos, thus enhancing the concealment of malicious information within the training data of retrieval models. Our poisoning model can identify security flaws in video-text cross-modal retrieval models, providing insights into improving the security of retrieval models. The effectiveness of our model is validated on the MSR-VTT, LSMDC, and MSVD datasets. Ming Jin 0007, Wenbo Hu 0001, Richang Hong, Lei Zhu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | BiSeR-LMA: A Bidirectional Semantic Reasoning and Large Model Enhancement Approach for Text-Video Cross-Modal RetrievalabstractVideo, as an information carrier, provides a vast amount of important information to people. Therefore, the method of obtaining video becomes particularly important, which drives the research on text-video cross-modal retrieval technology. However, current text-video cross-modal retrieval models still face several issues. First, these models do not fully utilize the powerful reasoning and generative capabilities of large models to address the issues of missing critical objects and insufficient high-quality video-text paired training data. Second, existing retrieval models do not adequately research the bidirectional cross-modal semantic interaction and reasoning mechanism, which hinders the ability to fully capture and learn the implicit semantic features between different modalities. To address these issues, we propose an innovative bidirectional semantic reasoning and large model data augmentation cross-modal retrieval model (BiSeR-LMA). This model first leverages the strong reasoning and generative capabilities of large models to perform semantic reasoning on the textual descriptions of videos, then generates multiple semantically rich video frames, thereby compensating for the missing critical objects in the original video and improving the quality of video-text paired training data. Second, we design a bidirectional text-video semantic reasoning module, which uses features from one modality as auxiliary information to assist the model in reasoning the implicit semantic information of another modality. This enhances the model’s capability to establish semantic relationships and perform reasoning on implicit semantics, promoting text-video semantic alignment. Finally, we verify the effectiveness of the proposed cross-modal retrieval model on the MSR-VTT, LSMDC, and MSVD datasets. The source codes of our method are available at: https://anonymous.4open.science/r/BiSeR_LMA. Ming Jin 0007, Lei Zhu 0002, Richang Hong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Based on Spatial and Temporal Implicit Semantic Relational Inference for Cross-Modal RetrievalabstractTo meet users’ demands for video retrieval, text-video cross-modal retrieval technology continues to evolve. Methods based on pre-trained models and transfer learning are widely employed in designing cross-modal retrieval models, significantly enhancing the accuracy of video retrieval. However, these methods exhibit shortcomings when it comes to studying the relationships between video frames, preventing the model from fully establishing the hidden semantic relationships within video features. To further deduce the implicit semantic relationships among video frames, we propose a cross-modal retrieval model based on graph convolutional networks (GCN) and visual semantic inference (GVSI). The GCN is utilized to establish relationships between video frame features, facilitating the mining of hidden semantic information across video frames. In order to use text semantic features to help the model to infer temporal and implicit semantic information between video frames, we introduce a semantic mining and temporal space (SM&TS) inference module. Additionally, we design semantic alignment modules (SA_M) to align explicit and implicit object features present in both video and text. Finally, we analyze and validate the effectiveness of the model using MSR-VTT, MSVD, and LSMDC datasets. Ming Jin 0007, Wenbo Hu 0001, Lei Zhu 0002, Xiang Wang 0010, Richang Hong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Video Sampled Frame Category Aggregation and Consistent Representation for Cross-Modal RetrievalabstractMany current video and text cross-modal retrieval research works focus on narrowing the semantic gap between video and text, but ignore the semantic difference between different sampled frames in the same video and the correlation of feature distribution of objects contained in different sampled frames in the same video, as a result, the features of the sampled frames in the final learned video cannot well represent the semantic features of the whole video. To overcome the shortcomings of existing studies, we first use a pre-trained video frame classification-aggregation network to make the object categories contained in different sampled frames in the same video be more close to the important object categories contained in the whole video, so as to promote the feature distribution of different sampled frames in the same video to be consistent, and increase the relevance of object features in different frames. Then we propose a video internal frame aggregation loss module to solve the problem of inconsistent feature distribution between different frame features encoded by video encoder in the same video and the aggregation feature of the sampled frame, thus enhancing the ability of video sampled frame aggregation feature representation. Experiments conducted on three common datasets MSVD, MSR-VTT and DiDeMo demonstrate the validity of the proposed approach. Ming Jin 0007, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Coarse-to-fine dual-level attention for video-text cross modal retrieval
Ming Jin 0007, Huaxiang Zhang 0001, Lei Zhu 0002, Jiande Sun 0001, Li Liu 0031 |
Knowl. Based Syst. | 1 |