VLDB 2026 Research / reviewers in the wild / expert
Jin Wang 0039
dblp:92/1375-39
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0002-0533-4523ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic ConsistencyabstractCross-modal Knowledge Distillation has demonstrated promising performance on paired modalities with strong semantic connections, referred to as Symmetric Cross-modal Knowledge Distillation (SCKD). However, implementing SCKD becomes exceedingly constrained in real-world scenarios due to the limited availability of paired modalities. To this end, we investigate a general and effective knowledge learning concept under weak semantic consistency, dubbed Asymmetric Cross-modal Knowledge Distillation (ACKD), aiming to bridge modalities with limited semantic overlap. Nevertheless, the shift from strong to weak semantic consistency improves flexibility but exacerbates challenges in knowledge transmission costs, which we rigorously verified based on optimal transport theory. To mitigate the issue, we further propose a framework, namely SemBridge, integrating a Student-Friendly Matching module and a Semantic-aware Knowledge Alignment module. The former leverages self-supervised learning to acquire semantic-based knowledge and provide personalized instruction for each student sample by dynamically selecting the relevant teacher samples. The latter seeks the optimal transport path by employing Lagrangian optimization. To facilitate the research, we curate a benchmark dataset derived from two modalities, namely Multi-Spectral (MS) and asymmetric RGB images, tailored for remote sensing scene classification. Comprehensive experiments exhibit that our framework achieves state-of-the-art performance compared with 7 existing approaches on 6 different model architectures across various datasets. Riling Wei, Kelu Yao, Chuanguang Yang, Jin Wang 0039, Zhuoyan Gao, Chao Li 0028 |
AAAI | 4 |
| 2025 | Forensics-Bench: A Comprehensive Forgery Detection Benchmark Suite for Large Vision Language ModelsabstractRecently, the rapid development of AIGC has significantly boosted the diversities of fake media spread in the Internet, posing unprecedented threats to social security, politics, law, and etc. To detect the ever-increasingly diverse malicious fake media in the new era of AIGC, recent studies have proposed to exploit Large Vision Language Models (LVLMs) to design robust forgery detectors due to their impressive performance on a wide range of multimodal tasks. However, it still lacks a comprehensive benchmark designed to comprehensively assess LVLMs' discerning capabilities on forgery media. To fill this gap, we present Forensics-Bench, a new forgery detection evaluation benchmark suite to assess LVLMs across massive forgery detection tasks, requiring comprehensive recognition, location and reasoning capabilities on diverse forgeries. Forensics-Bench comprises 63, 292 meticulously curated multi-choice visual questions, covering 112 unique forgery detection types from 5 perspectives: forgery semantics, forgery modalities, forgery tasks, forgery types and forgery models. We conduct thorough evaluations on 22 open-sourced LVLMs and 3 proprietary models GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet, highlighting the significant challenges of comprehensive forgery detection posed by Forensics-Bench. We anticipate that Forensics-Bench will motivate the community to advance the frontier of LVLMs, striving for all-around forgery detectors in the era of AIGC. The deliverables will be updated here. Jin Wang 0039, Chenghui Lv, Shichao Dong 0001, Kelu Yao, Chao Li 0028, Wenqi Shao |
CVPR | 1 |
| 2024 | Sparse Beats Dense: Rethinking Supervision in Radar-Camera Depth Completion
Minhao Jing, Jin Wang 0039, Shichao Dong 0001, Jiajun Liang, Haoqiang Fan, Renhe Ji |
ECCV (49) | 3 |
| 2024 | Diagnosing the Compositional Knowledge of Vision Language Models from a Game-Theoretic ViewabstractCompositional reasoning capabilities are usually considered as fundamental skills to characterize human perception. Recent studies show that current Vision Language Models (VLMs) surprisingly lack sufficient knowledge with respect to such capabilities. To this end, we propose to thoroughly diagnose the composition representations encoded by VLMs, systematically revealing the potential cause for this weakness. Specifically, we propose evaluation methods from a novel game-theoretic view to assess the vulnerability of VLMs on different aspects of compositional understanding, e.g., relations and attributes. Extensive experimental results demonstrate and validate several insights to understand the incapabilities of VLMs on compositional reasoning, which provide useful and reliable guidance for future studies. The deliverables will be updated here. Jin Wang 0039, Shichao Dong 0001, Yapeng Zhu, Kelu Yao, Chao Li 0028 |
ICML | 1 |
| 2023 | Implicit Identity Leakage: The Stumbling Block to Improving Deepfake Detection GeneralizationabstractIn this paper, we analyse the generalization ability of binary classifiers for the task of deepfake detection. We find that the stumbling block to their generalization is caused by the unexpected learned identity representation on images. Termed as the Implicit Identity Leakage, this phenomenon has been qualitatively and quantitatively verified among various DNNs. Furthermore, based on such understanding, we propose a simple yet effective method named the ID-unaware Deepfake Detection Model to reduce the influence of this phenomenon. Extensive experimental results demonstrate that our method outperforms the state-of-the-art in both in-dataset and cross-dataset evaluation. The code is available at https://github.com/megvii-research/CADDM. Shichao Dong 0001, Jin Wang 0039, Renhe Ji, Jiajun Liang, Haoqiang Fan, Zheng Ge |
CVPR | 2 |
| 2023 | Towards Understanding the Generalization of Deepfake Detectors from a Game-Theoretical ViewabstractThis paper aims to explain the generalization of deep-fake detectors from the novel perspective of multi-order interactions among visual concepts. Specifically, we propose three hypotheses: 1. Deepfake detectors encode multi-order interactions among visual concepts, in which the low-order interactions usually have substantially negative contributions to deepfake detection. 2. Deepfake detectors with better generalization abilities tend to encode low-order interactions with fewer negative contributions. 3. Generalized deepfake detectors usually weaken the negative contributions of low-order interactions by suppressing their strength. Accordingly, we design several mathematical metrics to evaluate the effect of low-order interaction for deepfake detectors. Extensive comparative experiments are conducted, which verify the soundness of our hypotheses. Based on the analyses, we further propose a generic method, which directly reduces the toxic effects of low-order interactions to improve the generalization of deepfake detectors to some extent. Kelu Yao, Jin Wang 0039, Boyu Diao, Chao Li 0028 |
ICCV | 2 |
| 2022 | Interpretable Generative Adversarial NetworksabstractLearning a disentangled representation is still a challenge in the field of the interpretability of generative adversarial networks (GANs). This paper proposes a generic method to modify a traditional GAN into an interpretable GAN, which ensures that filters in an intermediate layer of the generator encode disentangled localized visual concepts. Each filter in the layer is supposed to consistently generate image regions corresponding to the same visual concept when generating different images. The interpretable GAN learns to automatically discover meaningful visual concepts without any annotations of visual concepts. The interpretable GAN enables people to modify a specific visual concept on generated images by manipulating feature maps of the corresponding filters in the layer. Our method can be broadly applied to different types of GANs. Experiments have demonstrated the effectiveness of our method. Chao Li 0028, Kelu Yao, Jin Wang 0039, Boyu Diao, Yongjun Xu 0001, Quanshi Zhang |
AAAI | 3 |
| 2022 | Explaining Deepfake Detection by Analysing Image Matching
Shichao Dong 0001, Jin Wang 0039, Jiajun Liang, Haoqiang Fan, Renhe Ji |
ECCV (14) | 2 |