VLDB 2026 Research / reviewers in the wild / expert
Jiaxin Ai
dblp:324/2291
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-4303-7749ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language ModelsabstractMultimodal large language models (MLLMs), which integrate language and visual cues for problem-solving, are crucial for advancing artificial general intelligence (AGI). However, current benchmarks for measuring the intelligence of MLLMs suffer from limited scale, narrow coverage, and unstructured knowledge, offering only static and undifferentiated evaluations. To bridge this gap, we introduce MDK12-Bench, a large-scale multidisciplinary benchmark built from real-world K–12 exams spanning six disciplines with 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. Covering five question formats with difficulty and year annotations, it enables comprehensive evaluation to capture the extent to which MLLMs perform over four dimensions: 1) difficulty levels, 2) temporal (cross-year) shifts, 3) contextual shifts, and 4) knowledge-driven reasoning. We propose a novel dynamic evaluation framework that introduces unfamiliar visual, textual, and question form shifts to challenge model generalization while improving benchmark objectivity and longevity by mitigating data contamination. We further evaluate knowledge-point reference-augmented generation (KP-RAG) to examine the role of knowledge in reasoning. Key findings reveal limitations in current MLLMs in multiple aspects and provide guidance for enhancing model reasoning, robustness, and AI-assisted education. Xiaopeng Peng 0001, Fanrui Zhang, Zhaopan Xu, Jiaxin Ai, Yansheng Qiu, Wangbo Zhao, Jiajun Song, Chuanhao Li 0001, Weidong Tang, Zhen Li 0026, Haoquan Zhang, Zizhen Li, Xiaofeng Mao, Yukang Feng, Kai Wang 0036, Xiaojun Chang, Wenqi Shao, Yang You 0001, Kaipeng Zhang |
AAAI | 5 |
| 2026 | MeepleLM: A Virtual Playtester Simulating Diverse Subjective ExperiencesabstractZizhen Li, Chuanhao Li, Yibin Wang, Jianwen Sun, Yukang Feng, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Yifei Huang, Kaipeng Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zizhen Li, Chuanhao Li 0001, Yukang Feng, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Kaipeng Zhang |
ACL (1) | 6 |
| 2026 | Navigating Truth in Multimodal Fact-checking via Retrieval- and Reasoning-Enhanced Large Language ModelsabstractRecent studies show that claims incorporating both text and images spread more effectively than those with text alone, presenting significant challenges for multimodal fact-checking. The rapid development of Multi-modal Large Language Models (MLLMs) has greatly advanced research in this field, enabling stronger performance. However, existing MLLM-based fact-checking methods fail to fully exploit visual evidence, and their reliance on rigid fine-tuning templates limits context-aware explanations and leads to weak deep reasoning. To address these limitations, we propose FACTCOMPASS, a novel framework that combines reasoning-aware fine-tuning with large-scale rule-based reinforcement learning and incorporates a semantic- and knowledge-enhanced retrieval module to strengthen deep reasoning and improve evidence utilization. This framework enhances evidence retrieval by obtaining semantically relevant evidence images, enriching the contextual understanding of claim-related images, and refining textual evidence at the knowledge level. To further enhance reasoning, we introduce a self-refining reinforcement fine-tuning strategy: (1) distilling GPT-4o's reasoning from partially fact-checking data for cold-start Chain-of-Thought learning; (2) activating reasoning across broader datasets using prior knowledge and rejection sampling; (3) applying Group Relative Policy Optimization to explore diverse reasoning paths and optimize factual consistency. Extensive experiments have demonstrated the effectiveness of the proposed framework. Fanrui Zhang, Qiang Zhang 0051, Chuanhao Li 0001, Jiaxin Ai, Yukang Feng, Zizhen Li, Kaipeng Zhang, Jiawei Liu 0001, Zhengjun Zha |
WWW | 5 |
| 2026 | IDRetracor: Towards Visual Forensics against Malicious Face SwappingabstractThe deepfake-based face swapping technique poses significant risks to personal identity security. Although many detection methods have been proposed to counter malicious face swapping, they typically provide only binary labels (Fake/Real), lacking reliable and interpretable evidence. To address this limitation, we introduce a novel task called face retracing, which aims to visually trace back the original target face from a given fake one through inverse mapping. This task is based on the observation that current face swapping methods are neither flawless nor entirely random, leaving recoverable traces of the original identity. To this end, we propose IDRetracor, a model designed to recover arbitrary original target identities from fake faces generated by various face swapping techniques. Specifically, we first employ a mapping resolver to estimate the possible solution space of the original face for inverse mapping. Then, we introduce Mapping-Aware Convolutions (MACs), which consist of multiple dynamically combined kernels guided by the mapping resolver to adaptively handle diverse face swapping patterns. Extensive experiments demonstrate that IDRetracor achieves strong performance in retracing original faces, validated by both quantitative metrics and qualitative assessments. Jikang Cheng, Jiaxin Ai, Zhen Han 0002, Chao Liang 0001, Qin Zou 0001, Zhongyuan Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Stacking Brick by Brick: Aligned Feature Isolation for Incremental Face Forgery DetectionabstractThe rapid advancement of face forgery techniques has introduced a growing variety of forgeries. Incremental Face Forgery Detection (IFFD), involving gradually adding new forgery data to fine-tune the previously trained model, has been introduced as a promising strategy to deal with evolving forgery methods. However, a naively trained IFFD model is prone to catastrophic forgetting when new forgeries are integrated, as treating all forgeries as a single “Fake” class in the Real/Fake classification can cause different forgery types overriding one another, thereby resulting in the forgetting of unique characteristics from earlier tasks and limiting the model’s effectiveness in learning forgery specificity and generality. In this paper, we propose to stack the latent feature distributions of previous and new tasks brick by brick, i.e., achieving aligned feature isolation. In this manner, we aim to preserve learned forgery information and accumulate new knowledge by minimizing distribution overriding, thereby mitigating catastrophic forgetting. To achieve this, we first introduce Sparse Uniform Replay (SUR) to obtain the representative subsets that could be treated as the uniformly sparse versions of the previous global distributions. We then propose a Latent-space Incremental Detector (LID) that leverages SUR data to isolate and align distributions. For evaluation, we construct a more advanced and comprehensive benchmark tailored for IFFD. The leading experimental results validate the superiority of our method. Code is available at https://github.com/beautyremain/SUR-LID . Jikang Cheng, Zhiyuan Yan 0002, Ying Zhang 0021, Jiaxin Ai, Qin Zou 0001, Chen Li 0031, Zhongyuan Wang 0001 |
CVPR | 5 |
| 2025 | InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning StylesabstractZizhen Li, Chuanhao Li, Yibin Wang, Qi Chen, Diping Song, Yukang Feng, Jianwen Sun, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Kaipeng Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Zizhen Li, Chuanhao Li 0001, Diping Song, Yukang Feng, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Kaipeng Zhang |
EMNLP | 8 |
| 2025 | ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for Mllm-Based Process Judges
Jiaxin Ai, Zhaopan Xu, Fanrui Zhang, Zizhen Li, Yukang Feng, Baojin Huang, Zhongyuan Wang 0001, Kaipeng Zhang |
ICCV | 1 |
| 2025 | Sekai: A Video Dataset towards World ExplorationabstractVideo generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration.However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world.In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories.Comprehensive analyses and experiments demonstrate the dataset’s scale, diversity, annotation quality, and effectiveness for training video generation models.We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications. Zhen Li 0026, Chuanhao Li 0001, Xiaofeng Mao, Shaoheng Lin, Ming Li 0010, Shitian Zhao, Zhaopan Xu, Xinyue Li 0001, Yukang Feng, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Yuwei Wu 0001, Tong He 0001, Yunde Jia, Kaipeng Zhang |
NeurIPS | 13 |
| 2025 | Luminance decomposition and reconstruction for high dynamic range Video Quality Assessment
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Jing Xiao 0004, Zixiang Xiong |
Pattern Recognit. | 4 |
| 2024 | STA: Enhancing Spatio-temporal Crowd Flow Prediction Using Attention-based Deep Learning and Feature Similarity
Xiujuan Xu, RenJie Liu, Jiaxin Ai, Yu Liu 0035, Xiaowei Zhao 0003 |
ADMA (3) | 3 |
| 2024 | Joint Distortion Restoration and Quality Feature Learning for No-reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) methods, inspired by the free energy principle, improve the accuracy of image quality prediction by simulating the human brain’s repair process for distorted images. However, existing methods use separate optimization schemes for distortion restoration and quality prediction, which undermines the accurate mapping of feature representations to quality scores. To address this issue, we propose a joint restoration and quality feature learning NR-IQA (RQFL-IQA) method to jointly tackle distortion image restoration and quality prediction within a unified framework. To accurately establish the quality reconstruction relationship between distorted and restored images, a hybrid loss function based on pixel-wise and structure-wise representations is used to improve the restoration capability of the image restoration network. The proposed RQFL-IQA exploits rich labels, including restored images and quality scores, to enable the model to learn more discriminative features and establish a more accurate mapping from feature representation to quality scores. In addition, to avoid the impact of poor restoration on quality prediction, we propose a module with a cleaning function to reweight the fusion of restored and primitive features to achieve more perceptual consistency in feature fusion. Experimental results on public IQA datasets show that the proposed RQFL-IQA is superior over existing methods. Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Zixiang Xiong |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Implicit Identity Driven Deepfake Face Swapping DetectionabstractIn this paper, we consider the face swapping detection from the perspective of face identity. Face swapping aims to replace the target face with the source face and generate the fake face that the human cannot distinguish between real and fake. We argue that the fake face contains the explicit identity and implicit identity, which respectively corresponds to the identity of the source face and target face during face swapping. Note that the explicit identities of faces can be extracted by regular face recognizers. Particularly, the implicit identity of real face is consistent with the its explicit identity. Thus the difference between explicit and implicit identity of face facilitates face swapping detection. Following this idea, we propose a novel implicit identity driven framework for face swapping detection. Specifically, we design an explicit identity contrast (EIC) loss and an implicit identity exploration (IIE) loss, which supervises a CNN backbone to embed face images into the implicit identity space. Under the guidance of EIC, real samples are pulled closer to their explicit identities, while fake samples are pushed away from their explicit identities. More-over, IIE is derived from the margin-based classification loss function, which encourages the fake faces with known target identities to enjoy intra-class compactness and inter-class diversity. Extensive experiments and visualizations on several datasets demonstrate the generalization of our method against the state-of-the-art counterparts. Baojin Huang, Zhongyuan Wang 0001, Jifan Yang, Jiaxin Ai, Qin Zou 0001, Qian Wang 0002, Dengpan Ye |
CVPR | 4 |
| 2023 | Deepfake Face Provenance for Proactive ForensicsabstractMalicious deepfake face not only violates the privacy of personal identities, but also confuses the public and causes huge social harm. The current deepfake detection only stays at the level of distinguishing between true and false, but cannot trace the original genuine face corresponding to the fake face, that is, it does not have the ability to trace the source of evidence. The deepfake countermeasure technology for judicial forensics urgently calls for deepfake inversion. This paper pioneers an interesting question about face deepfake, active forensics that "know what it is and how it happened". Given that deepfake faces do not completely discard the features of original faces, especially facial expressions and poses, we argue that original faces can be approximately speculated from their deepfake counterparts. Correspondingly, we design a disentangling reversing network that decouples latent space features of deepfake faces under the supervision of real-fake face pair samples to infer original faces in reverse. Jiaxin Ai, Zhongyuan Wang 0001, Baojin Huang, Zhen Han 0002, Qin Zou 0001 |
ICIP | 1 |
| 2023 | DeepReversion: Reversely Inferring the Original Face from the DeepFake FaceabstractDeepfake techniques can generate realistic fake images and videos. Malicious fake facial images quickly spread through the Internet, posing a potential threat to personal privacy and judicial forensics. However, the defense methods against deepfake proposed so far mainly focus on the discrimination of authenticity, but cannot identify the true source of the forged face, i.e., the original genuine face corresponding to the face-swapped fake face. This paper poses an interesting issue for face deepfake, which is the proactive forensics of “knowing what and knowing how”. In view of the fact that the fake face exhibits high similarity with the original face, especially the facial expression and pose, we argue that the original face can be approximately estimated from the deepfake counterpart. Accordingly, we advocate a deep-learning-based face inversion approach, so-called DeepReversion, which learns the inverse mapping from the deepfake face to the original face. Based on UNet, we design a specific end-to-end DeepReversion network, and conduct comprehensive experiments on public deepfake datasets. The experimental results show that the speculated face is highly consistent with the original face in terms of visual effects, PSNR, SSIM and similarity given by face recognizers. Jiaxin Ai, Zhongyuan Wang 0001, Baojin Huang, Zhen Han 0002 |
IJCNN | 1 |