Kailin Jiang

dblp:309/9739 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-0742-4872ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 70% Trustworthy machine learning · 15% Vision and language · 15%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
knowledge editing
1.722025
In-Context Editing: Learning Knowledge from Self-Induced Distributions · ICLR 2025
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge · ICLR 2025
Machine learning › Trustworthy machine learning › multimodal trustworthiness
multimodal knowledge conflict
1.012026
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models · AAAI 2026
Natural language and speech › Language models and text generation › knowledge editing
in-context editing
0.912025
In-Context Editing: Learning Knowledge from Self-Induced Distributions · ICLR 2025
Natural language and speech › Language models and text generation
large language model fine-tuning
0.912025
In-Context Editing: Learning Knowledge from Self-Induced Distributions · ICLR 2025
Natural language and speech › Language models and text generation › knowledge editing
multimodal knowledge editing
0.912025
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge · ICLR 2025
Information retrieval › retrieval-augmented generation
multimodal retrieval-augmented generation
0.312026
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models · AAAI 2026
Information retrieval
retrieval-augmented generation
0.312026
Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models · AAAI 2026
Natural language and speech › Language models and text generation
large language model
0.312025
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge · ICLR 2025

Methods — techniques the papers use, named apart from their topics

conflict detection evaluation · 2.0benchmark construction · 2.0knowledge editing · 0.9in-context learning · 0.9gradient-based tuning · 0.9benchmark evaluation · 0.9
YearPublicationVenuePosition
2026 Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
abstract
Large Multimodal Models (LMMs) face notable challenges when encountering multimodal knowledge conflicts, particularly under retrieval-augmented generation (RAG) frameworks, where the contextual information from external sources may contradict the model’s internal parametric knowledge, leading to unreliable outputs. However, existing benchmarks fail to reflect such realistic conflict scenarios. Most focus solely on intra-memory conflicts, while context-memory and inter-context conflicts remain largely unaddressed. Furthermore, commonly used factual knowledge-based evaluations are often overlooked, and existing datasets lack a thorough investigation into conflict detection capabilities.To bridge this gap, we propose MMKC-Bench, a benchmark designed to evaluate factual knowledge conflicts in both context-memory and inter-context scenarios. MMKC-Bench encompasses four types of multimodal knowledge conflicts and includes 1,881 knowledge instances and 3,997 images across 32 broad types, collected through automated pipelines with human verification. We evaluate four representative series of LMMs on both model behavior analysis and conflict detection tasks. Our findings show that while current LMMs are capable of recognizing knowledge conflicts, they tend to favor internal parametric knowledge over external evidence. We hope MMKC-Bench will foster further research in multimodal knowledge conflict and enhance the development of multimodal RAG systems.
Yuntao Du 0001, Kailin Jiang, Yuyang Liang, Qihan Ren, Yi Xin 0003, Fenze Feng, Mingcai Chen, Hengyang Lu, Haozhe Wang 0002, Xiaoye Qu, Qian Li 0043, Dongrui Liu
AAAI3
2025 MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge
abstract
Knowledge editing techniques have emerged as essential tools for updating the factual knowledge of large language models (LLMs) and multimodal models (LMMs), allowing them to correct outdated or inaccurate information without retraining from scratch. However, existing benchmarks for multimodal knowledge editing primarily focus on entity-level knowledge represented as simple triplets, which fail to capture the complexity of real-world multimodal information. To address this issue, we introduce MMKE-Bench, a comprehensive **M**ulti**M**odal **K**nowledge **E**diting Benchmark, designed to evaluate the ability of LMMs to edit diverse visual knowledge in real-world scenarios. MMKE-Bench addresses these limitations by incorporating three types of editing tasks: visual entity editing, visual semantic editing, and user-specific editing. Besides, MMKE-Bench uses free-form natural language to represent and edit knowledge, offering a more flexible and effective format. The benchmark consists of 2,940 pieces of knowledge and 8,363 images across 33 broad categories, with evaluation questions automatically generated and human-verified. We assess five state-of-the-art knowledge editing methods on three prominent LMMs, revealing that no method excels across all criteria, and that visual and user-specific edits are particularly challenging. MMKE-Bench sets a new standard for evaluating the robustness of multimodal knowledge editing techniques, driving progress in this rapidly evolving field.
Yuntao Du 0001, Kailin Jiang, Zhi Gao 0002, Chenrui Shi, Zilong Zheng, Siyuan Qi, Qing Li 0003
ICLR2
2025 In-Context Editing: Learning Knowledge from Self-Induced Distributions
abstract
In scenarios where language models must incorporate new information efficiently without extensive retraining, traditional fine-tuning methods are prone to overfitting, degraded generalization, and unnatural language generation. To address these limitations, we introduce Consistent In-Context Editing (ICE), a novel approach leveraging the model's in-context learning capability to optimize towards a contextual distribution rather than a one-hot target. ICE introduces a simple yet effective optimization framework for the model to internalize new knowledge by aligning its output distributions with and without additional context. This method enhances the robustness and effectiveness of gradient-based tuning methods, preventing overfitting and preserving the model's integrity. We analyze ICE across four critical aspects of knowledge editing: accuracy, locality, generalization, and linguistic quality, demonstrating its advantages. Experimental results confirm the effectiveness of ICE and demonstrate its potential for continual editing, ensuring that the integrity of the model is preserved while updating information.
Siyuan Qi, Bangcheng Yang, Kailin Jiang, Xiaobo Wang 0004, Jiaqi Li 0021, Yifan Zhong, Yaodong Yang 0001, Zilong Zheng
ICLR3
2024 A classification method of marine mammal calls based on two-channel fusion network
abstract
Abstract Marine mammals are an important part of marine ecosystems, and human intervention seriously threatens their living environments. Few studies exist on the marine mammal call recognition task, and the accuracy of current research needs to improve. In this paper, a novel MG-ResFormer two-channel fusion network architecture is proposed, which can extract local features and global timing information from sound signals almost perfectly. Second, in the input stage of the model, we propose an improved acoustic feature energy fingerprint, which is different from the traditional single feature approach. This feature also contains frequency, energy, time sequence and other speech information and has a strong identity. Additionally, to achieve more reliable accuracy in the multiclass call recognition task, we propose a multigranular joint layer to capture the family and genus relationships between classes. In the experimental section, the proposed method is compared with the existing feature extraction methods and recognition methods. In addition, this paper also compares with the latest research, and the proposed method is the most advanced algorithm thus far. Ultimately, our proposed method achieves an accuracy of 99.39% in the marine mammal call recognition task.
Kailin Jiang, Haibo Pu, Jun Li 0114
Appl. Intell.4
2023 SDSCNet: an instance segmentation network for efficient monitoring of goose breeding conditions
Houcheng Su, Tianyu Xie 0002, Jianan Yuan, Kailin Jiang, Xuliang Duan
Appl. Intell.7