Kangcong Li

dblp:399/6665 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Vision and language · 23% Language models and text generation · 23% Deep learning architectures and training · 23%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
feedforward neural network
0.912025
PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding · NeurIPS 2025
Computer vision › Vision and language
image captioning
0.912025
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning · ICCV 2025
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context understanding
0.912025
PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding · NeurIPS 2025
Machine learning › Reinforcement learning
reward design
0.312025
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning · ICCV 2025

Methods — techniques the papers use, named apart from their topics

scene graph parsing · 0.9reinforcement learning · 0.9persistent activity mechanism · 0.9expert clustering · 0.9direct preference optimization · 0.9
YearPublicationVenuePosition
2025 SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
abstract
We propose SC-Captioner, a reinforcement learning framework that enables the self-correcting capability of image caption models. Our crucial technique lies in the design of the reward function to incentivize accurate caption corrections. Specifically, the predicted and reference captions are decomposed into object, attribute, and relation sets using scene-graph parsing algorithms. We calculate the set difference between sets of initial and self-corrected captions to identify added and removed elements. These elements are matched against the reference sets to calculate correctness bonuses for accurate refinements and mistake punishments for wrong additions and removals, thereby forming the final reward. For image caption quality assessment, we propose a set of metrics refined from CAPTURE that alleviate its incomplete precision evaluation and inefficient relation matching problems. Furthermore, we collect a fine-grained annotated image caption dataset, RefinedCaps, consisting of 6.5K diverse images from COCO dataset. Experiments show that applying SC-Captioner on large visual-language models can generate better image captions across various scenarios, significantly outperforming the direct preference optimization training strategy.
Lin Zhang 0055, Xianfang Zeng, Kangcong Li, Gang Yu 0002, Tao Chen 0003
ICCV3
2025 PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding
abstract
While Large Language Models (LLMs) demonstrate strong performance across domains, their long-context capabilities are limited by transient neural activations causing information decay and unstructured feed-forward network (FFN) weights leading to semantic fragmentation. Inspired by the brain’s working memory and cortical modularity, we propose PaceLLM, featuring two innovations: (1) a Persistent Activity (PA) Mechanism that mimics prefrontal cortex (PFC) neurons’ persistent firing by introducing an activation-level memory bank to dynamically retrieve, reuse, and update critical FFN states, addressing contextual decay; and (2) Cortical Expert (CE) Clustering that emulates task-adaptive neural specialization to reorganize FFN weights into semantic modules, establishing cross-token dependencies and mitigating fragmentation. Extensive evaluations show that PaceLLM achieves 6% improvement on LongBench’s Multi-document QA and 12.5–17.5% performance gains on $\infty$-Bench tasks, while extending measurable context length to 200K tokens in Needle-In-A-Haystack (NIAH) tests. This work pioneers brain-inspired LLM optimization and is complementary to other works. Besides, it can be generalized to any model and enhance their long-context performance and interpretability without structural overhauls.
Kangcong Li, Peng Ye 0006, Chongjun Tu, Lin Zhang 0055, Chunfeng Song, Qihao Zheng, Tao Chen 0003
NeurIPS1