Jiannan Cao

dblp:361/2118 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 45% Efficient and distributed learning · 14% Question answering and dialogue systems · 14%
Network and information security
1 paper
Privacy and data protection · 100%
Software engineering, system software, and programming languages
1 paper
Program verification · 100%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
coreference resolution
0.912025
Bridging Context Gaps: Leveraging Coreference Resolution for Long Contextual Understanding · ICLR 2025
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
long-context question answering
0.912025
Bridging Context Gaps: Leveraging Coreference Resolution for Long Contextual Understanding · ICLR 2025
Machine learning › Efficient and distributed learning › memory-efficient training
memory-efficient fine-tuning
0.912025
DP-MemArc: Differential Privacy Transfer Learning for Memory Efficient Language Models · AAAI 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
task planning
0.912025
Tool-Planner: Task Planning with Clusters across Multiple Tools · ICLR 2025
Natural language and speech › Language models and text generation › LLM agents
tool learning
0.912025
Tool-Planner: Task Planning with Clusters across Multiple Tools · ICLR 2025
Natural language and speech › Language models and text generation › LLM agents › tool use
tool selection
0.912025
Tool-Planner: Task Planning with Clusters across Multiple Tools · ICLR 2025
Privacy and data protection
differential privacy
0.912025
DP-MemArc: Differential Privacy Transfer Learning for Memory Efficient Language Models · AAAI 2025
Program verification
equivalence checking
0.912025
EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking · EMNLP 2025
Natural language and speech › Language models and text generation
large language model
0.312025
DP-MemArc: Differential Privacy Transfer Learning for Memory Efficient Language Models · AAAI 2025

Methods — techniques the papers use, named apart from their topics

side network · 1.7reversible network · 1.7benchmarking · 1.7mention replacement · 0.9mention distance computation · 0.9large language model planning · 0.9
YearPublicationVenuePosition
2025 DP-MemArc: Differential Privacy Transfer Learning for Memory Efficient Language Models
abstract
Large language models have repeatedly shown outstanding performance across diverse applications. However, deploying these models can inadvertently risk user privacy. The significant memory demands during training pose a major challenge in terms of resource consumption. This substantial size places a heavy load on memory resources, raising considerable practical concerns. In this paper, we introduce DP-MemArc, a novel training framework aimed at reducing the memory costs of large language models while emphasizing the protection of user data privacy. DP-MemArc incorporates side network or reversible network designs to support a variety of differential privacy memory-efficient fine-tuning schemes. Our approach not only achieves about 2.5 times in memory optimization but also ensures robust privacy protection, keeping user data secure and confidential. Extensive experiments have demonstrated that DP-MemArc effectively provides differential privacy-efficient fine-tuning across different task scenarios.
Yanming Liu 0003, Xinyue Peng, Xiaolan Ke, Songhang Deng, Jiannan Cao, Mengchen Fu, Xuhong Zhang 0002, Jianwei Yin, Tianyu Du
AAAI6
2025 EquiBench: Benchmarking Large Language Models' Reasoning about Program Semantics via Equivalence Checking
abstract
Anjiang Wei, Jiannan Cao, Ran Li, Hongyu Chen, Yuhui Zhang, Ziheng Wang, Yuan Liu, Thiago S. F. X. Teixeira, Diyi Yang, Ke Wang, Alex Aiken. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Anjiang Wei, Jiannan Cao, Thiago S. F. X. Teixeira, Diyi Yang, Alex Aiken
EMNLP2
2025 Bridging Context Gaps: Leveraging Coreference Resolution for Long Contextual Understanding
abstract
Large language models (LLMs) have shown remarkable capabilities in natural language processing; however, they still face difficulties when tasked with understanding lengthy contexts and executing effective question answering. These challenges often arise due to the complexity and ambiguity present in longer texts. To enhance the performance of LLMs in such scenarios, we introduce the Long Question Coreference Adaptation (LQCA) method. This innovative framework focuses on coreference resolution tailored to long contexts, allowing the model to identify and manage references effectively. The LQCA method encompasses four key steps: resolving coreferences within sub-documents, computing the distances between mentions, defining a representative mention for coreference, and answering questions through mention replacement. By processing information systematically, the framework provides easier-to-handle partitions for LLMs, promoting better understanding. Experimental evaluations on a range of LLMs and datasets have yielded positive results, with a notable improvements on OpenAI-o1-mini and GPT-4o models, highlighting the effectiveness of leveraging coreference resolution to bridge context gaps in question answering. Our code is public at https://github.com/OceannTwT/LQCA.
Yanming Liu 0003, Xinyue Peng, Jiannan Cao, Shi Bo, Yanxin Shen, Tianyu Du, Jianwei Yin, Xuhong Zhang 0002
ICLR3
2025 Tool-Planner: Task Planning with Clusters across Multiple Tools
abstract
Large language models (LLMs) have demonstrated exceptional reasoning capabilities, enabling them to solve various complex problems. Recently, this ability has been applied to the paradigm of tool learning. Tool learning involves providing examples of tool usage and their corresponding functions, allowing LLMs to formulate plans and demonstrate the process of invoking and executing each tool. LLMs can address tasks that they cannot complete independently, thereby enhancing their potential across different tasks. However, this approach faces two key challenges. First, redundant error correction leads to unstable planning and long execution time. Additionally, designing a correct plan among multiple tools is also a challenge in tool learning. To address these issues, we propose Tool-Planner, a task-processing framework based on toolkits. Tool-Planner groups tools based on the API functions with the same function into a toolkit and allows LLMs to implement planning across the various toolkits. When a tool error occurs, the language model can reselect and adjust tools based on the toolkit. Experiments show that our approach demonstrates a high pass and win rate across different datasets and optimizes the planning scheme for tool learning in models such as GPT-4 and Claude 3, showcasing the potential of our method. Our code is public at https://github.com/OceannTwT/Tool-Planner.
Yanming Liu 0003, Xinyue Peng, Jiannan Cao, Shi Bo, Xuhong Zhang 0002, Jianwei Yin, Tianyu Du
ICLR3
2024 SPA: Towards A Computational Friendly Cloud-Base and On-Devices Collaboration Seq2seq Personalized Generation with Causal Inference
Yanming Liu 0003, Xinyue Peng, Jiannan Cao, Le Dai, Xingzu Liu, Ruilin Nong, Songhang Deng
PRICAI (2)3