Jason Cai

dblp:371/9818 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0002-9190-0252ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 8 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 87% Reinforcement learning · 13%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 87% Recommender systems · 13%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
LLM agents
1.722025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025
SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection · EMNLP 2025
Natural language and speech › Language models and text generation
text summarization
1.622025
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages · ACL (1) 2025
FineSurE: Fine-grained Summarization Evaluation using LLMs · ACL (1) 2024
Natural language and speech › Language models and text generation
memory augmentation
0.912025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025
Machine learning › Reinforcement learning
self-improving agent
0.912025
SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection · EMNLP 2025
Information retrieval › retrieval-augmented generation
memory retrieval
0.912025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025
Information retrieval
retrieval-augmented generation
0.912025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.812024
FineSurE: Fine-grained Summarization Evaluation using LLMs · ACL (1) 2024
Natural language and speech › Language models and text generation › text summarization
summarization evaluation
0.812024
FineSurE: Fine-grained Summarization Evaluation using LLMs · ACL (1) 2024
Natural language and speech › Language models and text generation › text summarization › neural summarization
LLM-based summarization
0.312025
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages · ACL (1) 2025
Recommender systems › interactive recommendation
conversational recommendation
0.312025
MemInsight: Autonomous Memory Augmentation for LLM Agents · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

autonomous memory augmentation · 1.7retrospective language model · 0.9multi-level reflection synthesis · 0.9fine-tuning · 0.9large language model · 0.8faithfulness assessment · 0.8completeness assessment · 0.8
YearPublicationVenuePosition
2025 Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
abstract
Hyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng, Nicole Hee-Yeon Kim, Taewon Yun, Hang Su, Jason Cai, Hwanjun Song. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hyangsuk Min, Yuho Lee, Minjeong Ban, Nicole Hee-Yeon Kim, Taewon Yun, Jason Cai, Hwanjun Song
ACL (1)8
2025 SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection
abstract
Despite the rapid advancements in LLM agents, they still face the challenge of generating meaningful reflections due to inadequate error analysis and a reliance on rare successful trajectories, especially in complex tasks.In this work, we propose SAMULE, a new framework for self-learning agents powered by a retrospective language model that is trained based on Multi-Level Reflection Synthesis.It first synthesizes high-quality reflections across three complementary levels: Single-Trajectory Learning (micro-level) for detailed error correction; Intra-Task Learning (meso-level) to build error taxonomies across multiple trials of the same task, and Inter-Task Learning (macrolevel) to extract transferable insights based on same typed errors from diverse task failures.Then we fine-tune a language model serving as the retrospective model to generate reflections during inference.We further extend our framework to interactive settings through a foresightbased reflection mechanism, enabling agents to proactively reflect and adapt during user interactions by comparing predicted and actual responses.Extensive experiments on three challenging benchmarks-TravelPlanner, NAT-URAL PLAN, and Tau-bench-demonstrate that our approach significantly outperforms reflection-based baselines.Our results highlight the critical role of well-designed reflection synthesis and failure-centric learning in building self-improving LLM agents.
Yubin Ge, Salvatore Romeo, Jason Cai, Monica Sunkara
EMNLP3
2025 MemInsight: Autonomous Memory Augmentation for LLM Agents
abstract
Large language model (LLM) agents have evolved to intelligently process information, make decisions, and interact with users or tools. A key capability is the integration of long-term memory capabilities, enabling these agents to draw upon historical interactions and knowledge. However, the growing memory size and need for semantic structuring pose significant challenges. In this work, we propose an autonomous memory augmentation approach, MemInsight, to enhance semantic data representation and retrieval mechanisms. By leveraging autonomous augmentation to historical interactions, LLM agents are shown to deliver more accurate and contextualized responses. We empirically validate the efficacy of our proposed approach in three task scenarios; conversational recommendation, question answering and event summarization. On the LLM-REDIAL dataset, MemInsight boosts persuasiveness of recommendations by up to 14%. Moreover, it outperforms a RAG baseline by 34% in recall for LoCoMo retrieval. Our empirical results show the potential of MemInsight to enhance the contextual performance of LLM agents across multiple tasks.
Rana Salama, Jason Cai, Michelle Yuan, Anna Currey, Monica Sunkara, Yassine Benajiba
EMNLP2
2025 Learning to Summarize from LLM-generated Feedback
abstract
Hwanjun Song, Taewon Yun, Yuho Lee, Jihwan Oh, Gihun Lee, Jason Cai, Hang Su. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hwanjun Song, Taewon Yun, Yuho Lee, Jihwan Oh, Gihun Lee, Jason Cai
NAACL (Long Papers)6
2024 FineSurE: Fine-grained Summarization Evaluation using LLMs
abstract
Automated evaluation is crucial for streamlining text summarization benchmarking and model development, given the costly and timeconsuming nature of human evaluation.Traditional methods like ROUGE do not correlate well with human judgment, while recently proposed LLM-based metrics provide only summary-level assessment using Likertscale scores.This limits deeper model analysis, e.g., we can only assign one hallucination score at the summary level, while at the sentence level, we can count sentences containing hallucinations.To remedy those limitations, we propose FineSurE, a fine-grained evaluator specifically tailored for the summarization task using large language models (LLMs).It also employs completeness and conciseness criteria, in addition to faithfulness, enabling multi-dimensional assessment.We compare various open-source and proprietary LLMs as backbones for FineSurE.In addition, we conduct extensive benchmarking of FineSurE against SOTA methods including NLI-, QA-, and LLM-based methods, showing improved performance especially on the completeness and conciseness dimensions.The code is available at https://github.com/ DISL-Lab/FineSurE-ACL24.
Hwanjun Song, Igor Shalyminov, Jason Cai, Saab Mansour
ACL (1)4
2024 Zero-Shot End-to-End Spoken Language Understanding via Cross-Modal Selective Self-Training
abstract
Jianfeng He, Julian Salazar, Kaisheng Yao, Haoqi Li, Jason Cai. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Julian Salazar, Kaisheng Yao, Jason Cai
EACL (1)5
2024 CERET: Cost-Effective Extrinsic Refinement for Text Generation
abstract
Jason Cai, Hang Su, Monica Sunkara, Igor Shalyminov, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jason Cai, Monica Sunkara, Igor Shalyminov, Saab Mansour
NAACL-HLT1
2024 Semi-Supervised Dialogue Abstractive Summarization via High-Quality Pseudolabel Selection
abstract
Jianfeng He, Hang Su, Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Jason Cai, Igor Shalyminov, Hwanjun Song, Saab Mansour
NAACL-HLT3