Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Keer Lu

dblp:364/7890 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0005-5966-8309ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 76% Learning paradigms · 18% Multi-agent systems · 6%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 57% Computational social science and digital humanities · 43%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 50% Recommender systems · 50%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
Med-R2: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine · WWW 2026
Medical and health informatics
clinical decision support
1.012026
Med-R2: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine · WWW 2026
Machine learning › Learning paradigms › continual learning › catastrophic forgetting
catastrophic forgetting mitigation
0.912025
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs · EMNLP 2025
Natural language and speech › Language models and text generation
instruction tuning
0.912025
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning · ICLR 2025
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.912025
VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs · EMNLP 2025
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models
0.912025
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning · ICLR 2025
Recommender systems
multi-objective optimization
0.912025
DataSculpt: A Holistic Data Management Framework for Long-Context LLMs Training · ICDE 2025
Machine learning and data management
training data management
0.912025
DataSculpt: A Holistic Data Management Framework for Long-Context LLMs Training · ICDE 2025
Visualization and visual analytics
visual analytics
0.812024
LiberRoad: Probing into the Journey of Chinese Classics Through Visual Analytics · IEEE Trans. Vis. Comput. Graph. 2024
Knowledge, reasoning and agents › Multi-agent systems
multi-agent environments
0.312025
Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning · ICLR 2025
Visualization and visual analytics
uncertainty visualization
0.212024
LiberRoad: Probing into the Journey of Chinese Classics Through Visual Analytics · IEEE Trans. Vis. Comput. Graph. 2024

Methods — techniques the papers use, named apart from their topics

retrieval-augmented generation · 2.0reasoning · 2.0lens magnification · 1.5clustering · 1.5circle packing · 1.5synthetic data generation · 0.9semantic clustering · 0.9multi-objective greedy search · 0.9multi-agent simulation · 0.9dynamic domain weighting · 0.9data composition · 0.9
YearPublicationVenuePosition
2026 Med-R2: Crafting Trustworthy LLM Physicians via Retrieval and Reasoning of Evidence-Based Medicine
Keer Lu, Da Pan 0003, Shusen Zhang, Guosheng Dong, Huang Leng, Bin Cui 0001, Zhonghai Wu, Wentao Zhang 0001
WWW1
2025 VersaTune: An Efficient Data Composition Framework for Training Multi-Capability LLMs
abstract
As demonstrated by the proprietary Large Language Models (LLMs) such as GPT and Claude series, LLMs have the potential to achieve remarkable proficiency across a wide range of domains, including law, medicine, finance, science, code, etc., all within a single model. These capabilities are further augmented during the Supervised Fine-Tuning (SFT) phase. Despite their potential, existing work mainly focuses on domain-specific enhancements during fine-tuning, the challenge of which lies in catastrophic forgetting of knowledge across other domains. In this study, we introduce VersaTune, a novel data composition framework designed for enhancing LLMs’ overall multi-domain capabilities during training. We begin with detecting the distribution of domain-specific knowledge within the base model, followed by the training data composition that aligns with the model’s existing knowledge distribution. During the subsequent training process, domain weights are dynamically adjusted based on their learnable potential and forgetting degree. Experimental results indicate that VersaTune is effective in multi-domain fostering, with an improvement of 29.77% in the overall multi-ability performances compared to uniform domain weights. Furthermore, we find that Qwen-2.5-32B + VersaTune even surpasses frontier models, including GPT-4o, Claude3.5-Sonnet and DeepSeek-V3 by 0.86%, 4.76% and 4.60%. Additionally, in scenarios where flexible expansion of a specific domain is required, VersaTune reduces the performance degradation in other domains by 38.77%, while preserving the training efficacy of the target domain.
Keer Lu, Keshi Zhao, Bin Cui 0001, Wentao Zhang 0001
EMNLP1
2025 DataSculpt: A Holistic Data Management Framework for Long-Context LLMs Training
abstract
In recent years, foundation models, particularly large language models (LLMs), have demonstrated significant improvements across a variety of tasks. One of their most important features is long-context capability, which enables them to generate extended text with high semantic coherence, retrieving relevant information, and handling tasks with substantial amounts of text efficiently. The key to improving long-context performance lies in effective data organization and management strategies that integrate data from multiple domains and optimize the context window during training. Through extensive experimental analysis, we identified three key challenges in designing effective data management strategies that enable the model to achieve long-context capability without sacrificing performance in other tasks: (1) a shortage of long documents across multiple domains, (2) effective construction of context windows, and (3) efficient organization of large-scale datasets. To address these challenges, we introduce DataSculpt, a novel data management framework designed for long-context training. We first formulate the organization of training data as a multi-objective optimization problem, focusing on attributes including the relevance among documents within the same training sequence, the quantity of concatenated instances, individual document integrity, and computational cost. Specifically, our approach utilizes a coarse-to-fine method to optimize training data organization effectively. We begin by clustering the data based on semantic similarity (coarse), followed by a multi-objective greedy search within each cluster to score and concatenate documents into various context windows (fine). We have deployed DataSculpt as the data management backend for long-context training in Baichuan Inc. Extensive experiments with diverse downstream tasks show that DataSculpt enhances the model's long-context performance by an average of 15.73%, while maintaining the general capabilities with a 4.63% improvement.
Keer Lu, Xiaonan Nie, Da Pan 0003, Shusen Zhang, Keshi Zhao, Weipeng Chen, Zenan Zhou, Guosheng Dong, Bin Cui 0001, Wentao Zhang 0001
ICDE1
2025 Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction Tuning
abstract
Large Language Models (LLMs) have exhibited significant potential in performing diverse tasks, including the ability to call functions or use external tools to enhance their performance. While current research on function calling by LLMs primarily focuses on single-turn interactions, this paper addresses the overlooked necessity for LLMs to engage in multi-turn function calling—critical for handling compositional, real-world queries that require planning with functions but not only use functions. To facilitate this, we introduce an approach, BUTTON, which generates synthetic compositional instruction tuning data via bottom-up instruction construction and top-down trajectory generation. In the bottom-up phase, we generate simple atomic tasks based on real-world scenarios and build compositional tasks using heuristic strategies based on atomic tasks. Corresponding function definitions are then synthesized for these compositional tasks. The top-down phase features a multi-agent environment where interactions among simulated humans, assistants, and tools are utilized to gather multi-turn function calling trajectories. This approach ensures task compositionality and allows for effective function and trajectory generation by examining atomic tasks within compositional tasks. We produce a dataset BUTTONInstruct comprising 8k data points and demonstrate its effectiveness through extensive experiments across various LLMs.
Mingyang Chen 0002, Haoze Sun, Tianpeng Li, Fan Yang 0132, Hao Liang 0017, Keer Lu, Bin Cui 0001, Wentao Zhang 0001, Zenan Zhou, Weipeng Chen
ICLR6
2024 LiberRoad: Probing into the Journey of Chinese Classics Through Visual Analytics
abstract
Books act as a crucial carrier of cultural dissemination in ancient times. This work involves joint efforts between visualization and humanities researchers, aiming at building a holistic view of the cultural exchange and integration between China and Japan brought about by the overseas circulation of Chinese classics. Book circulation data consist of uncertain spatiotemporal trajectories, with multiple dimensions, and movement across hierarchical spaces forms a compound network. LiberRoad visualizes the circulation of books collected in the Imperial Household Agency of Japan, and can be generalized to other book movement data. The LiberRoad system enables a smooth transition between three views (Location Graph, map, and timeline) according to the desired perspectives (spatial or temporal), as well as flexible filtering and selection. The Location Graph is a novel uncertainty-aware visualization method that employs improved circle packing to represent spatial hierarchy. The map view intuitively shows the overall circulation by clustering and allows zooming into single book trajectory with lenses magnifying local movements. The timeline view ranks dynamically in response to user interaction to facilitate the discovery of temporal events. The evaluation and feedback from the expert users demonstrate that LiberRoad is helpful in revealing movement patterns and comparing circulation characteristics of different times and spaces.
Yuhan Guo 0004, Yuchu Luo, Keer Lu, Linfang Li, Haizheng Yang, Xiaoru Yuan
IEEE Trans. Vis. Comput. Graph.3