Kexin Ma 0008

dblp:40/4446-8 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0000-0001-9309-9174ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 29% Knowledge representation and reasoning · 25% Planning, search and constraint satisfaction · 25%
Databases, data mining, and information retrieval
1 paper
Data models and query languages · 67% Information retrieval · 33%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › agent planning
embodied planning
1.012026
Conflict-Aware Memory for Embodied Agents: Enhancing Vector Data Quality via Detection Rules · ACL (1) 2026
Natural language and speech › Language models and text generation › natural language understanding
ambiguity detection
0.912025
CLEAR: A Parser-Independent Disambiguation Framework for NL2SQL · ICDE 2025
Natural language and speech › Question answering and dialogue systems › interactive question answering
interactive clarification
0.912025
CLEAR: A Parser-Independent Disambiguation Framework for NL2SQL · ICDE 2025
Information retrieval
query reformulation
0.912025
CLEAR: A Parser-Independent Disambiguation Framework for NL2SQL · ICDE 2025
Data models and query languages › SQL
SQL query generation
0.912025
CLEAR: A Parser-Independent Disambiguation Framework for NL2SQL · ICDE 2025
Data models and query languages › natural language interface › natural language interface to database
text-to-SQL
0.912025
CLEAR: A Parser-Independent Disambiguation Framework for NL2SQL · ICDE 2025
Natural language and speech › Language models and text generation › LLM agents
large language model planning
0.312026
Conflict-Aware Memory for Embodied Agents: Enhancing Vector Data Quality via Detection Rules · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

rewriting rules · 1.7large language model · 1.7interactive selection · 1.7vector similarity search · 1.0conflict detection rules · 1.0
YearPublicationVenuePosition
2026 Conflict-Aware Memory for Embodied Agents: Enhancing Vector Data Quality via Detection Rules
abstract
Embodied agents have successfully leveraged large language models (LLMs) to better transform human instructions and images into executable task plans.Furthermore, memories of agents can be leveraged to achieve continual self-learning and optimization.However, vector data quality problems emerge in memories when they are projected into vector space, especially in discerning contextually similar but semantically conflicting sentences and highly similar images.This is particularly detrimental to embodied AI as it potentially distorts the robot's actions.To address this challenge, we propose Conflict Detection Rules (CDRs) to identify and manage data quality issues in vector knowledge bases, which assist in correcting the index structure and further improving the answer quality.Experimental results show that planners with CDRs exceed the basic LLM planner by 15.25% and 14.25% in grammatical accuracy (GA) and interpretation accuracy (IA) on average, respectively.Moreover, the entire workflow has been successfully integrated into various scenarios, demonstrating its practical applicability and robustness in the real world 1 .
Kexin Ma 0008, Haotian Wang 0001, Shenglin Chen, Yishuai Cai, Ruochun Jin
ACL (1)1
2026 Neuro-symbolic Hierarchical Learning for Long-Horizon Robotic Tasks
abstract
Recent advances in foundation models have motivated hybrid programming that integrates natural language descriptions, formal specifications, and executable code. A critical challenge in such systems lies in achieving semantic alignment across heterogeneous representations at different abstraction levels. This challenge is particularly pressing in programmatic reinforcement learning (PRL) for robotics, where long-horizon, sparse-reward tasks demand tight coordination between symbolic reasoning and continuous control. Existing approaches either rely on manually engineered symbolic representations or on LLM-generated plans that might be untrustworthy or infeasible to execute, leaving such fundamental gaps unaddressed. We present a closed-loop, counterexample-guided synthesis framework that unifies LLM-based planning, formal verification, and differentiable behavior tree (BT) synthesis for neuro-symbolic policy learning. The framework first converts natural language task descriptions into PDDL, then generates valid high-level plans via a Guess-Check-Critique loop that interleaves LLM generation with SMT-based verification. We automatically compile symbolic abstractions in verified plans into parameterized termination conditions, and co-optimize them with low-level policies to enforce semantic consistency. The learned sub-policies are composed into an integrated BT and fine‑tuned to ensure task-level executability. Furthermore, we perform closed-loop iterations of abstractions and compilations using feedback from verification and learning, while incrementally building a reusable skill library for efficient knowledge transfer. Experiments on challenging long-horizon robotic tasks show that our method exceeds state-of-the-art methods while providing interpretability and generalization.
Ziji Wu, Zhengyi Ma, Kexin Ma 0008, Ji Wang 0001
Proc. ACM Program. Lang.4
2025 CLEAR: A Parser-Independent Disambiguation Framework for NL2SQL
abstract
Parsing Natural Language to SQL (NL2SQL) helps users who are not proficient in databases to efficiently query desired data through natural language. Although existing NL2SQL parsers demonstrate good capabilities in processing clear queries, ambiguity still remains an unresolved issue which makes parsers produce unstable outputs that deviate from the user's actual intent. To bridge the gap, this paper introduces the CLEAR framework, a systematic study of disambiguation for NL2SQL, including ambiguity detection, clarification, and reformulation, which benefits any NL2SQL parsers. Firstly, CLEAR employs a pipeline using Large Language Models (LLMs) and a series of rules to detect ambiguities, thus obtaining the “candidate mapping” for ambiguity representation. Secondly, an interactive selection module is employed to collect the clarification information from users through multiple-choice questions, thus obtaining the “selection mapping”. Finally, rewriting rules are employed to reformulate the question and schema, thus obtaining a clear input for parsers to generate clear SQLs. Furthermore, we construct CLAMBSQL, a novel benchmark for systematic evaluation for NL2SQL disambiguation, which contains fine-grained ambiguity and clarification annotations. Experiments on various datasets and baselines demonstrate that CLEAR can successfully address seven types of ambiguity. When parsers are integrated with CLEAR, the performance of ambiguous SQLs detection achieves a significant improvement of 30.5 % on AMBROSIA in the AllFound metric and 21.1 % on AmbiQT in the BothInTop-5 metric, the performance of ambiguity clarification achieves a remarkable improvement of 16.2 % on CLAMBSQL in the CEX metric, and the performance of the general prediction achieves an increase of 1.6 % in the EX metric and 7.7 % in the CSR metric on BIRD. The CLEAR code and CLAMBSQL dataset are available at https://github.com/mengzhang18/CLEAR.
Kexin Ma 0008, Kedi Zhang, Yuanxi Peng, Ruochun Jin
ICDE2
2025 Enhancing Vector Data Quality through Negative Learning for Retrieval-augmented Large Models
abstract
Retrieval-augmented Large Models (RALMs) have emerged as a promising paradigm to enhance large language models (LLMs) by integrating external knowledge. However, the inherent complexity of vector-based retrieval often introduces noise and inaccuracies that can compromise the model’s performance. While existing approaches primarily focus on filtering retrieved contexts, our study reveals a novel perspective: seemingly irrelevant or contradictory knowledge can serve as valuable learning signals for LLMs to improve the data quality of vector database. We propose the NDIC (Negatives Driven Index Correction) framework, a context-aware retrieval method inspired by relational database theory. By introducing Contexts Clear Matching Dependence (CCMDs), our approach enables LLM to precisely categorize retrieval contexts into three distinct types: positively matched, negatively matched, and unclear. Unlike traditional methods, we strategically utilize both positively and negatively matched contexts to refine the learning process and enhance knowledge utilization. Experiments across diverse LLMs demonstrate a 6.17% improvement in accuracy in challenging question-answering tasks, effectively transforming potentially noisy retrievals into structured learning opportunities.
Limei Yao, Kexin Ma 0008, Ruochun Jin, Haoqi Zheng
IJCNN2