VLDB 2026 Research / reviewers in the wild / expert
Ercong Nie
dblp:336/4767
· DBLP profile ↗
12ranked-venue papers
1as first author
12since 2021 · last 2026
0000-0003-1453-4460ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | XToM: Exploring the Multilingual Theory of Mind for Large Language ModelsabstractTheory of Mind (ToM), the ability to infer mental states in others, is pivotal for human social cognition. Existing evaluations of ToM in LLMs are largely limited to English, neglecting the linguistic diversity that shapes human cognition. This limitation raises a critical question: can LLMs exhibit Multilingual Theory of Mind, which is the capacity to reason about mental states across diverse linguistic contexts? To address this gap, we present XToM, a rigorously validated multilingual benchmark that evaluates ToM across five languages and incorporates diverse, contextually rich task scenarios. Using XToM, we systematically evaluate LLMs (e.g., DeepSeek R1), revealing a pronounced dissonance: while models excel in multilingual language understanding, their ToM performance varies across languages. Our findings expose limitations in LLMs' ability to replicate human-like mentalizing across linguistic contexts. Chunkit Chan, Yauwai Yim, Hongchuan Zeng, Zhiying Zou, Xinyuan Cheng, Zhifan Sun, Zheye Deng, Kawai Chung, Yuzhuo Ao, Yixiang Fan, Cheng Jiayang, Ercong Nie, Ginny Y. Wong, Helmut Schmid, Hinrich Schütze, Simon See, Yangqiu Song |
ACL (1) | 12 |
| 2026 | Look Within or Beyond? A Theoretical Comparison Between Parameter-Efficient and Full Fine-TuningabstractYongKang Liu, Xingle Xu, Ercong Nie, Zijing Wang, Shi Feng, Daling Wang, Qian Li, Hinrich Schuetze. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yongkang Liu 0002, Xingle Xu, Ercong Nie, Shi Feng 0001, Daling Wang, Qian Li 0043, Hinrich Schütze |
ACL (1) | 3 |
| 2026 | SAD: A Large-Scale Strategic Argumentative Dialogue DatasetabstractYongKang Liu, Jiayang Yu, Mingyang Wang, Yiqun Zhang, Ercong Nie, Shi Feng, Daling Wang, Kaisong Song, Hinrich Schuetze. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yongkang Liu 0002, Jiayang Yu, Mingyang Wang 0003, Ercong Nie, Shi Feng 0001, Daling Wang, Kaisong Song, Hinrich Schütze |
ACL (1) | 5 |
| 2026 | Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement LearningabstractSikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schuetze, Volker Tresp, Yunpu Ma. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma 0001, Jinhe Bi, Kristian Kersting, Jeff Z. Pan, Hinrich Schütze, Volker Tresp, Yunpu Ma |
ACL (1) | 4 |
| 2026 | CoDAE: Adapting Large Language Models for Education via Chain-of-Thought Data AugmentationabstractLarge Language Models (LLMs) are increasingly employed as AI tutors due to their scalability and potential for personalized instruction. However, off-the-shelf LLMs often underperform in educational settings: they frequently reveal answers too readily, fail to adapt their responses to student uncertainty, and remain vulnerable to emotionally manipulative prompts. To address these challenges, we introduce CoDAE, a framework that adapts LLMs for educational use through Chain-of-Thought (CoT) data augmentation. We collect real-world dialogues between students and a ChatGPT-based tutor and enrich them using CoT prompting to promote step-by-step reasoning and pedagogically aligned guidance. Furthermore, we design targeted dialogue cases to explicitly mitigate three key limitations: over-compliance, low response adaptivity, and threat vulnerability. We fine-tune four open-source LLMs on different variants of the augmented datasets and evaluate them in simulated educational scenarios using both automatic metrics and LLM-as-a-judge assessments. Our results show that models fine-tuned with CoDAE deliver more pedagogically appropriate guidance, better support reasoning processes, and effectively resist premature answer disclosure. Shuzhou Yuan, William LaCroix, Hardik Ghoshal, Ercong Nie, Michael Färber 0001 |
LREC | 4 |
| 2026 | Query Expansion in the Age of Pre-trained and Large Language Models: A Comprehensive SurveyabstractModern information retrieval (IR) must reconcile short, ambiguous queries with increasingly diverse and dynamic corpora. Query expansion (QE) remains a core technique for mitigating vocabulary mismatch, but its design space has been reshaped by pre-trained and large language models (PLMs/LLMs). This survey reviews QE methods in the PLM/LLM era and provides a unified view of the emerging landscape. We first summarize how different model families enable new expansion behaviors, including stronger contextualization, more controllable generation, and instruction-following. We then organize recent techniques along four complementary design dimensions: where expansion is injected in the pipeline, how it is grounded and interacts with corpus evidence, how it is learned or aligned, and how structured knowledge, such as knowledge graphs, is incorporated. Beyond taxonomy, we synthesize application patterns and deployment considerations across representative retrieval settings, highlighting practical tradeoffs among effectiveness, controllability, grounding quality, and operating cost. Finally, we outline open challenges and future directions toward more reliable, safe, efficient, and continually adaptive QE under real-world constraints (resources are available at https://github.com/lmh0921/QueryExpansion-PLM-LLM-Survey-paperList ). Minghan Li 0003, Xinxuan Lv, Junjie Zou, Tongna Chen, Suchao An, Ercong Nie, Guodong Zhou 0001 |
ACM Trans. Inf. Syst. | 7 |
| 2025 | Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language ModelsabstractMingyang Wang, Heike Adel, Lukas Lange, Yihong Liu, Ercong Nie, Jannik Strötgen, Hinrich Schuetze. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Mingyang Wang 0003, Heike Adel, Lukas Lange, Yihong Liu 0001, Ercong Nie, Jannik Strötgen, Hinrich Schütze |
ACL (1) | 5 |
| 2025 | BMIKE-53: Investigating Cross-Lingual Knowledge Editing with In-Context LearningabstractThis paper introduces BMIKE-53, a comprehensive benchmark for cross-lingual in-context knowledge editing (IKE) across 53 languages, unifying three knowledge editing (KE) datasets: zsRE, CounterFact, and WikiFactDiff.Crosslingual KE, which requires knowledge edited in one language to generalize across others while preserving unrelated knowledge, remains underexplored.To address this gap, we systematically evaluate IKE under zero-shot, oneshot, and few-shot setups, incorporating tailored metric-specific demonstrations.Our findings reveal that model scale and demonstration alignment critically govern cross-lingual IKE efficacy, with larger models and tailored demonstrations significantly improving performance.Linguistic properties, particularly script type, strongly influence performance variation across languages, with non-Latin languages underperforming due to issues like language confusion. Ercong Nie, Mingyang Wang 0003, Zifeng Ding, Helmut Schmid, Hinrich Schütze |
ACL (1) | 1 |
| 2025 | Why Lift so Heavy? Slimming Large Language Models by Cutting Off the LayersabstractLarge Language Models (LLMs) demonstrate exceptional language understanding and generation capabilities by learning from context. Leveraging the strong in-context learning (ICL) abilities of LLMs, prompt-based fine-tuning has proven to be effective for enhancing the adaptability and alignment of LLMs, especially in low-data scenarios. However, the billions of parameters resulting from layer stacking in LLMs present significant computational challenges, limiting the practicality of fine-tuning. To tackle this problem, we explore the application of layer-wise model pruning in prompt-based fine-tuning of LLMs for few-shot learning scenarios. Our approach involves dropping certain model layers and fine-tuning the model with the remaining layers. Surprisingly, we observe that even with fewer layers, LLMs maintain similar or better performance levels, particularly in prompt-based fine-tuning for text classification tasks. Remarkably, in certain cases, models with a single layer outperform their fully layered counterparts. These findings offer valuable insights for future work aimed at mitigating the size constraints of LLMs while preserving their performance, thereby opening avenues for significantly more efficient use of LLMs. Shuzhou Yuan, Ercong Nie, Bolei Ma, Michael Färber 0001 |
IJCNN | 2 |
| 2024 | Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models Using Minimal PairsabstractInspired by cognitive neuroscience studies, we introduce a novel “decoding probing” method that uses minimal pairs benchmark (BLiMP) to probe internal linguistic characteristics in neural language models layer by layer. By treating the language model as the brain and its representations as “neural activations”, we decode grammaticality labels of minimal pairs from the intermediate layers’ representations. This approach reveals: 1) Self-supervised language models capture abstract linguistic structures in intermediate layers that GloVe and RNN language models cannot learn. 2) Information about syntactic grammaticality is robustly captured through the first third layers of GPT-2 and also distributed in later layers. As sentence complexity increases, more layers are required for learning grammatical capabilities. 3) Morphological and semantics/syntax interface-related features are harder to capture than syntax. 4) For Transformer-based models, both embeddings and attentions capture grammatical features but show distinct patterns. Different attention heads exhibit similar tendencies toward various linguistic phenomena, but with varied contributions. Linyang He, Peili Chen, Ercong Nie, Yuanning Li, Jonathan Brennan |
LREC/COLING | 3 |
| 2024 | ToPro: Token-Level Prompt Decomposition for Cross-Lingual Sequence Labeling TasksabstractBolei Ma, Ercong Nie, Shuzhou Yuan, Helmut Schmid, Michael Färber, Frauke Kreuter, Hinrich Schuetze. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Bolei Ma, Ercong Nie, Shuzhou Yuan, Helmut Schmid, Michael Färber 0001, Frauke Kreuter, Hinrich Schütze |
EACL (1) | 2 |
| 2024 | A Unified Data Augmentation Framework for Low-Resource Multi-domain Dialogue Generation
Yongkang Liu 0002, Ercong Nie, Shi Feng 0001, Zifeng Ding, Daling Wang, Yifei Zhang 0003, Hinrich Schütze |
ECML/PKDD (2) | 2 |