VLDB 2026 Research / reviewers in the wild / expert
Jingwei Ni
dblp:304/2832
· DBLP profile ↗
12ranked-venue papers
5as first author
12since 2021 · last 2026
0009-0008-8988-301XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsabstractAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Frank Ďurech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan, Skander Moalla, Tiancheng Chen, Vinko Sabolčec, Yixuan Xu, Michael Aerni, Badr AlKhamissi, Inés Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Milan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush Kumar Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Hao Zhao, Alexander Ilic, Ana Klimovic, Andreas Krause, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Alejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou, Antoni-Joan Solergibert i Llaquet, Barna Pásztor, Bettina Messmer, Dhia Garbaya, Eduard Durech, Ido Hakimi, Juan Garcia Giraldo, Mete Ismayilzada, Negar Foroutan Eghlidi, Skander Moalla, Tiancheng Chen, Vinko Sabolcec, Yixuan Even Xu, Michael Aerni, Badr AlKhamissi, Ines Altemir Marinas, Mohammad Hossein Amani, Matin Ansaripour, Ilia Badanin, Harold Benoit, Emanuela Boros, Nicholas John Browning, Fabian Bösch, Maximilian Böther, Niklas Canova, Camille Challier, Clément Charmillot, Jonathan Coles, Jan Deriu, Arnout Devos, Lukas Drescher, Daniil Dzenhaliou, Maud Ehrmann, Dongyang Fan, Simin Fan, Silin Gao, Miguel Gila, María Grandury, Diba Hashemi, Alexander Miserlis Hoyle, Jiaming Jiang, Mark Klein 0002, Andrei Kucharavy, Anastasiia Kucherenko, Frederike Lübeck, Roman Machacek, Theofilos Ioannis Manitaras, Andreas Marfurt, Kyle Matoba, Simon Matrenok, Henrique Mendonça, Fawzi Roberto Mohamed, Syrielle Montariol, Luca Mouchel, Sven Najem-Meyer, Jingwei Ni, Gennaro Oliva, Matteo Pagliardini, Elia Palme, Andrei Panferov, Léo Paoletti, Marco Passerini, Ivan Pavlov, Auguste Poiroux, Kaustubh Ponkshe, Nathan Ranchin, Javier Rando, Mathieu Sauser, Jakhongir Saydaliev, Mukhammadali Sayfiddinov, Marian Schneider, Stefano Schuppli, Marco Scialanga, Andrei Semenov, Kumar Shridhar, Raghav Singhal, Anna Sotnikova, Alexander Sternfeld, Ayush K. Tarun, Paul Teiletche, Jannis Vamvas, Xiaozhe Yao, Alexander Ilic, Ana Klimovic, Andreas Krause 0001, Caglar Gulcehre, David Rosenthal, Elliott Ash, Florian Tramèr, Joost VandeVondele, Livio Veraldi, Martin Rajman, Thomas C. Schulthess, Torsten Hoefler, Antoine Bosselut, Martin Jaggi, Imanol Schlag |
ACL (1) | 60 |
| 2026 | Efficient Test-Time Scaling of Multi-Step Reasoning by Probing Internal States of Large Language Models
Jingwei Ni, Ekaterina Fadeeva, Mubashara Akhtar, Jiaheng Zhang, Elliott Ash, Markus Leippold, Timothy Baldwin, See-Kiong Ng, Artem Shelmanov, Mrinmaya Sachan |
ACL (1) | 1 |
| 2026 | Tackling the Root of Misinformation by Teaching Laypeople about Logical Fallacies via Socratic Questioning and Critical ArgumentationabstractIdentifying logical fallacies in everyday discourse is challenging for many people.This challenge is amplified in the era of Large Language Models (LLMs), where malicious agents can deploy fallacious arguments to disseminate misinformation at scale.In this work, we explore the potential of LLMs as part of the solution.We introduce LFTutor, an intelligent tutoring system which uses LLMs to tutor laypeople and help them learn about logical fallacies.LFTutor integrates intent-driven Socratic questioning and critical argumentation principles to actively engage learners to reflect on their reasoning.Through both automatic and human evaluations, we demonstrate that LFTutor significantly outperforms baseline LLMs lacking these pedagogical strategies.This work highlights the promise of combining LLMs with pedagogical scaffolding to foster critical thinking and argument literacy in the age of AI.ing university students' ability to recognize argument structures and fallacies. Minjing Shi, Junling Wang 0001, Jingwei Ni, Sankalan Pal Chowdhury, Mrinmaya Sachan |
ACL (1) | 3 |
| 2026 | WeatherArchive: A Benchmark for Retrieval-Augmented Reasoning over Historical Weather ArchivesabstractHistorical news segments on weather events are collections of enduring primary source records that offer rich, untapped narratives of how societies have experienced and responded to extreme weather events. These qualitative accounts provide insights into societal vulnerability and resilience that are largely absent from meteorological records, making them valuable for climate scientists to understand societal responses. However, their large scale, noise in optical character recognition (OCR), and archaic language make it difficult to transform them into structured knowledge for climate research. To address this challenge, we introduce øurmethod, the first large-scale benchmark for evaluating end-to-end retrieval-augmented generation (RAG) systems on historical weather archives. WeatherArchive-Bench comprises two tasks: WeatherArchive-Retrieval, which measures a system's ability to locate historically relevant news segments from over one million archival news segments, and WeatherArchive-Assessment, which evaluates whether Large Language Models (LLMs) can classify societal vulnerability and resilience indicators from extreme weather narratives and answer queries using the segments retrieved. Extensive experiments across sparse, dense, and re-ranking retrievers, as well as a diverse set of LLMs, reveal that dense retrievers often fail on historical terminology, while LLMs frequently misinterpret vulnerability and resilience concepts. These findings highlight key limitations in reasoning about complex societal indicators and provide insights for designing more robust climate-focused RAG systems from archival contexts. The constructed dataset and evaluation framework are available at: https://github.com/Weather-Archival-Rescue/WeatherArchive-Bench. Yongan Yu, Xianda Du, Qingchen Hu, Jingwei Ni, Dan Qiang, Grant McKenzie, Renée Sieber, Fengran Mo |
SIGIR | 5 |
| 2025 | DIRAS: Efficient LLM Annotation of Document Relevance for Retrieval Augmented GenerationabstractJingwei Ni, Tobias Schimanski, Meihong Lin, Mrinmaya Sachan, Elliott Ash, Markus Leippold. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Jingwei Ni, Tobias Schimanski, Meihong Lin, Mrinmaya Sachan, Elliott Ash, Markus Leippold |
NAACL (Long Papers) | 1 |
| 2024 | AFaCTA: Assisting the Annotation of Factual Claim Detection with Reliable LLM AnnotatorsabstractJingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan, Elliott Ash, Markus Leippold. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Jingwei Ni, Minjing Shi, Dominik Stammbach, Mrinmaya Sachan, Elliott Ash, Markus Leippold |
ACL (1) | 1 |
| 2024 | Towards Faithful and Robust LLM Specialists for Evidence-Based Question-AnsweringabstractAdvances towards more faithful and traceable answers of Large Language Models (LLMs) are crucial for various research and practical endeavors.One avenue in reaching this goal is basing the answers on reliable sources.However, this Evidence-Based QA has proven to work insufficiently with LLMs in terms of citing the correct sources (source quality) and truthfully representing the information within sources (answer attributability).In this work, we systematically investigate how to robustly fine-tune LLMs for better source quality and answer attributability.Specifically, we introduce a data generation pipeline with automated data quality filters, which can synthesize diversified high-quality training and testing data at scale.We further introduce four test sets to benchmark the robustness of fine-tuned specialist models.Extensive evaluation shows that fine-tuning on synthetic data improves performance on both in-and out-of-distribution.Furthermore, we show that data quality, which can be drastically improved by proposed quality filters, matters more than quantity in improving Evidence-Based QA. Tobias Schimanski, Jingwei Ni, Mathias Kraus, Elliott Ash, Markus Leippold |
ACL (1) | 2 |
| 2024 | ClimRetrieve: A Benchmarking Dataset for Information Retrieval from Corporate Climate DisclosuresabstractTo handle the vast amounts of qualitative data produced in corporate climate communication, stakeholders increasingly rely on Retrieval Augmented Generation (RAG) systems.However, a significant gap remains in evaluating domain-specific information retrieval -the basis for answer generation.To address this challenge, this work simulates the typical tasks of a sustainability analyst by examining 30 sustainability reports with 16 detailed climate-related questions.As a result, we obtain a dataset with over 8.5K unique question-source-answer pairs labeled by different levels of relevance.Furthermore, we develop a use case with the dataset to investigate the integration of expert knowledge into information retrieval with embeddings.Although we show that incorporating expert knowledge works, we also outline the critical limitations of embeddings in knowledge-intensive downstream domains like climate change communication.121 All the data and code for this project is available on https://github.com/tobischimanski/ClimRetrieve.2 We thank the expert annotators Aysha Emmerson, Emily Hsu, and Capucine Le Meur for their work on this project.3 For example, companies must describe the processes they use to identify, assess, and manage these risks and opportuni- Tobias Schimanski, Jingwei Ni, Roberto Martín, Nicola Ranger, Markus Leippold |
EMNLP | 2 |
| 2023 | When Does Aggregating Multiple Skills with Multi-Task Learning Work? A Case Study in Financial NLPabstractMulti-task learning (MTL) aims at achieving a better model by leveraging data and knowledge from multiple tasks.However, MTL does not always work -sometimes negative transfer occurs between tasks, especially when aggregating loosely related skills, leaving it an open question when MTL works.Previous studies show that MTL performance can be improved by algorithmic tricks.However, what tasks and skills should be included is less well explored.In this work, we conduct a case study in Financial NLP where multiple datasets exist for skills relevant to the domain, such as numeric reasoning and sentiment analysis.Due to the task difficulty and data scarcity in the Financial NLP domain, we explore when aggregating such diverse skills from multiple datasets with MTL can work.Our findings suggest that the key to MTL success lies in skill diversity, relatedness between tasks, and choice of aggregation size and shared capacity.Specifically, MTL works well when tasks are diverse but related, and when the size of the task aggregation and the shared capacity of the model are balanced to avoid overwhelming certain tasks. 1 Jingwei Ni, Zhijing Jin 0001, Mrinmaya Sachan, Markus Leippold |
ACL (1) | 1 |
| 2022 | Original or Translated? A Causal Analysis of the Impact of Translationese on Machine Translation PerformanceabstractJingwei Ni, Zhijing Jin, Markus Freitag, Mrinmaya Sachan, Bernhard Schölkopf. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Jingwei Ni, Zhijing Jin 0001, Markus Freitag, Mrinmaya Sachan, Bernhard Schölkopf |
NAACL-HLT | 1 |
| 2022 | Teachers cooperation: team-knowledge distillation for multiple cross-domain few-shot learning
Zhong Ji, Jingwei Ni, Xiyao Liu 0002, Yanwei Pang |
Frontiers Comput. Sci. | 2 |
| 2021 | Causal Direction of Data Collection Matters: Implications of Causal and Anticausal Learning for NLPabstractZhijing Jin, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, Bernhard Schoelkopf. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021. Zhijing Jin 0001, Julius von Kügelgen, Jingwei Ni, Tejas Vaidhya, Ayush Kaushal, Mrinmaya Sachan, Bernhard Schölkopf |
EMNLP (1) | 3 |