EDBT 2026 Demo / reviewers in the wild / expert
Ruoxi Ning
dblp:320/5732
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Language models and text generation · 62% Question answering and dialogue systems · 29% Information extraction and text analysis · 9% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › large language model training
language model pretraining |
1.0 | 1 | 2026 | On the Effect of Hyperparameters in Language Modeling for Computational Linguistics · ACL (1) 2026 |
Empirical software engineering
reproducibility |
1.0 | 1 | 2026 | On the Effect of Hyperparameters in Language Modeling for Computational Linguistics · ACL (1) 2026 |
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
long-context question answering |
0.9 | 1 | 2025 | NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens · ICLR 2025 |
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context understanding |
0.9 | 1 | 2025 | NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
replication study · 2.0multi-hop reasoning evaluation · 0.9manual annotation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | On the Effect of Hyperparameters in Language Modeling for Computational LinguisticsabstractTraining language models and examining their linguistic behaviors have been a common protocol in computational linguistics for studying linguistic phenomena and modeling human language processing.However, work in this area is often limited to proof-of-concept demonstrations with arbitrary model configurations, without considering hyperparameter sensitivity, an important source of variation in model performance.In this work, we replicate three prior studies (Chang and Bergen, 2022; Hu et al., 2020b;Kuribayashi et al., 2024) with hyperparameters varied within a practical range, and show that modest hyperparameter changes can alter some qualitative conclusions about models' linguistic abilities and even reverse the ranking of model performance.Our results highlight the risk that prior work may have reflected optimization artifacts rather than the genuine inductive biases of model classes, and that hyperparameter sensitivity should receive more attention as a factor that can meaningfully influence model behavior.We suggest future work to report the variation of performance across the configuration space to enhance the reliability and generalizability of conclusions. Ruoxi Ning, Yongpeng Zhu, Qingcheng Zeng, Tatsuki Kuribayashi, Freda Shi |
ACL (1) | 1 |
| 2025 | NovelQA: Benchmarking Question Answering on Documents Exceeding 200K TokensabstractRecent advancements in Large Language Models (LLMs) have pushed the boundaries of natural language processing, especially in long-context understanding. However, the evaluation of these models' long-context abilities remains a challenge due to the limitations of current benchmarks. To address this gap, we introduce NovelQA, a benchmark tailored for evaluating LLMs with complex, extended narratives. NovelQA, constructed from English novels, offers a unique blend of complexity, length, and narrative coherence, making it an ideal tool for assessing deep textual understanding in LLMs. This paper details the design and construction of NovelQA, focusing on its comprehensive manual annotation process and the variety of question types aimed at evaluating nuanced comprehension. Our evaluation of long-context LLMs on NovelQA reveals significant insights into their strengths and weaknesses. Notably, the models struggle with multi-hop reasoning, detail-oriented questions, and handling extremely long inputs, averaging over 200,000 tokens. Results highlight the need for substantial advancements in LLMs to enhance their long-context comprehension and contribute effectively to computational literary analysis. Cunxiang Wang, Ruoxi Ning, Boqi Pan, Tonghui Wu, Qipeng Guo, Cheng Deng 0001, Guangsheng Bao, Xiangkun Hu, Zheng Zhang 0001, Yue Zhang 0004 |
ICLR | 2 |