VLDB 2026 Research / reviewers in the wild / expert
Jabo Serge Byusa
dblp:395/6488
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Artificial intelligence
1 paper |
Question answering and dialogue systems · 77% Language models and text generation · 23% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems
domain-specific question answering |
0.9 | 1 | 2025 | Evaluating LLM Reasoning in the Operations Research Domain with ORQA · AAAI 2025 |
Mathematical optimization › optimization modeling
LLM-based optimization modeling |
0.9 | 1 | 2025 | Evaluating LLM Reasoning in the Operations Research Domain with ORQA · AAAI 2025 |
Mathematical optimization
optimization modeling |
0.9 | 1 | 2025 | Evaluating LLM Reasoning in the Operations Research Domain with ORQA · AAAI 2025 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2025 | Evaluating LLM Reasoning in the Operations Research Domain with ORQA · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
expert annotation · 1.7benchmark construction · 1.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Evaluating LLM Reasoning in the Operations Research Domain with ORQAabstractIn this paper, we introduce and apply Operations Research Question Answering (ORQA), a new benchmark, to assess the generalization capabilities of Large Language Models (LLMs) in the specialized technical domain of Operations Research (OR). This benchmark is designed to evaluate whether LLMs can emulate the knowledge and reasoning skills of OR experts when given diverse and complex optimization problems. The dataset, crafted by OR experts, presents real-world optimization problems that require multistep reasoning to build their mathematical models. Our evaluations of various open-source LLMs, such as LLaMA 3.1, DeepSeek, and Mixtral reveal their modest performance, indicating a gap in their aptitude to generalize to specialized technical domains. This work contributes to the ongoing discourse on LLMs’ generalization capabilities, providing insights for future research in this area. The dataset and evaluation code are publicly available. Mahdi Mostajabdaveh, Timothy T. L. Yu, Samarendra Chandan Bindu Dash, Rindranirina Ramamonjison, Jabo Serge Byusa, Giuseppe Carenini, Zirui Zhou, Yong Zhang 0004 |
AAAI | 5 |