Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jabo Serge Byusa

dblp:395/6488 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
1 paper
Mathematical optimization · 100%
Artificial intelligence
1 paper
Question answering and dialogue systems · 77% Language models and text generation · 23%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
domain-specific question answering
0.912025
Evaluating LLM Reasoning in the Operations Research Domain with ORQA · AAAI 2025
Mathematical optimization › optimization modeling
LLM-based optimization modeling
0.912025
Evaluating LLM Reasoning in the Operations Research Domain with ORQA · AAAI 2025
Mathematical optimization
optimization modeling
0.912025
Evaluating LLM Reasoning in the Operations Research Domain with ORQA · AAAI 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.312025
Evaluating LLM Reasoning in the Operations Research Domain with ORQA · AAAI 2025

Methods — techniques the papers use, named apart from their topics

expert annotation · 1.7benchmark construction · 1.7
YearPublicationVenuePosition
2025 Evaluating LLM Reasoning in the Operations Research Domain with ORQA
abstract
In this paper, we introduce and apply Operations Research Question Answering (ORQA), a new benchmark, to assess the generalization capabilities of Large Language Models (LLMs) in the specialized technical domain of Operations Research (OR). This benchmark is designed to evaluate whether LLMs can emulate the knowledge and reasoning skills of OR experts when given diverse and complex optimization problems. The dataset, crafted by OR experts, presents real-world optimization problems that require multistep reasoning to build their mathematical models. Our evaluations of various open-source LLMs, such as LLaMA 3.1, DeepSeek, and Mixtral reveal their modest performance, indicating a gap in their aptitude to generalize to specialized technical domains. This work contributes to the ongoing discourse on LLMs’ generalization capabilities, providing insights for future research in this area. The dataset and evaluation code are publicly available.
Mahdi Mostajabdaveh, Timothy T. L. Yu, Samarendra Chandan Bindu Dash, Rindranirina Ramamonjison, Jabo Serge Byusa, Giuseppe Carenini, Zirui Zhou, Yong Zhang 0004
AAAI5