VLDB 2026 Research / reviewers in the wild / expert
Anubhav Shrestha
dblp:408/3297
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 61% Reinforcement learning · 30% Transfer learning and domain adaptation · 9% |
Topics — the 4 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › chain-of-thought reasoning
chain-of-thought distillation |
0.9 | 1 | 2025 | Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings · EMNLP 2025 |
Natural language and speech › Language models and text generation › large language model
reasoning model |
0.9 | 1 | 2025 | Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings · EMNLP 2025 |
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
0.9 | 1 | 2025 | Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings · EMNLP 2025 |
Machine learning › Transfer learning and domain adaptation
domain generalization |
0.3 | 1 | 2025 | Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained Settings · EMNLP 2025 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.9distillation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Warm Up Before You Train: Unlocking General Reasoning in Resource-Constrained SettingsabstractDesigning effective reasoning-capable LLMs typically requires training using Reinforcement Learning with Verifiable Rewards (RLVR) or distillation with carefully curated Long Chain of Thoughts (CoT), both of which depend heavily on extensive training data.This creates a major challenge when the amount of quality training data is scarce.We propose a sampleefficient, two-stage training strategy to develop reasoning LLMs under limited supervision.In the first stage, we "warm up" the model by distilling Long CoTs from a toy domain, namely, Knights & Knaves (K&K) logic puzzles to acquire general reasoning skills.In the second stage, we apply RLVR to the warmed-up model using a limited set of target-domain examples.Our experiments demonstrate that this two-phase approach offers several benefits: (i) the warmup phase alone facilitates generalized reasoning, leading to performance improvements across a range of tasks, including MATH, HumanEval + , and MMLU-Pro; (ii) When both the base model and the warmed-up model are RLVR trained on the same small dataset (≤ 100 examples), the warmed-up model consistently outperforms the base model; (iii) Warming up before RLVR training allows a model to maintain cross-domain generalizability even after training on a specific domain; (iv) Introducing warmup in the pipeline improves not only accuracy but also overall sample efficiency during RLVR training.The results in this paper highlight the promise of warmup for building robust reasoning LLMs in data-scarce environments. Safal Shrestha, Minwu Kim, Aadim Nepal, Anubhav Shrestha, Keith W. Ross |
EMNLP | 4 |