Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Ivan Moshkov

dblp:368/8197 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 79% Generative modeling · 21%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
instruction tuning
1.622025
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data · ICLR 2025
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset · NeurIPS 2024
Natural language and speech › Language models and text generation
mathematical reasoning
1.622025
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data · ICLR 2025
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset · NeurIPS 2024
Machine learning › Generative modeling
synthetic data generation
1.122025
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data · ICLR 2025
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.912025
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data · ICLR 2025

Methods — techniques the papers use, named apart from their topics

large language model · 0.9data ablation · 0.9prompting · 0.8model distillation · 0.8code-interpreter solutions · 0.8
YearPublicationVenuePosition
2025 OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
abstract
Mathematical reasoning continues to be a critical challenge in large language model (LLM) development with significant interest. However, most of the cutting-edge progress in mathematical reasoning with LLMs has become closed-source due to lack of access to training data. This lack of data access limits researchers from understanding the impact of different choices for synthesizing and utilizing the data. With the goal of creating a high-quality finetuning (SFT) dataset for math reasoning, we conduct careful ablation experiments on data synthesis using the recently released Llama3.1 family of models. Our experiments show that: (a) solution format matters, with excessively verbose solutions proving detrimental to SFT performance, (b) data generated by a strong teacher outperforms on-policy data generated by a weak student model, (c) SFT is robust to low-quality solutions, allowing for imprecise data filtering, and (d) question diversity is crucial for achieving data scaling gains. Based on these insights, we create the OpenMathInstruct-2 dataset which consists of 14M question-solution pairs (≈ 600K unique questions), making it nearly eight times larger than the previous largest open-source math reasoning dataset. Finetuning the Llama-3.1-8B-Base using OpenMathInstruct-2 outperforms Llama3.1-8B-Instruct on MATH by an absolute 15.9% (51.9% → 67.8%). Finally, to accelerate the open-source efforts, we release the code, the finetuned models, and the OpenMathInstruct-2 dataset under a commercially permissive license.
Shubham Toshniwal, Ivan Moshkov, Branislav Kisacanin, Alexan Ayrapetyan, Igor Gitman
ICLR3
2024 OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
abstract
Recent work has shown the immense potential of synthetically generated datasets for training large language models (LLMs), especially for acquiring targeted skills. Current large-scale math instruction tuning datasets such as MetaMathQA (Yu et al., 2024) and MAmmoTH (Yue et al., 2024) are constructed using outputs from closed-source LLMs with commercially restrictive licenses. A key reason limiting the use of open-source LLMs in these data generation pipelines has been the wide gap between the mathematical skills of the best closed-source LLMs, such as GPT-4, and the best open-source LLMs. Building on the recent progress in open-source LLMs, our proposed prompting novelty, and some brute-force scaling, we construct OpenMathInstruct-1, a math instruction tuning dataset with 1.8M problem-solution pairs. The dataset is constructed by synthesizing code-interpreter solutions for GSM8K and MATH, two popular math reasoning benchmarks, using the recently released and permissively licensed Mixtral model. Our best model, OpenMath-CodeLlama-70B, trained on a subset of OpenMathInstruct-1, achieves a score of 84.6% on GSM8K and 50.7% on MATH, which is competitive with the best gpt-distilled models. We will release our code, models, and the OpenMathInstruct-1 dataset under a commercially permissive license.
Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, Igor Gitman
NeurIPS2