Shubham Toshniwal

dblp:160/4302 · DBLP profile ↗
← Back
13ranked-venue papers
9as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 54% Information extraction and text analysis · 28% Deep learning architectures and training · 9%

Topics — the 17 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
coreference resolution
1.632024
Major Entity Identification: A Generalizable Alternative to Coreference Resolution · EMNLP 2024
Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks · EMNLP (1) 2020
PeTra: A Sparsely Supervised Memory Model for People Tracking · ACL 2020
Natural language and speech › Language models and text generation
instruction tuning
1.622025
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data · ICLR 2025
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset · NeurIPS 2024
Natural language and speech › Language models and text generation
mathematical reasoning
1.622025
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data · ICLR 2025
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset · NeurIPS 2024
Machine learning › Generative modeling
synthetic data generation
1.122025
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data · ICLR 2025
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset · NeurIPS 2024
Natural language and speech › Language models and text generation › large language model › large language model adaptation
supervised fine-tuning
0.912025
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data · ICLR 2025
Natural language and speech › Information extraction and text analysis
named entity recognition
0.812024
Major Entity Identification: A Generalizable Alternative to Coreference Resolution · EMNLP 2024
Natural language and speech › Language models and text generation
chain-of-thought reasoning
0.712023
Learning to Reason and Memorize with Self-Notes · NeurIPS 2023
Natural language and speech › Language models and text generation › large language model reasoning
multi-step reasoning
0.712023
Learning to Reason and Memorize with Self-Notes · NeurIPS 2023
Natural language and speech › Language models and text generation
large language model evaluation
0.612022
Chess as a Testbed for Language Model State Tracking · AAAI 2022
Machine learning › Deep learning architectures and training › sequence modeling
state tracking
0.612022
Chess as a Testbed for Language Model State Tracking · AAAI 2022
Natural language and speech › Language models and text generation › language modeling › language model architecture
transformer language model
0.612022
Chess as a Testbed for Language Model State Tracking · AAAI 2022
Machine learning › Deep learning architectures and training
memory-augmented neural networks
0.622020
Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks · EMNLP (1) 2020
PeTra: A Sparsely Supervised Memory Model for People Tracking · ACL 2020
Natural language and speech › Information extraction and text analysis
entity tracking
0.412020
PeTra: A Sparsely Supervised Memory Model for People Tracking · ACL 2020
Natural language and speech › Information extraction and text analysis › coreference resolution
long document coreference resolution
0.412020
Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis › coreference resolution
pronoun resolution
0.412020
PeTra: A Sparsely Supervised Memory Model for People Tracking · ACL 2020
Natural language and speech › Language models and text generation › in-context learning
few-shot prompting
0.212024
Major Entity Identification: A Generalizable Alternative to Coreference Resolution · EMNLP 2024
Natural language and speech › Language models and text generation
memory augmentation
0.212023
Learning to Reason and Memorize with Self-Notes · NeurIPS 2023

Methods — techniques the papers use, named apart from their topics

memory-augmented neural network · 0.9large language model · 0.9data ablation · 0.9supervised classification · 0.8prompting · 0.8model distillation · 0.8large language model prompting · 0.8code-interpreter solutions · 0.8scratchpad · 0.7chain-of-thought · 0.7
YearPublicationVenuePosition
2025 OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
abstract
Mathematical reasoning continues to be a critical challenge in large language model (LLM) development with significant interest. However, most of the cutting-edge progress in mathematical reasoning with LLMs has become closed-source due to lack of access to training data. This lack of data access limits researchers from understanding the impact of different choices for synthesizing and utilizing the data. With the goal of creating a high-quality finetuning (SFT) dataset for math reasoning, we conduct careful ablation experiments on data synthesis using the recently released Llama3.1 family of models. Our experiments show that: (a) solution format matters, with excessively verbose solutions proving detrimental to SFT performance, (b) data generated by a strong teacher outperforms on-policy data generated by a weak student model, (c) SFT is robust to low-quality solutions, allowing for imprecise data filtering, and (d) question diversity is crucial for achieving data scaling gains. Based on these insights, we create the OpenMathInstruct-2 dataset which consists of 14M question-solution pairs (≈ 600K unique questions), making it nearly eight times larger than the previous largest open-source math reasoning dataset. Finetuning the Llama-3.1-8B-Base using OpenMathInstruct-2 outperforms Llama3.1-8B-Instruct on MATH by an absolute 15.9% (51.9% → 67.8%). Finally, to accelerate the open-source efforts, we release the code, the finetuned models, and the OpenMathInstruct-2 dataset under a commercially permissive license.
Shubham Toshniwal, Ivan Moshkov, Branislav Kisacanin, Alexan Ayrapetyan, Igor Gitman
ICLR1
2024 Major Entity Identification: A Generalizable Alternative to Coreference Resolution
abstract
The limited generalization of coreference resolution (CR) models has been a major bottleneck in the task's broad application.Prior work has identified annotation differences, especially for mention detection, as one of the main reasons for the generalization gap and proposed using additional annotated target domain data.Rather than relying on this additional annotation, we propose an alternative referential task, Major Entity Identification (MEI), where we: (a) assume the target entities to be specified in the input, and (b) limit the task to only the frequent entities.Through extensive experiments, we demonstrate that MEI models generalize well across domains on multiple datasets with supervised models and LLM-based few-shot prompting.Additionally, MEI fits the classification framework, which enables the use of robust and intuitive classification-based metrics.Finally, MEI is also of practical use as it allows a user to search for all mentions of a particular entity or a group of entities of interest.
Kawshik Sundar, Shubham Toshniwal, Makarand Tapaswi, Vineet Gandhi
EMNLP2
2024 OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset
abstract
Recent work has shown the immense potential of synthetically generated datasets for training large language models (LLMs), especially for acquiring targeted skills. Current large-scale math instruction tuning datasets such as MetaMathQA (Yu et al., 2024) and MAmmoTH (Yue et al., 2024) are constructed using outputs from closed-source LLMs with commercially restrictive licenses. A key reason limiting the use of open-source LLMs in these data generation pipelines has been the wide gap between the mathematical skills of the best closed-source LLMs, such as GPT-4, and the best open-source LLMs. Building on the recent progress in open-source LLMs, our proposed prompting novelty, and some brute-force scaling, we construct OpenMathInstruct-1, a math instruction tuning dataset with 1.8M problem-solution pairs. The dataset is constructed by synthesizing code-interpreter solutions for GSM8K and MATH, two popular math reasoning benchmarks, using the recently released and permissively licensed Mixtral model. Our best model, OpenMath-CodeLlama-70B, trained on a subset of OpenMathInstruct-1, achieves a score of 84.6% on GSM8K and 50.7% on MATH, which is competitive with the best gpt-distilled models. We will release our code, models, and the OpenMathInstruct-1 dataset under a commercially permissive license.
Shubham Toshniwal, Ivan Moshkov, Sean Narenthiran, Daria Gitman, Fei Jia, Igor Gitman
NeurIPS1
2023 Learning to Reason and Memorize with Self-Notes
abstract
Large language models have been shown to struggle with multi-step reasoning, and do not retain previous reasoning steps for future use. We propose a simple method for solving both of these problems by allowing the model to take Self-Notes. Unlike recent chain-of-thought or scratchpad approaches, the model can deviate from the input context at any time to explicitly think and write down its thoughts. This allows the model to perform reasoning on the fly as it reads the context and even integrate previous reasoning steps, thus enhancing its memory with useful information and enabling multi-step reasoning. Experiments across a wide variety of tasks demonstrate that our method can outperform chain-of-thought and scratchpad methods by taking Self-Notes that interleave the input text.
Jack Lanchantin, Shubham Toshniwal, Jason Weston, Arthur Szlam, Sainbayar Sukhbaatar
NeurIPS2
2022 Chess as a Testbed for Language Model State Tracking
abstract
Transformer language models have made tremendous strides in natural language understanding tasks. However, the complexity of natural language makes it challenging to ascertain how accurately these models are tracking the world state underlying the text. Motivated by this issue, we consider the task of language modeling for the game of chess. Unlike natural language, chess notations describe a simple, constrained, and deterministic domain. Moreover, we observe that the appropriate choice of chess notation allows for directly probing the world state, without requiring any additional probing-related machinery. We find that: (a) With enough training data, transformer language models can learn to track pieces and predict legal moves with high accuracy when trained solely on move sequences. (b) For small training sets providing access to board state information during training can yield significant improvements. (c) The success of transformer language models is dependent on access to the entire game history i.e. “full attention”. Approximating this full attention results in a significant performance drop. We propose this testbed as a benchmark for future work on the development and analysis of transformer language models.
Shubham Toshniwal, Sam Wiseman, Karen Livescu, Kevin Gimpel
AAAI1
2020 PeTra: A Sparsely Supervised Memory Model for People Tracking
abstract
We propose PeTra, a memory-augmented neural network designed to track entities in its memory slots.PeTra is trained using sparse annotation from the GAP pronoun resolution dataset and outperforms a prior memory model on the task while using a simpler architecture.We empirically compare key modeling choices, finding that we can simplify several aspects of the design of the memory module while retaining strong performance.To measure the people tracking capability of memory models, we (a) propose a new diagnostic evaluation based on counting the number of unique entities in text, and (b) conduct a small scale human evaluation to compare evidence of people tracking in the memory logs of PeTra relative to a previous approach.PeTra is highly effective in both evaluations, demonstrating its ability to track people in its memory despite being trained with limited annotation.
Shubham Toshniwal, Allyson Ettinger, Kevin Gimpel, Karen Livescu
ACL1
2020 Learning to Ignore: Long Document Coreference with Bounded Memory Neural Networks
abstract
Long document coreference resolution remains a challenging task due to the large memory and runtime requirements of current models.Recent work doing incremental coreference resolution using just the global representation of entities shows practical benefits but requires keeping all entities in memory, which can be impractical for long documents.We argue that keeping all entities in memory is unnecessary, and we propose a memoryaugmented neural network that tracks only a small bounded number of entities at a time, thus guaranteeing a linear runtime in length of document.We show that (a) the model remains competitive with models with high memory and computational requirements on OntoNotes and LitBank, and (b) the model learns an efficient memory management strategy easily outperforming a rule-based strategy.
Shubham Toshniwal, Sam Wiseman, Allyson Ettinger, Karen Livescu, Kevin Gimpel
EMNLP (1)1
2019 Pre-Trained Text Embeddings for Enhanced Text-to-Speech Synthesis
Tomoki Hayashi, Shinji Watanabe 0001, Tomoki Toda, Kazuya Takeda, Shubham Toshniwal, Karen Livescu
INTERSPEECH5
2018 Multilingual Speech Recognition with a Single End-to-End Model
abstract
Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models are well suited for multilingual ASR because they encapsulate an acoustic, pronunciation and language model jointly in a single network. In this work we present a single sequence-to-sequence ASR model trained on 9 different Indian languages, which have very little overlap in their scripts. Specifically, we take a union of language-specific grapheme sets and train a grapheme-based sequence-to-sequence model jointly on data from all languages. We find that this model, which is not explicitly given any information about language identity, improves recognition performance by 21% relative compared to analogous sequence-to-sequence models trained on each language individually. By modifying the model to accept a language identifier as an additional input feature, we further improve performance by an additional 7% relative and eliminate confusion between different languages.
Shubham Toshniwal, Tara N. Sainath, Ron J. Weiss, Bo Li 0028, Pedro J. Moreno 0001, Eugene Weinstein, Kanishka Rao
ICASSP1
2018 Parsing Speech: a Neural Approach to Integrating Lexical and Acoustic-Prosodic Information
abstract
Trang Tran, Shubham Toshniwal, Mohit Bansal, Kevin Gimpel, Karen Livescu, Mari Ostendorf. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Trang Tran 0001, Shubham Toshniwal, Mohit Bansal, Kevin Gimpel, Karen Livescu, Mari Ostendorf
NAACL-HLT2
2018 A Comparison of Techniques for Language Model Integration in Encoder-Decoder Speech Recognition
abstract
Attention-based recurrent neural encoder-decoder models present an elegant solution to the automatic speech recognition problem. This approach folds the acoustic model, pronunciation model, and language model into a single network and requires only a parallel corpus of speech and text for training. However, unlike in conventional approaches that combine separate acoustic and language models, it is not clear how to use additional (unpaired) text. While there has been previous work on methods addressing this problem, a thorough comparison among methods is still lacking. In this paper, we compare a suite of past methods and some of our own proposed methods for using unpaired text data to improve encoder-decoder models. For evaluation, we use the medium-sized Switchboard data set and the large-scale Google voice search and dictation data sets. Our results confirm the benefits of using unpaired text across a range of methods and data sets. Surprisingly, for first-pass decoding, the rather simple approach of shallow fusion performs best across data sets. However, for Google data sets we find that cold fusion has a lower oracle error rate and outperforms other approaches after second-pass rescoring on the Google voice search data set.
Shubham Toshniwal, Anjuli Kannan, Chung-Cheng Chiu, Tara N. Sainath, Karen Livescu
SLT1
2017 Multitask Learning with Low-Level Auxiliary Tasks for Encoder-Decoder Based Speech Recognition
abstract
End-to-end training of deep learning-based models allows for implicit learning of intermediate representations based on the final task loss. However, the end-to-end approach ignores the useful domain knowledge encoded in explicit intermediate-level supervision. We hypothesize that using intermediate representations as auxiliary supervision at lower levels of deep networks may be a good way of combining the advantages of end-to-end training and more traditional pipeline approaches. We present experiments on conversational speech recognition where we use lower-level tasks, such as phoneme recognition, in a multitask training approach with an encoder-decoder model for direct character transcription. We compare multiple types of lower-level tasks and analyze the effects of the auxiliary tasks. Our results on the Switchboard corpus show that this approach improves recognition accuracy over a standard encoder-decoder model on the Eval2000 test set.
Shubham Toshniwal, Hao Tang 0002, Liang Lu 0001, Karen Livescu
INTERSPEECH1
2016 Jointly learning to align and convert graphemes to phonemes with neural attention models
abstract
We propose an attention-enabled encoder-decoder model for the problem of grapheme-to-phoneme conversion. Most previous work has tackled the problem via joint sequence models that require explicit alignments for training. In contrast, the attention-enabled encoder-decoder model allows for jointly learning to align and convert characters to phonemes. We explore different types of attention models, including global and local attention, and our best models achieve state-of-the-art results on three standard data sets (CMU-Dict, Pronlex, and NetTalk).
Shubham Toshniwal, Karen Livescu
SLT1