Nickvash Kani

dblp:153/6088 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-5223-5069ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Integrating Arithmetic Learning Improves Mathematical Reasoning in Smaller Models
abstract
While large models pre-trained on high-quality data exhibit excellent performance on mathematical reasoning (e.g., GSM8k, MultiArith), it remains challenging to specialize smaller models for these tasks. Common approaches to address this challenge include knowledge distillation from large teacher models and data augmentation (e.g., rephrasing questions and generating synthetic solutions). Despite these efforts, smaller models struggle with arithmetic computations, leading to errors in mathematical reasoning. In this work, we leverage a synthetic arithmetic dataset generated programmatically to enhance the reasoning capabilities of smaller models. We investigate two key approaches to incorporate this dataset: (1) intermediate fine-tuning, in which a model is fine-tuned on the arithmetic dataset before training it on a reasoning dataset, and (2) integrating the arithmetic dataset into an instruction-tuning mixture, allowing the model to learn arithmetic skills alongside general instruction-following abilities. Our experiments on multiple reasoning benchmarks demonstrate that incorporating an arithmetic dataset, whether through targeted fine-tuning or within an instruction-tuning mixture, enhances models' arithmetic capabilities, thereby improving their mathematical reasoning performance.
Neeraj Gangwar, Suma P. Bhat, Nickvash Kani
LREC3
2025 Advancing Math Formula Search Using Diverse Structural and Symbolic Representations
Sumedh Vemuganti, Ayu Seiya, Nickvash Kani
ECIR (1)3
2025 E-Gen: Leveraging E-Graphs to Improve Continuous Representations of Symbolic Expressions
abstract
Hongbo Zheng, Suyuan Wang, Neeraj Gangwar, Nickvash Kani. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Hongbo Zheng, Suyuan Wang, Neeraj Gangwar, Nickvash Kani
NAACL (Long Papers)4
2024 Factors That Influence Automatic Recognition of African-American Vernacular English in Machine-Learning Models
abstract
Racial bias is a well-documented problem in natural language processing (NLP). The dialectal language used by marginalized groups is often misclassified or mischaracterized by language models, which in turn can further disenfranchise these populations. Previous works have noted that some popular language identification (LID) models perform worse when classifying tweets that contain African-American Vernacular English (AAVE) than when classifying tweets that contain White-Aligned English (WAE). This work examines the factors that contribute to racial bias in language models for the LID task. The contributions of this work are two-fold. First, a thorough analysis demonstrates that a lack of “unique” language-specific n-gram features in an LID model can lead to poor performance on dialectal data, especially on shorter-length inputs like those typically found on social media. Second, based on these findings, this work introduces and illustrates the efficacy of two simple yet accurate solutions: i.) mining “unique” n-gram features and ii.) including examples of dialectal English in training data. These solutions mitigate the accuracy gap between WAE and AAVE which some language identification models demonstrate when classifying shorter inputs. Mining for unique features and training with a more diverse dataset can improve the disparity on short-length sequences by 6% and 9.8% respectively.
Emma Hamel, Nickvash Kani
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Highlighting Named Entities in Input for Auto-formulation of Optimization Problems
Neeraj Gangwar, Nickvash Kani
CICM2
2022 An Evaluation of NLP Methods to Extract Mathematical Token Descriptors
Emma Hamel, Hongbo Zheng, Nickvash Kani
CICM3