VLDB 2026 Research / reviewers in the wild / expert
Nickvash Kani
dblp:153/6088
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0000-0002-5223-5069ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Theory of computation · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Integrating Arithmetic Learning Improves Mathematical Reasoning in Smaller ModelsabstractWhile large models pre-trained on high-quality data exhibit excellent performance on mathematical reasoning (e.g., GSM8k, MultiArith), it remains challenging to specialize smaller models for these tasks. Common approaches to address this challenge include knowledge distillation from large teacher models and data augmentation (e.g., rephrasing questions and generating synthetic solutions). Despite these efforts, smaller models struggle with arithmetic computations, leading to errors in mathematical reasoning. In this work, we leverage a synthetic arithmetic dataset generated programmatically to enhance the reasoning capabilities of smaller models. We investigate two key approaches to incorporate this dataset: (1) intermediate fine-tuning, in which a model is fine-tuned on the arithmetic dataset before training it on a reasoning dataset, and (2) integrating the arithmetic dataset into an instruction-tuning mixture, allowing the model to learn arithmetic skills alongside general instruction-following abilities. Our experiments on multiple reasoning benchmarks demonstrate that incorporating an arithmetic dataset, whether through targeted fine-tuning or within an instruction-tuning mixture, enhances models' arithmetic capabilities, thereby improving their mathematical reasoning performance. Neeraj Gangwar, Suma P. Bhat, Nickvash Kani |
LREC | 3 |
| 2025 | Advancing Math Formula Search Using Diverse Structural and Symbolic Representations
Sumedh Vemuganti, Ayu Seiya, Nickvash Kani |
ECIR (1) | 3 |
| 2025 | E-Gen: Leveraging E-Graphs to Improve Continuous Representations of Symbolic ExpressionsabstractHongbo Zheng, Suyuan Wang, Neeraj Gangwar, Nickvash Kani. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hongbo Zheng, Suyuan Wang, Neeraj Gangwar, Nickvash Kani |
NAACL (Long Papers) | 4 |
| 2024 | Factors That Influence Automatic Recognition of African-American Vernacular English in Machine-Learning ModelsabstractRacial bias is a well-documented problem in natural language processing (NLP). The dialectal language used by marginalized groups is often misclassified or mischaracterized by language models, which in turn can further disenfranchise these populations. Previous works have noted that some popular language identification (LID) models perform worse when classifying tweets that contain African-American Vernacular English (AAVE) than when classifying tweets that contain White-Aligned English (WAE). This work examines the factors that contribute to racial bias in language models for the LID task. The contributions of this work are two-fold. First, a thorough analysis demonstrates that a lack of “unique” language-specific n-gram features in an LID model can lead to poor performance on dialectal data, especially on shorter-length inputs like those typically found on social media. Second, based on these findings, this work introduces and illustrates the efficacy of two simple yet accurate solutions: i.) mining “unique” n-gram features and ii.) including examples of dialectal English in training data. These solutions mitigate the accuracy gap between WAE and AAVE which some language identification models demonstrate when classifying shorter inputs. Mining for unique features and training with a more diverse dataset can improve the disparity on short-length sequences by 6% and 9.8% respectively. Emma Hamel, Nickvash Kani |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Highlighting Named Entities in Input for Auto-formulation of Optimization Problems
Neeraj Gangwar, Nickvash Kani |
CICM | 2 |
| 2022 | An Evaluation of NLP Methods to Extract Mathematical Token Descriptors
Emma Hamel, Hongbo Zheng, Nickvash Kani |
CICM | 3 |