Shyam Upadhyay

dblp:161/0014 · DBLP profile ↗
← Back
22ranked-venue papers
8as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
12 papers
Language models and text generation · 39% Information extraction and text analysis · 32% Knowledge representation and reasoning · 8%
Databases, data mining, and information retrieval
1 paper
Machine learning and data management · 100%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model reasoning
efficient reasoning
1.012026
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model reasoning
overthinking
1.012026
Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking · ACL (1) 2026
Natural language and speech › Language models and text generation
model routing
0.812024
AutoMix: Automatically Mixing Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning › verification
self-verification
0.812024
AutoMix: Automatically Mixing Language Models · NeurIPS 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › temporal reasoning
temporal commonsense reasoning
0.512021
TIMEDIAL: Temporal Commonsense Reasoning in Dialog · ACL/IJCNLP (1) 2021
Natural language and speech › Information extraction and text analysis › entity linking
cross-lingual entity linking
0.422018
Joint Multilingual Supervision for Cross-lingual Entity Linking · EMNLP 2018
Bootstrapping Transliteration with Guided Discovery for Low-Resource Languages · EMNLP 2018
Natural language and speech › Information extraction and text analysis › text classification
topic classification
0.412019
Toward any-language zero-shot topic classification of textual documents · Artif. Intell. 2019
Natural language and speech › Information extraction and text analysis
entity linking
0.312018
Joint Multilingual Supervision for Cross-lingual Entity Linking · EMNLP 2018
Natural language and speech › Information extraction and text analysis › multilingual NLP
cross-lingual classification
0.212016
Cross-Lingual Dataless Classification for Many Languages · IJCAI 2016
Machine learning › Representation and self-supervised learning › word representation › word embedding
cross-lingual word embedding
0.212016
Cross-lingual Models of Word Embeddings: An Empirical Comparison · ACL (1) 2016
Natural language and speech › Information extraction and text analysis › text classification › transfer learning for text classification
dataless classification
0.212016
Cross-Lingual Dataless Classification for Many Languages · IJCAI 2016
Natural language and speech › Question answering and dialogue systems
math word problem solving
0.212016
Learning from Explicit and Implicit Supervision Jointly For Algebra Word Problems · EMNLP 2016
Natural language and speech › Information extraction and text analysis
text classification
0.212016
Cross-Lingual Dataless Classification for Many Languages · IJCAI 2016
Machine learning › Optimization for machine learning › coordinate descent
dual coordinate ascent
0.212015
Structural Learning with Amortized Inference · AAAI 2015
Machine learning and data management
structured prediction
0.212015
Structural Learning with Amortized Inference · AAAI 2015
Machine learning › Deep learning architectures and training
transformer
0.212022
TableFormer: Robust Transformer Modeling for Table-Text Encoding · ACL (1) 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.112021
TIMEDIAL: Temporal Commonsense Reasoning in Dialog · ACL/IJCNLP (1) 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
temporal reasoning
0.112021
TIMEDIAL: Temporal Commonsense Reasoning in Dialog · ACL/IJCNLP (1) 2021
Machine learning › Transfer learning and domain adaptation
zero-shot learning
0.112019
Toward any-language zero-shot topic classification of textual documents · Artif. Intell. 2019
Machine learning › Transfer learning and domain adaptation › cross-lingual transfer
multilingual transfer
0.112018
Joint Multilingual Supervision for Cross-lingual Entity Linking · EMNLP 2018
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
structured output learning
0.112016
Learning from Explicit and Implicit Supervision Jointly For Algebra Word Problems · EMNLP 2016

Methods — techniques the papers use, named apart from their topics

reasoning trace analysis · 1.0few-shot self-verification · 0.8POMDP · 0.8transformer · 0.6table-text encoding · 0.6attention bias · 0.6benchmark construction · 0.5joint multilingual supervision · 0.3constrained discovery · 0.3bootstrapping · 0.3structured perceptron · 0.2structured SVM · 0.2dual coordinate descent · 0.2
YearPublicationVenuePosition
2026 Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
abstract
Xinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang, Shyam Upadhyay. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinliang Frederick Zhang, Anhad Mohananey, Alexandra Chronopoulou, Pinelopi Papalampidi, Somit Gupta, Tsendsuren Munkhdalai, Lu Wang 0008, Shyam Upadhyay
ACL (1)8
2025 Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation
abstract
Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey, Steven Schwarcz, Adam Stambler, Shyam Upadhyay, Manaal Faruqui. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Satyapriya Krishna, Kalpesh Krishna, Anhad Mohananey, Steven Schwarcz, Adam Stambler, Shyam Upadhyay, Manaal Faruqui
NAACL (Long Papers)6
2024 AutoMix: Automatically Mixing Language Models
abstract
Large language models (LLMs) are now available from cloud API providers in various sizes and configurations. While this diversity offers a broad spectrum of choices, effectively leveraging the options to optimize computational cost and performance remains challenging. In this work, we present AutoMix, an approach that strategically routes queries to larger LMs, based on the approximate correctness of outputs from a smaller LM. Central to AutoMix are two key technical contributions. First, it has a few-shot self-verification mechanism, which estimates the reliability of its own outputs without requiring extensive training. Second, given that self-verification can be noisy, it employs a POMDP based router that can effectively select an appropriately sized model, based on answer confidence. Experiments across five language models and five challenging datasets show that Automix consistently surpasses strong baselines, reducing computational cost by over 50\% for comparable performance.
Pranjal Aggarwal, Aman Madaan, Ankit Anand, Srividya Pranavi Potharaju, Swaroop Mishra, Aditya Gupta 0001, Dheeraj Rajagopal, Karthik Kappaganthu, Yiming Yang 0002, Shyam Upadhyay, Manaal Faruqui, Mausam
NeurIPS11
2023 Efficient Encoders for Streaming Sequence Tagging
abstract
A naive application of state-of-the-art bidirectional encoders for streaming sequence tagging would require re-encoding all tokens from scratch whenever a new token appears in an incremental streaming input (like transcribed speech).The lack of re-usability of previous computation leads to a higher number of Floating Point Operations (or FLOPs) and higher number of unnecessary label flips.Increased FLOPs consequently lead to higher wall-clock time and increased label flipping leads to poorer streaming performance.In this work, we present Hybrid Encoder with Adaptive Restart (HEAR) that addresses these issues while maintaining the performance of bidirectional encoders over offline (or complete) inputs and improving performance on streaming (or incomplete) inputs.HEAR uses a HYBRID unidirectional-bidirectional encoder architecture to perform sequence tagging, along with an Adaptive Restart Module (ARM) to selectively guide the restart of bidirectional portion of the encoder.Across four sequence tagging tasks, HEAR offers FLOPs savings in streaming settings upto 71.1% and also outperforms bidirectional encoders for streaming predictions by upto +10% streaming exact match.
Ayush Kaushal, Aditya Gupta 0001, Shyam Upadhyay, Manaal Faruqui
EACL3
2022 TableFormer: Robust Transformer Modeling for Table-Text Encoding
abstract
Understanding tables is an important aspect of natural language understanding.Existing models for table understanding require linearization of the table structure, where row or column order is encoded as an unwanted bias.Such spurious biases make the model vulnerable to row and column order perturbations.Additionally, prior work has not thoroughly modeled the table structures or table-text alignments, hindering the table-text understanding ability.In this work, we propose a robust and structurally aware table-text encoding architecture TABLEFORMER, where tabular structural biases are incorporated completely through learnable attention biases.TABLEFORMER is (1) strictly invariant to row and column orders, and, (2) could understand tables better due to its tabular inductive biases.Our evaluations showed that TABLEFORMER outperforms strong baselines in all settings on SQA, WTQ and TABFACT table reasoning datasets, and achieves state-of-the-art performance on SQA, especially when facing answer-invariant row and column order perturbations (6% improvement over the best baseline), because previous SOTA models' performance drops by 4% -6% when facing such perturbations while TABLEFORMER is not affected.1
Jingfeng Yang 0001, Aditya Gupta 0001, Shyam Upadhyay, Luheng He, Rahul Goel, Shachi Paul
ACL (1)3
2022 Streaming Intended Query Detection using E2E Modeling for Continued Conversation
abstract
In voice-enabled applications, a predetermined hotword is usually used to activate a device in order to attend to the query.However, speaking queries followed by a hotword each time introduces a cognitive burden in continued conversations.To avoid repeating a hotword, we propose a streaming end-to-end (E2E) intended query detector that identifies the utterances directed towards the device and filters out other utterances not directed towards device.The proposed approach incorporates the intended query detector into the E2E model that already folds different components of the speech recognition pipeline into one neural network.The E2E modeling on speech decoding and intended query detection also allows us to declare a quick intended query detection based on early partial recognition result, which is important to decrease latency and make the system responsive.We demonstrate that the proposed E2E approach yields a 22% relative improvement on equal error rate (EER) for the detection accuracy and 600 ms latency improvement compared with an independent intended query detector.In our experiment, the proposed model detects whether the user is talking to the device with a 8.7% EER within 1.4 seconds of median latency after user starts speaking.
Shuo-Yiin Chang, Guru Prakash Arumugam, Zelin Wu, Tara N. Sainath, Bo Li 0028, Qiao Liang 0001, Adam Stambler, Shyam Upadhyay, Manaal Faruqui, Trevor Strohman
INTERSPEECH8
2021 TIMEDIAL: Temporal Commonsense Reasoning in Dialog
abstract
Lianhui Qin, Aditya Gupta, Shyam Upadhyay, Luheng He, Yejin Choi, Manaal Faruqui. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Lianhui Qin, Aditya Gupta 0001, Shyam Upadhyay, Luheng He, Yejin Choi 0001, Manaal Faruqui
ACL/IJCNLP (1)3
2019 A General-Purpose Algorithm for Constrained Sequential Inference
abstract
Inference in structured prediction involves finding the best output structure for an input, subject to certain constraints.Many current approaches use sequential inference, which constructs the output in a left-to-right manner.However, there is no general framework to specify constraints in these approaches.We present a principled approach for incorporating constraints into sequential inference algorithms.Our approach expresses constraints using an automaton, which is traversed in lockstep during inference, guiding the search to valid outputs.We show that automata can express commonly used constraints and are easily incorporated into sequential inference.When it is more natural to represent constraints as a set of automata, our algorithm uses an active set method for demonstrably fast and efficient inference.We experimentally show the benefits of our algorithm on constituency parsing and semantic role labeling.For parsing, unlike unconstrained approaches, our algorithm always generates valid output, incurring only a small drop in performance.For semantic role labeling, imposing constraints using our algorithm corrects common errors, improving F 1 by 1.5 points.These benefits increase in low-resource settings.Our active set method achieves a 5.2x relative speedup over a naive approach.1
Daniel Deutsch, Shyam Upadhyay, Dan Roth 0001
CoNLL2
2019 Toward any-language zero-shot topic classification of textual documents
Yangqiu Song, Shyam Upadhyay, Haoruo Peng, Stephen Mayhew 0001, Dan Roth 0001
Artif. Intell.2
2018 Joint Multilingual Supervision for Cross-lingual Entity Linking
abstract
Cross-lingual Entity Linking (XEL) aims to ground entity mentions written in any language to an English Knowledge Base (KB), such as Wikipedia.XEL for most languages is challenging, owing to limited availability of resources as supervision.We address this challenge by developing the first XEL approach that combines supervision from multiple languages jointly.This enables our approach to: (a) augment the limited supervision in the target language with additional supervision from a high-resource language (like English), and (b) train a single entity linking model for multiple languages, improving upon individually trained models for each language.Extensive evaluation on three benchmark datasets across 8 languages shows that our approach significantly improves over the current state-of-theart.We also provide analyses in two limited resource settings: (a) zero-shot setting, when no supervision in the target language is available, and in (b) low-resource setting, when some supervision in the target language is available.Our analysis provides insights into the limitations of zero-shot XEL approaches in realistic scenarios, and shows the value of joint supervision in low-resource settings.1
Shyam Upadhyay, Nitish Gupta, Dan Roth 0001
EMNLP1
2018 Bootstrapping Transliteration with Guided Discovery for Low-Resource Languages
abstract
Generating the English transliteration of a name written in a foreign script is an important and challenging step in multilingual knowledge acquisition and information extraction.Existing approaches to transliteration generation require a large (>5000) number of training examples.This difficulty contrasts with transliteration discovery, a somewhat easier task that involves picking a plausible transliteration from a given list.In this work, we present a bootstrapping algorithm that uses constrained discovery to improve generation, and can be used with as few as 500 training examples, which we show can be sourced from annotators in a matter of hours.This opens the task to languages for which large number of training examples are unavailable.We evaluate transliteration generation performance itself, as well the improvement it brings to crosslingual candidate generation for entity linking, a typical downstream task.We present a comprehensive evaluation of our approach on nine languages, each written in a unique script. 1
Shyam Upadhyay, Jordan Kodner, Dan Roth 0001
EMNLP1
2018 (Almost) Zero-Shot Cross-Lingual Spoken Language Understanding
abstract
Spoken language understanding (SLU) is a component of goal-oriented dialogue systems that aims to interpret user's natural language queries in system's semantic representation format. While current state-of-the-art SLU approaches achieve high performance for English domains, the same is not true for other languages. Approaches in the literature for extending SLU models and grammars to new languages rely primarily on machine translation. This poses a challenge in scaling to new languages, as machine translation systems may not be reliable for several (especially low resource) languages. In this work, we examine different approaches to train a SLU component with little supervision for two new languages - Hindi and Turkish, and show that with only a few hundred labeled examples we can surpass the approaches proposed in the literature. Our experiments show that training a model bilingually (i.e., jointly with English), enables faster learning, in that the model requires fewer labeled instances in the target language to generalize. Qualitative analysis shows that rare slot types benefit the most from the bilingual training.
Shyam Upadhyay, Manaal Faruqui, Gökhan Tür, Dilek Hakkani-Tür, Larry Heck
ICASSP1
2018 CogCompNLP: Your Swiss Army Knife for NLP
Daniel Khashabi, Mark Sammons, Ben Zhou, Tom Redman, Christos Christodoulopoulos 0001, Vivek Srikumar, Nick Rizzolo, Lev-Arie Ratinov, Guanheng Luo, Quang Do, Chen-Tse Tsai, Subhro Roy, Stephen Mayhew 0001, Zhili Feng, John Wieting, Xiaodong Yu 0003, Yangqiu Song, Shashank Gupta 0007, Shyam Upadhyay, Naveen Arivazhagan, Qiang Ning, Shaoshi Ling, Dan Roth 0001
LREC19
2018 Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple Sentences
abstract
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, Dan Roth. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth 0001, Shyam Upadhyay, Dan Roth 0001
NAACL-HLT4
2018 Robust Cross-Lingual Hypernymy Detection Using Dependency Context
abstract
Shyam Upadhyay, Yogarshi Vyas, Marine Carpuat, Dan Roth. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Shyam Upadhyay, Yogarshi Vyas, Marine Carpuat, Dan Roth 0001
NAACL-HLT1
2017 Annotating Derivations: A New Evaluation Strategy and Dataset for Algebra Word Problems
abstract
We propose a new evaluation for automatic solvers for algebra word problems, which can identify mistakes that existing evaluations overlook.Our proposal is to evaluate such solvers using derivations, which reflect how an equation system was constructed from the word problem.To accomplish this, we develop an algorithm for checking the equivalence between two derivations, and show how derivation annotations can be semi-automatically added to existing datasets.To make our experiments more comprehensive, we include the derivation annotation for DRAW-1K, a new dataset containing 1000 general algebra word problems.In our experiments, we found that the annotated derivations enable a more accurate evaluation of automatic solvers than previously used metrics.We release derivation annotations for over 2300 algebra word problems for future evaluations.
Shyam Upadhyay, Ming-Wei Chang
EACL (1)1
2016 Cross-lingual Models of Word Embeddings: An Empirical Comparison
abstract
Despite interest in using cross-lingual knowledge to learn word embeddings for various tasks, a systematic comparison of the possible approaches is lacking in the literature.We perform an extensive evaluation of four popular approaches of inducing cross-lingual embeddings, each requiring a different form of supervision, on four typologically different language pairs.Our evaluation setup spans four different tasks, including intrinsic evaluation on mono-lingual and cross-lingual similarity, and extrinsic evaluation on downstream semantic and syntactic applications.We show that models which require expensive cross-lingual knowledge almost always perform better, but cheaply supervised models often prove competitive on certain tasks.
Shyam Upadhyay, Manaal Faruqui, Chris Dyer, Dan Roth 0001
ACL (1)1
2016 Revisiting the Evaluation for Cross Document Event Coreference
abstract
Cross document event coreference (CDEC) is an important task that aims at aggregating event-related information across multiple documents. We revisit the evaluation for CDEC, and discover that past works have adopted different, often inconsistent, evaluation settings, which either overlook certain mistakes in coreference decisions, or make assumptions that simplify the coreference task considerably. We suggest a new evaluation methodology which overcomes these limitations, and allows for an accurate assessment of CDEC systems. Our new evaluation setting better reflects the corpus-wide information aggregation ability of CDEC systems by separating event-coreference decisions made across documents from those made within a document. In addition, we suggest a better baseline for the task and semi-automatically identify several inconsistent annotations in the evaluation dataset.
Shyam Upadhyay, Nitish Gupta, Christos Christodoulopoulos 0001, Dan Roth 0001
COLING1
2016 Equation Parsing : Mapping Sentences to Grounded Equations
abstract
Identifying mathematical relations expressed in text is essential to understanding a broad range of natural language text from election reports, to financial news, to sport commentaries to mathematical word problems.This paper focuses on identifying and understanding mathematical relations described within a single sentence.We introduce the problem of Equation Parsing -given a sentence, identify noun phrases which represent variables, and generate the mathematical equation expressing the relation described in the sentence.We introduce the notion of projective equation parsing and provide an efficient algorithm to parse text to projective equations.Our system makes use of a high precision lexicon of mathematical expressions and a pipeline of structured predictors, and generates correct equations in 70% of the cases.In 60% of the time, it also identifies the correct noun phrase → variables mapping, significantly outperforming baselines.We also release a new annotated dataset for task evaluation.
Subhro Roy, Shyam Upadhyay, Dan Roth 0001
EMNLP2
2016 Learning from Explicit and Implicit Supervision Jointly For Algebra Word Problems
abstract
Automatically solving algebra word problems has raised considerable interest recently.Existing state-of-the-art approaches mainly rely on learning from human annotated equations.In this paper, we demonstrate that it is possible to efficiently mine algebra problems and their numerical solutions with little to no manual effort.To leverage the mined dataset, we propose a novel structured-output learning algorithm that aims to learn from both explicit (e.g., equations) and implicit (e.g., solutions) supervision signals jointly.Enabled by this new algorithm, our model gains 4.6% absolute improvement in accuracy on the ALG-514 benchmark compared to the one without using implicit supervision.The final model also outperforms the current state-of-the-art approach by 3%.
Shyam Upadhyay, Ming-Wei Chang, Kai-Wei Chang 0001, Scott Yih
EMNLP1
2016 Cross-Lingual Dataless Classification for Many Languages
Yangqiu Song, Shyam Upadhyay, Haoruo Peng, Dan Roth 0001
IJCAI2
2015 Structural Learning with Amortized Inference
abstract
Training a structured prediction model involves performing several loss-augmented inference steps. Over the lifetime of the training, many of these inference problems, although different, share the same solution. We propose AI-DCD, an Amortized Inference framework for Dual Coordinate Descent method, an approximate learning algorithm, that accelerates the training process by exploiting this redundancy of solutions, without compromising the performance of the model. We show the efficacy of our method by training a structured SVM using dual coordinate descent for an entityrelation extraction task. Our method learns the same model as an exact training algorithm would, but call the inference engine only in 10% – 24% of the inference problems encountered during training. We observe similar gains on a multi-label classification task and with a Structured Perceptron model for the entity-relation task.
Kai-Wei Chang 0001, Shyam Upadhyay, Gourab Kundu, Dan Roth 0001
AAAI2