Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Riyaz A. Bhat

dblp:146/3952 · also Riyaz Ahmad Bhat · DBLP profile ↗
← Back
14ranked-venue papers
6as first author
4since 2021 · last 2026
0000-0002-8327-2882ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Artificial intelligence
2 papers
Language models and text generation · 78% Question answering and dialogue systems · 22%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › ranking
learning to rank
1.012026
Logit Inflation in ListMLE: Theoretical Analysis and Mitigation Strategies · SIGIR 2026
Information retrieval › ranking › learning to rank
listwise ranking
1.012026
Logit Inflation in ListMLE: Theoretical Analysis and Mitigation Strategies · SIGIR 2026
Information retrieval › ranking
ranking calibration
1.012026
Logit Inflation in ListMLE: Theoretical Analysis and Mitigation Strategies · SIGIR 2026
Natural language and speech › Language models and text generation
instruction following
0.712023
Prompting with Pseudo-Code Instructions · EMNLP 2023
Natural language and speech › Language models and text generation
prompting
0.712023
Prompting with Pseudo-Code Instructions · EMNLP 2023
Natural language and speech › Question answering and dialogue systems
evidence selection
0.612022
DocInfer: Document-level Natural Language Inference using Optimal Evidence Selection · EMNLP 2022
Natural language and speech › Language models and text generation › natural language understanding › sentence pair modeling
natural language inference
0.612022
DocInfer: Document-level Natural Language Inference using Optimal Evidence Selection · EMNLP 2022
Natural language and speech › Language models and text generation › natural language understanding
long document understanding
0.212022
DocInfer: Document-level Natural Language Inference using Optimal Evidence Selection · EMNLP 2022

Methods — techniques the papers use, named apart from their topics

ListMLE · 1.0large language model · 0.7subgraph pooling · 0.6hierarchical document graph · 0.6REINFORCE · 0.6BERT · 0.6
YearPublicationVenuePosition
2026 Logit Inflation in ListMLE: Theoretical Analysis and Mitigation Strategies
abstract
Modern learning-to-rank methods often rely on listwise objectives that directly model and optimize relative document order over entire permutations. While these objectives improve ranking quality, they frequently produce models with highly inflated relevance scores whose magnitudes exceed what is necessary for meaningful document separation, leading to poor probabilistic calibration.
Riyaz A. Bhat, Jaydeep Sen
SIGIR1
2023 Prompting with Pseudo-Code Instructions
abstract
Prompting with natural language instructions has recently emerged as a popular method of harnessing the capabilities of large language models (LLM).Given the inherent ambiguity present in natural language, it is intuitive to consider the possible advantages of prompting with less ambiguous prompt styles, like pseudocode.In this paper, we explore if prompting via pseudo-code instructions helps improve the performance of pre-trained language models.We manually create a dataset 1 of pseudo-code prompts for 132 different tasks spanning classification, QA, and generative language tasks, sourced from the Super-NaturalInstructions dataset (Wang et al., 2022b).Using these prompts along with their counterparts in natural language, we study their performance on two LLM families -BLOOM (Scao et al., 2023), CodeGen (Nijkamp et al., 2023).Our experiments show that using pseudo-code instructions leads to better results, with an average increase (absolute) of 7-16 points in F1 scores for classification tasks and an improvement (relative) of 12-38% in aggregate ROUGE-L scores across all tasks.We include detailed ablation studies which indicate that code comments, docstrings, and the structural clues encoded in pseudo-code all contribute towards the improvement in performance.To the best of our knowledge, our work is the first to demonstrate how pseudocode prompts can be helpful in improving the performance of pre-trained LMs.* Equal contribution 1 Code and dataset available at https://github.com/ mayank31398/pseudo-code-instructions Listing 1 An example pseudo-code instruction for the task from Wang et al. (2022b).A successful model is expected to use the provided pseudo-code instructions and output responses to a pool of evaluation instances.1 def generate_sentiment(sentence: str) -> str: 2 """For the given sentence, the task is to 3 predict the sentiment.For positive 4 sentiment return "positive" else return 5 "negative".
Riyaz A. Bhat, Rudra Murthy V, Danish Contractor, Srikanth Tamilselvam
EMNLP3
2022 DocInfer: Document-level Natural Language Inference using Optimal Evidence Selection
abstract
We present DocInfer -a novel, end-to-end Document-level Natural Language Inference model that builds a hierarchical document graph enriched through inter-sentence relations (topical, entity-based, concept-based), performs paragraph pruning using the novel SubGraph Pooling layer, followed by optimal evidence selection based on REINFORCE algorithm to identify the most important context sentences for a given hypothesis.Our evidence selection mechanism allows it to transcend the input length limitation of modern BERT-like Transformer models while presenting the entire evidence together for inferential reasoning.We show this is an important property needed to reason on large documents where the evidence may be fragmented and located arbitrarily far from each other.Extensive experiments on popular corpora -DocNLI, ContractNLI, and ConTRoL datasets, and our new proposed dataset called CaseHoldNLI on the task of legal judicial reasoning, demonstrate significant performance gains of 8-12% over SOTA methods.Our ablation studies validate the impact of our model.Performance improvement of ∼ 3 -6% on annotation-scarce downstream tasks of fact verification, multiple-choice QA, and contract clause retrieval demonstrates the usefulness of DocInfer beyond primary NLI tasks.
Puneet Mathur, Gautam Kunapuli, Riyaz A. Bhat, Manish Shrivastava 0001, Dinesh Manocha, Maneesh Kumar Singh 0001
EMNLP3
2022 Is My Model Using The Right Evidence? Systematic Probes for Examining Evidence-Based Tabular Reasoning
abstract
Abstract Neural models command state-of-the-art performance across NLP tasks, including ones involving “reasoning”. Models claiming to reason about the evidence presented to them should attend to the correct parts of the input while avoiding spurious patterns therein, be self-consistent in their predictions across inputs, and be immune to biases derived from their pre-training in a nuanced, context- sensitive fashion. Do the prevalent *BERT- family of models do so? In this paper, we study this question using the problem of reasoning on tabular data. Tabular inputs are especially well-suited for the study—they admit systematic probes targeting the properties listed above. Our experiments demonstrate that a RoBERTa-based model, representative of the current state-of-the-art, fails at reasoning on the following counts: it (a) ignores relevant parts of the evidence, (b) is over- sensitive to annotation artifacts, and (c) relies on the knowledge encoded in the pre-trained language model rather than the evidence presented in its tabular inputs. Finally, through inoculation experiments, we show that fine- tuning the model on perturbed data does not help it overcome the above challenges.
Vivek Gupta 0001, Riyaz A. Bhat, Atreya Ghosal, Manish Shrivastava 0001, Maneesh Kumar Singh 0001, Vivek Srikumar
Trans. Assoc. Comput. Linguistics2
2019 Neural Transition Systems for Modeling Hierarchical Semantic Representations
Riyaz A. Bhat, John Chen 0001, Rashmi Prasad, Srinivas Bangalore
INTERSPEECH1
2018 Universal Dependency Parsing for Hindi-English Code-Switching
abstract
Irshad Bhat, Riyaz A. Bhat, Manish Shrivastava, Dipti Sharma. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Irshad Ahmad Bhat, Riyaz A. Bhat, Manish Shrivastava 0001, Dipti Misra Sharma
NAACL-HLT2
2017 Improving Transition-Based Dependency Parsing of Hindi and Urdu by Modeling Syntactically Relevant Phenomena
abstract
In recent years, transition-based parsers have shown promise in terms of efficiency and accuracy. Though these parsers have been extensively explored for multiple Indian languages, there is still considerable scope for improvement by properly incorporating syntactically relevant information. In this article, we enhance transition-based parsing of Hindi and Urdu by redefining the features and feature extraction procedures that have been previously proposed in the parsing literature of Indian languages. We propose and empirically show that properly incorporating syntactically relevant information like case marking, complex predication and grammatical agreement in an arc-eager parsing model can significantly improve parsing accuracy. Our experiments show an absolute improvement of ∼2% LAS for parsing of both Hindi and Urdu over a competitive baseline which uses rich features like part-of-speech (POS) tags, chunk tags, cluster ids and lemmas. We also propose some heuristics to identify ezafe constructions in Urdu texts which show promising results in parsing these constructions.
Riyaz A. Bhat, Irshad Ahmad Bhat, Dipti Misra Sharma
ACM Trans. Asian Low Resour. Lang. Inf. Process.1
2016 A House United: Bridging the Script and Lexical Barrier between Hindi and Urdu
abstract
In Computational Linguistics, Hindi and Urdu are not viewed as a monolithic entity and have received separate attention with respect to their text processing. From part-of-speech tagging to machine translation, models are separately trained for both Hindi and Urdu despite the fact that they represent the same language. The reasons mainly are their divergent literary vocabularies and separate orthographies, and probably also their political status and the social perception that they are two separate languages. In this article, we propose a simple but efficient approach to bridge the lexical and orthographic differences between Hindi and Urdu texts. With respect to text processing, addressing the differences between the Hindi and Urdu texts would be beneficial in the following ways: (a) instead of training separate models, their individual resources can be augmented to train single, unified models for better generalization, and (b) their individual text processing applications can be used interchangeably under varied resource conditions. To remove the script barrier, we learn accurate statistical transliteration models which use sentence-level decoding to resolve word ambiguity. Similarly, we learn cross-register word embeddings from the harmonized Hindi and Urdu corpora to nullify their lexical divergences. As a proof of the concept, we evaluate our approach on the Hindi and Urdu dependency parsing under two scenarios: (a) resource sharing, and (b) resource augmentation. We demonstrate that a neural network-based dependency parser trained on augmented, harmonized Hindi and Urdu resources performs significantly better than the parsing models trained separately on the individual resources. We also show that we can achieve near state-of-the-art results when the parsers are used interchangeably.
Riyaz A. Bhat, Irshad Ahmad Bhat, Naman Jain, Dipti Misra Sharma
COLING1
2016 A Proposition Bank of Urdu
Maaz Anwar, Riyaz A. Bhat, Dipti Misra Sharma, Ashwini Vaidya, Martha Palmer, Tafseer Ahmed
LREC2
2014 Towards building a Kashmiri Treebank: Setting up the Annotation Pipeline
Riyaz A. Bhat, Shahid Musjtaq Bhat, Dipti Misra Sharma
LREC1
2013 Animacy Acquisition Using Morphological Case
Riyaz A. Bhat, Dipti Misra Sharma
IJCNLP1
2013 Exploring Semantic Information in Hindi WordNet for Hindi Dependency Parsing
Sambhav Jain, Naman Jain, Aniruddha Tammewar, Riyaz A. Bhat, Dipti Misra Sharma
IJCNLP4
2013 Automatic Clause Boundary Annotation in the Hindi Treebank
Soma Paul, Riyaz A. Bhat, Sambhav Jain
PACLIC3
2013 Transliteration Systems across Indian Languages Using Parallel Corpora
Rishabh Srivastava, Riyaz A. Bhat
PACLIC2