VLDB 2026 Research / reviewers in the wild / expert
Rajakrishnan Rajkumar
dblp:36/8159
· DBLP profile ↗
9ranked-venue papers
1as first author
4since 2021 · last 2025
0000-0003-2052-4082ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 59% Information extraction and text analysis · 32% Machine translation · 9% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 8 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis
syntactic parsing |
0.6 | 1 | 2022 | Discourse Context Predictability Effects in Hindi Word Order · EMNLP 2022 |
Natural language and speech › Language models and text generation
text generation |
0.2 | 2 | 2012 | Minimal Dependency Length in Realization Ranking · EMNLP-CoNLL 2012 Perceptron Reranking for CCG Realization · EMNLP 2009 |
Natural language and speech › Language models and text generation › text generation
surface realization |
0.2 | 2 | 2010 | Further Meta-Evaluation of Broad-Coverage Surface Realization · EMNLP 2010 Perceptron Reranking for CCG Realization · EMNLP 2009 |
Mathematical optimization
combinatorial optimization |
0.1 | 1 | 2012 | Minimal Dependency Length in Realization Ranking · EMNLP-CoNLL 2012 |
Natural language and speech › Machine translation
word reordering |
0.1 | 1 | 2011 | A Word Reordering Model for Improved Machine Translation · EMNLP 2011 |
Program synthesis and code generation
grammar-constrained generation |
0.1 | 1 | 2009 | Perceptron Reranking for CCG Realization · EMNLP 2009 |
Natural language and speech › Machine translation
neural machine translation |
0.0 | 1 | 2011 | A Word Reordering Model for Improved Machine Translation · EMNLP 2011 |
Natural language and speech › Language models and text generation
text generation evaluation |
0.0 | 1 | 2010 | Further Meta-Evaluation of Broad-Coverage Surface Realization · EMNLP 2010 |
Methods — techniques the papers use, named apart from their topics
syntactic priming · 0.6classifier-based prediction · 0.6LSTM · 0.6perceptron reranking · 0.2word reordering model · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Psycholinguistic Features Predict Word Duration in Hindi Read Aloud SpeechabstractReliable assessment of oral reading fluency (ORF) is of great importance in foundational literacy missions globally. For the design of level appropriate testing passages, text difficulty has traditionally been based on coarse-grained measures of readability like the Flesch–Kincaid score. We present a novel study where we deploy psycholinguistic measures of reading difficulty from Natural Language Processing to predict the duration of words in Hindi read-aloud speech. We test the hypotheses that expectation-based measures of linguistic complexity are significant predictors of word duration in Hindi read-aloud speech. We validate the stated hypotheses by estimating surprisal measures inspired from Surprisal Theory of sentence comprehension and introduce a novel measure of orthographic complexity to model the intricacies of the Hindi script. Cognitive modelling experiments were conducted on a dataset of six Hindi short stories read aloud by 5 expert readers, containing 2 measures of word duration. Our results show that both surprisal as well as the orthographic complexity measures are significant predictors of word duration. In contrast with long words, we find duration reducing with increased orthographic complexity in the case of short words. The variation between individual speakers in terms of word duration is very low and the variance in the data is caused by the properties of the words used in the text. Finally, we reflect on the implications of our work for cognitive models of language production and for ORF assessment. Rajakrishnan Rajkumar, Sneha Raman, Aadya Ranjan, Mildred Pereira, Nagesh Nayak, Preeti Rao |
ICASSP | 1 |
| 2022 | Interference and Case Marker Effects in Dependency Locality: Insights from Hindi
Sidharth Ranjan, Rajakrishnan Rajkumar, Sumeet Agarwal |
CogSci | 2 |
| 2022 | Linguistically Motivated Features for Classifying Shorter Text into Fiction and Non-Fiction GenreabstractThis work deploys linguistically motivated features to classify paragraph-level text into fiction and non-fiction genre using a logistic regression model and infers lexical and syntactic properties that distinguish the two genres. Previous works have focused on classifying document-level text into fiction and non-fiction genres, while in this work, we deal with shorter texts which are closer to real-world applications like sentiment analysis of tweets. Going beyond simple POS tag ratios proposed in Qureshi et al.(2019) for document-level classification, we extracted multiple linguistically motivated features belonging to four categories: Lexical features, POS ratio features, Syntactic features and Raw features. For the task of short-text classification, a model containing 28 best-features (selected via Recursive feature elimination with cross-validation; RFECV) confers an accuracy jump of 15.56 % over a baseline model consisting of 2 POS-ratio features found effective in previous work (cited above). The efficacy of the above model containing a linguistically motivated feature set also transfers over to another dataset viz, Baby BNC corpus. We also compared the classification accuracy of the logistic regression model with two deep-learning models. A 1D CNN model gives an increase of 2% accuracy over the logistic Regression classifier on both corpora. And the BERT-base-uncased model gives the best classification accuracy of 97% on Brown corpus and 98% on Baby BNC corpus. Although both the deep learning models give better results in terms of classification accuracy, the problem of interpreting these models remains unsolved. In contrast, regression model coefficients revealed that fiction texts tend to have more character-level diversity and have lower lexical density (quantified using content-function word ratios) compared to non-fiction texts. Moreover, subtle differences in word order exist between the two genres, i.e., in fiction texts Verbs precede Adverbs (inter-alia). Arman Kazmi, Sidharth Ranjan, Arpit Sharma 0002, Rajakrishnan Rajkumar |
COLING | 4 |
| 2022 | Discourse Context Predictability Effects in Hindi Word OrderabstractWe test the hypothesis that discourse predictability influences Hindi syntactic choice.While prior work has shown that a number of factors (e.g., information status, dependency length, and syntactic surprisal) influence Hindi word order preferences, the role of discourse predictability is underexplored in the literature.Inspired by prior work on syntactic priming, we investigate how the words and syntactic structures in a sentence influence the word order of the following sentences.Specifically, we extract sentences from the Hindi-Urdu Treebank corpus (HUTB), permute the preverbal constituents of those sentences, and build a classifier to predict which sentences actually occurred in the corpus against artificially generated distractors.The classifier uses a number of discourse-based features and cognitive features to make its predictions, including dependency length, surprisal, and information status.We find that information status and LSTM-based discourse predictability influence word order choices, especially for non-canonical objectfronted orders.We conclude by situating our results within the broader syntactic priming literature. Sidharth Ranjan, Marten van Schijndel, Sumeet Agarwal, Rajakrishnan Rajkumar |
EMNLP | 4 |
| 2012 | Minimal Dependency Length in Realization Ranking
Michael White 0001, Rajakrishnan Rajkumar |
EMNLP-CoNLL | 2 |
| 2011 | A Word Reordering Model for Improved Machine Translation
Karthik Visweswariah, Rajakrishnan Rajkumar, Ankur Gandhe, Ananthakrishnan Ramanathan, Jirí Navrátil 0001 |
EMNLP | 2 |
| 2010 | Further Meta-Evaluation of Broad-Coverage Surface Realization
Dominic Espinosa, Rajakrishnan Rajkumar, Michael White 0001, Shoshana Berleant |
EMNLP | 2 |
| 2009 | Perceptron Reranking for CCG Realization
Michael White 0001, Rajakrishnan Rajkumar |
EMNLP | 2 |
| 2009 | Eye tracking for the online evaluation of prosody in speech synthesis: not so fast!abstractThis paper presents an eye-tracking experiment comparing the processing of different accent patterns in unit selection synthesis and human speech. The synthetic speech results failed to replicate the facilitative effect of contextually appropriate accent patterns found with human speech, while producing a more robust intonational garden-path effect with contextually inappropriate patterns, both of which could be due to processing delays seen with the synthetic speech. As the synthetic speech was of high quality, the results indicate that eye tracking holds promise as a highly sensitive and objective method for the online evaluation of prosody in speech synthesis. Index Terms: speech synthesis, evaluation, prosody, eye tracking, unit selection Michael White 0001, Rajakrishnan Rajkumar, Kiwako Ito, Shari R. Speer |
INTERSPEECH | 2 |