VLDB 2026 Research / reviewers in the wild / expert
Abraham Ittycheriah
dblp:72/3792 · also Abe Ittycheriah
· DBLP profile ↗
27ranked-venue papers
4as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
9 papers |
Language models and text generation · 39% Machine translation · 17% Reinforcement learning · 17% | |
| Databases, data mining, and information retrieval
1 paper |
Web and social media mining · 100% |
Topics — the 16 heaviest of 19, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
alignment |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Natural language and speech › Language models and text generation › alignment
reward hacking |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Machine learning › Reinforcement learning › reward learning
reward model training |
0.9 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Natural language and speech › Machine translation
neural machine translation |
0.5 | 2 | 2016 | Supervised Attentions for Neural Machine Translation · EMNLP 2016 Coverage Embedding Models for Neural Machine Translation · EMNLP 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
structured information extraction |
0.4 | 1 | 2020 | Factoring Fact-Checks: Structured Information Extraction from Fact-Checking Articles · WWW 2020 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
0.3 | 1 | 2025 | RRM: Robust Reward Model Training Mitigates Reward Hacking · ICLR 2025 |
Machine learning › Deep learning architectures and training › attention mechanism
attention alignment |
0.2 | 1 | 2016 | Supervised Attentions for Neural Machine Translation · EMNLP 2016 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.2 | 1 | 2016 | Coverage Embedding Models for Neural Machine Translation · EMNLP 2016 |
Machine learning › Deep learning architectures and training › regularization › attention regularization
supervised attention |
0.2 | 1 | 2016 | Supervised Attentions for Neural Machine Translation · EMNLP 2016 |
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation |
0.2 | 1 | 2014 | Adaptive HTER Estimation for Document-Specific MT Post-Editing · ACL (1) 2014 |
Natural language and speech › Machine translation › statistical machine translation
word alignment |
0.1 | 1 | 2011 | A Correction Model for Word Alignments · EMNLP 2011 |
Natural language and speech › Machine translation › computer-assisted translation
post-editing |
0.1 | 1 | 2014 | Adaptive HTER Estimation for Document-Specific MT Post-Editing · ACL (1) 2014 |
Natural language and speech › Information extraction and text analysis
named entity recognition |
0.1 | 2 | 2003 | HowtogetaChineseName(Entity): Segmentation and Combination Issues · EMNLP 2003 tRuEcasIng · ACL 2003 |
Natural language and speech › Information extraction and text analysis
coreference resolution |
0.0 | 1 | 2004 | A Mention-Synchronous Coreference Resolution Algorithm Based On the Bell Tree · ACL 2004 |
Natural language and speech › Information extraction and text analysis › named entity recognition
chinese named entity recognition |
0.0 | 1 | 2003 | HowtogetaChineseName(Entity): Segmentation and Combination Issues · EMNLP 2003 |
Natural language and speech › Information extraction and text analysis
word segmentation |
0.0 | 1 | 2003 | HowtogetaChineseName(Entity): Segmentation and Combination Issues · EMNLP 2003 |
Methods — techniques the papers use, named apart from their topics
data augmentation · 0.9causal framework · 0.9sequence tagging · 0.9BERT fine-tuning · 0.9attention mechanism · 0.5coverage embedding · 0.2alignment supervision · 0.2regression · 0.2feature engineering · 0.2classification · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RRM: Robust Reward Model Training Mitigates Reward HackingabstractReward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. However, traditional RM training, which relies on response pairs tied to specific prompts, struggles to disentangle prompt-driven preferences from prompt-independent artifacts, such as response length and format. In this work, we expose a fundamental limitation of current RM training methods, where RMs fail to effectively distinguish between contextual signals and irrelevant artifacts when determining preferences. To address this, we introduce a causal framework that learns preferences independent of these artifacts and propose a novel data augmentation technique designed to eliminate them. Extensive experiments show that our approach successfully filters out undesirable artifacts, yielding a more robust reward model (RRM). Our RRM improves the performance of a pairwise reward model trained on Gemma-2-9b-it, on Reward-Bench, increasing accuracy from 80.61% to 84.15%. Additionally, we train two DPO policies using both the RM and RRM, demonstrating that the RRM significantly enhances DPO-aligned policies, improving MT-Bench scores from 7.27 to 8.31 and length-controlled win-rates in AlpacaEval-2 from 33.46% to 52.49%. Tianqi Liu 0002, Wei Xiong 0015, Jie Ren 0006, Lichang Chen, Rishabh Joshi, Zhen Qin 0001, Tianhe Yu, Daniel Sohn, Anastasia Makarova, Jeremiah Z. Liu, Bilal Piot, Abraham Ittycheriah, Aviral Kumar, Mohammad Saleh |
ICLR | 16 |
| 2020 | Factoring Fact-Checks: Structured Information Extraction from Fact-Checking ArticlesabstractFact-checking, which investigates claims made in public to arrive at a verdict supported by evidence and logical reasoning, has long been a significant form of journalism to combat misinformation in the news ecosystem. Most of the fact-checks share common structured information (called factors) such as claim, claimant, and verdict. In recent years, the emergence of ClaimReview as the standard schema for annotating those factors within fact-checking articles has led to wide adoption of fact-checking features by online platforms (e.g., Google, Bing). However, annotating fact-checks is a tedious process for fact-checkers and distracts them from their core job of investigating claims. As a result, less than half of the fact-checkers worldwide have adopted ClaimReview as of mid-2019. In this paper, we propose the task of factoring fact-checks for automatically extracting structured information from fact-checking articles. Exploring a public dataset of fact-checks, we empirically show that factoring fact-checks is a challenging task, especially for fact-checkers that are under-represented in the existing dataset. We then formulate the task as a sequence tagging problem and fine-tune the pre-trained BERT models with a modification made from our observations to approach the problem. Through extensive experiments, we demonstrate the performance of our models for well-known fact-checkers and promising initial results for under-represented fact-checkers. Shan Jiang 0008, Simon Baumgartner, Abraham Ittycheriah, Cong Yu 0001 |
WWW | 3 |
| 2016 | Sentence Similarity Learning by Lexical Decomposition and CompositionabstractMost conventional sentence similarity methods only focus on similar parts of two input sentences, and simply ignore the dissimilar parts, which usually give us some clues and semantic meanings about the sentences. In this work, we propose a model to take into account both the similarities and dissimilarities by decomposing and composing lexical semantics over sentences. The model represents each word as a vector, and calculates a semantic matching vector for each word based on all words in the other sentence. Then, each word vector is decomposed into a similar component and a dissimilar component based on the semantic matching vector. After this, a two-channel CNN model is employed to capture features by composing the similar and dissimilar components. Finally, a similarity score is estimated over the composed feature vectors. Experimental results show that our model gets the state-of-the-art performance on the answer sentence selection task, and achieves a comparable result on the paraphrase identification task. Haitao Mi, Abraham Ittycheriah |
COLING | 3 |
| 2016 | Semi-supervised Clustering for Short Text via Deep Representation LearningabstractIn this work, we propose a semisupervised method for short text clustering, where we represent texts as distributed vectors with neural networks, and use a small amount of labeled data to specify our intention for clustering.We design a novel objective to combine the representation learning process and the kmeans clustering process together, and optimize the objective with both labeled data and unlabeled data iteratively until convergence through three steps: (1) assign each short text to its nearest centroid based on its representation from the current neural networks; (2) re-estimate the cluster centroids based on cluster assignments from step (1); (3) update neural networks according to the objective by keeping centroids and cluster assignments fixed.Experimental results on four datasets show that our method works significantly better than several other text clustering methods. Haitao Mi, Abraham Ittycheriah |
CoNLL | 3 |
| 2016 | Coverage Embedding Models for Neural Machine TranslationabstractIn this paper, we enhance the attention-based neural machine translation (NMT) by adding explicit coverage embedding models to alleviate issues of repeating and dropping translations in NMT.For each source word, our model starts with a full coverage embedding vector to track the coverage status, and then keeps updating it with neural networks as the translation goes.Experiments on the large-scale Chinese-to-English task show that our enhanced model improves the translation quality significantly on various test sets over the strong large vocabulary NMT system. Haitao Mi, Baskaran Sankaran, Abraham Ittycheriah |
EMNLP | 4 |
| 2016 | Supervised Attentions for Neural Machine TranslationabstractIn this paper, we improve the attention or alignment accuracy of neural machine translation by utilizing the alignments of training sentence pairs.We simply compute the distance between the machine attentions and the "true" alignments, and minimize this cost in the training procedure.Our experiments on large-scale Chinese-to-English task show that our model improves both translation and alignment qualities significantly over the large-vocabulary neural machine translation system, and even beats a state-of-the-art traditional syntax-based system. Haitao Mi, Abraham Ittycheriah |
EMNLP | 3 |
| 2014 | Adaptive HTER Estimation for Document-Specific MT Post-EditingabstractWe present an adaptive translation quality estimation (QE) method to predict the human-targeted translation error rate (HTER) for a document-specific machine translation model.We first introduce features derived internal to the translation decoding process as well as externally from the source sentence analysis.We show the effectiveness of such features in both classification and regression of MT quality.By dynamically training the QE model for the document-specific MT model, we are able to achieve consistency and prediction quality across multiple documents, demonstrated by the higher correlation coefficient and F-scores in finding Good sentences.Additionally, the proposed method is applied to IBM English-to-Japanese MT post editing field study and we observe strong correlation with human preference, with a 10% increase in human translators' productivity. Fei Huang 0002, Jian-Ming Xu, Abraham Ittycheriah, Salim Roukos |
ACL (1) | 3 |
| 2014 | Improving Egyptian-to-English SMT by Mapping Egyptian into MSA
Nadir Durrani, Yaser Al-Onaizan, Abraham Ittycheriah |
CICLing (2) | 3 |
| 2014 | Improving MT post-editing productivity with adaptive confidence estimation for document-specific translation model
Fei Huang 0002, Jian-Ming Xu, Abraham Ittycheriah, Salim Roukos |
Mach. Transl. | 3 |
| 2012 | Document-Specific Statistical Machine Translation for Improving Human Translation Productivity
Salim Roukos, Abraham Ittycheriah, Jian-Ming Xu |
CICLing (2) | 2 |
| 2011 | A Correction Model for Word Alignments
J. Scott McCarley, Abraham Ittycheriah, Salim Roukos, Bing Xiang, Jian-Ming Xu |
EMNLP | 2 |
| 2010 | Decoding with shrinkage-based language models
Ahmad Emami, Stanley F. Chen, Abraham Ittycheriah, Hagen Soltau |
INTERSPEECH | 3 |
| 2007 | Direct Translation Model 2
Abraham Ittycheriah, Salim Roukos |
HLT-NAACL | 1 |
| 2004 | A Mention-Synchronous Coreference Resolution Algorithm Based On the Bell TreeabstractThis paper proposes a new approach for coreference resolution which uses the Bell tree to represent the search space and casts the coreference resolution problem as finding the best path from the root of the Bell tree to the leaf nodes. A Maximum Entropy model is used to rank these paths. The coreference performance on the 2002 and 2003 Automatic Content Extraction (ACE) data will be reported. We also train a coreference system using the MUC6 data and competitive results are obtained. Xiaoqiang Luo, Abraham Ittycheriah, Hongyan Jing, Nanda Kambhatla, Salim Roukos |
ACL | 2 |
| 2004 | A Statistical Model for Multilingual Entity Detection and Tracking
Radu Florian, Hany Hassan, Abraham Ittycheriah, Hongyan Jing, Nanda Kambhatla, Xiaoqiang Luo, Nicolas Nicolov, Salim Roukos |
HLT-NAACL | 3 |
| 2003 | tRuEcasIngabstractTruecasing is the process of restoring case information to badly-cased or non-cased text. This paper explores truecasing issues and proposes a statistical, language modeling based truecaser which achieves an accuracy of ~98% on news articles. Task based evaluation shows a 26% F-measure improvement in named entity recognition when using truecasing. In the context of automatic content extraction, mention detection on automatic speech recognition text is also improved by a factor of 8. Truecasing also enhances machine translation output legibility and yields a BLEU score improvement of 80.2%. This paper argues for the use of truecasing as a valuable component in text processing applications. Lucian Vlad Lita, Abraham Ittycheriah, Salim Roukos, Nanda Kambhatla |
ACL | 2 |
| 2003 | Named Entity Recognition through Classifier Combination
Radu Florian, Abraham Ittycheriah, Hongyan Jing, Tong Zhang 0001 |
CoNLL | 2 |
| 2003 | HowtogetaChineseName(Entity): Segmentation and Combination Issues
Hongyan Jing, Radu Florian, Xiaoqiang Luo, Tong Zhang 0001, Abraham Ittycheriah |
EMNLP | 5 |
| 2003 | In Question Answering, Two Heads Are Better Than One
Jennifer Chu-Carroll, Krzysztof Czuba, John M. Prager, Abraham Ittycheriah |
HLT-NAACL | 4 |
| 2003 | dentifying and Tracking Entity Mentions in a Maximum Entropy Framework
Abraham Ittycheriah, Lucian Vlad Lita, Nanda Kambhatla, Nicolas Nicolov, Salim Roukos, Margo Stys |
HLT-NAACL | 1 |
| 2003 | Automatic Derivation of Surface Text Patterns for a Maximum Entropy Based Question Answering System
Deepak Ravichandran, Abraham Ittycheriah, Salim Roukos |
HLT-NAACL | 2 |
| 2001 | Question Answering Using Maximum-Entropy Components
Abraham Ittycheriah, Martin Franz, Wei-Jing Zhu, Adwait Ratnaparkhi |
NAACL | 1 |
| 1999 | The IBM conversational telephony system for financial applicationsabstractWe describe our development work on a telephonebased conversational system in the domain of mutual fund transactions. This system uses several components including robust large vocabulary continuous speech recognition, natural language understanding, dialog management, and text-to-speech synthesis technologies. K. Davies, Robert E. Donovan, Mark Epstein, Martin Franz, Abraham Ittycheriah, Ea-Ee Jan, Jean-Michel LeRoux, David M. Lubensky, Chalapathy Neti, Mukund Padmanabhan, Kishore Papineni, Salim Roukos, Andrej Sakrajda, Jeffrey S. Sorensen, Borivoj Tydlitát, Todd Ward |
EUROSPEECH | 5 |
| 1999 | Detecting user speech in barge-in over prompts using speaker identification methods
Abraham Ittycheriah, Richard J. Mammone |
EUROSPEECH | 1 |
| 1999 | Acoustics-based baseform generation with pronunciation and/or phonotactic models
Bhuvana Ramabhadran, Sabine Deligne, Abraham Ittycheriah |
EUROSPEECH | 3 |
| 1998 | Time shift invariant speech recognition
Sankar Basu, Abraham Ittycheriah, Stéphane H. Maes |
ICSLP | 2 |
| 1998 | Phonological rules for enhancing acoustic enrollment of unknown words
Bhuvana Ramabhadran, Abraham Ittycheriah |
ICSLP | 2 |