VLDB 2026 Research / reviewers in the wild / expert
Shahram Khadivi
dblp:03/5199
· DBLP profile ↗
28ranked-venue papers
5as first author
5since 2021 · last 2024
0000-0001-5499-6542ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Efficient and distributed learning · 37% Machine translation · 18% Transfer learning and domain adaptation · 17% |
Topics — the 15 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
large language model fine-tuning |
0.8 | 1 | 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language Model · EMNLP 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language Model · EMNLP 2024 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.8 | 1 | 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language Model · EMNLP 2024 |
Machine learning › Efficient and distributed learning › model compression
quantization |
0.8 | 1 | 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language Model · EMNLP 2024 |
Machine learning › Representation and self-supervised learning › representation learning › invariant representation learning
domain-invariant representation |
0.7 | 1 | 2023 | Energy-based Self-Training and Normalization for Unsupervised Domain Adaptation · ICCV 2023 |
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation |
0.7 | 1 | 2023 | Energy-based Self-Training and Normalization for Unsupervised Domain Adaptation · ICCV 2023 |
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer |
0.4 | 1 | 2019 | Pivot-based Transfer Learning for Neural Machine Translation between Non-English Languages · EMNLP/IJCNLP (1) 2019 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
hybrid search |
0.3 | 1 | 2017 | Neural Machine Translation Leveraging Phrase-based Models in a Hybrid Search · EMNLP 2017 |
Natural language and speech › Machine translation
neural machine translation |
0.3 | 1 | 2017 | Neural Machine Translation Leveraging Phrase-based Models in a Hybrid Search · EMNLP 2017 |
Natural language and speech › Machine translation
computer-assisted translation |
0.1 | 2 | 2008 | Integration of Speech Recognition and Machine Translation in Computer-Assisted Translation · IEEE Trans. Speech Audio Process. 2008 Integration of Speech to Computer-Assisted Translation Using Finite-State Automata · ACL 2006 |
Natural language and speech › Machine translation › statistical machine translation
phrase-based translation |
0.1 | 1 | 2017 | Neural Machine Translation Leveraging Phrase-based Models in a Hybrid Search · EMNLP 2017 |
Natural language and speech › Machine translation
statistical machine translation |
0.1 | 1 | 2017 | Neural Machine Translation Leveraging Phrase-based Models in a Hybrid Search · EMNLP 2017 |
Natural language and speech › Machine translation › statistical machine translation
word alignment |
0.1 | 1 | 2007 | A Sequence Alignment Model Based on the Averaged Perceptron · EMNLP-CoNLL 2007 |
Natural language and speech › Machine translation
speech translation |
0.1 | 1 | 2006 | Integration of Speech to Computer-Assisted Translation Using Finite-State Automata · ACL 2006 |
Natural language and speech › Machine translation › speech translation
speech-to-speech translation |
0.0 | 1 | 2008 | Integration of Speech Recognition and Machine Translation in Computer-Assisted Translation · IEEE Trans. Speech Audio Process. 2008 |
Methods — techniques the papers use, named apart from their topics
quantization-aware initialization · 0.8LoRA · 0.8self-training · 0.7energy-based learning · 0.7energy normalization · 0.7pivot translation · 0.4neural machine translation · 0.4phrase-based features · 0.3beam search · 0.3n-best list rescoring · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language ModelabstractMemory-efficient finetuning of large language models (LLMs) has recently attracted huge attention with the increasing size of LLMs, primarily due to the constraints posed by GPU memory limitations and the effectiveness of these methods compared to full finetuning.Despite the advancements, current strategies for memory-efficient finetuning, such as QLoRA, exhibit inconsistent performance across diverse bit-width quantizations and multifaceted tasks.This inconsistency largely stems from the detrimental impact of the quantization process on preserved knowledge, leading to catastrophic forgetting and undermining the utilization of pretrained models for finetuning purposes.In this work, we introduce a novel quantization framework named ApiQ, designed to restore the lost information from quantization by concurrently initializing the LoRA components and quantizing the weights of LLMs.This approach ensures the maintenance of the original LLM's activation precision while mitigating the error propagation from shallower into deeper layers.Through comprehensive evaluations conducted on a spectrum of language tasks with various LLMs, ApiQ demonstrably minimizes activation error during quantization.Consequently, it consistently achieves superior finetuning results across various bit-widths.Notably, one can even finetune a 2-bit Llama-2-70b with ApiQ on a single NVIDIA A100-80GB GPU without any memory-saving techniques, and achieve promising results. Baohao Liao, Christian Herold, Shahram Khadivi, Christof Monz |
EMNLP | 3 |
| 2023 | Probabilistic Robustness for Data FilteringabstractWe introduce our probabilistic robustness rewarded data optimization (PRoDO) approach as a framework to enhance the model's generalization power by selecting training data that optimizes our probabilistic robustness metrics.We use proximal policy optimization (PPO) reinforcement learning to approximately solve the computationally intractable training subset selection problem.The PPO's reward is defined as our (α, ϵ, γ)-Robustness that measures performance consistency over multiple domains by simulating unknown test sets in real-world scenarios using a leaving-one-out strategy.We demonstrate that our PRoDO effectively filters data that lead to significantly higher prediction accuracy and robustness on unknown-domain test sets.Our experiments achieve up to +17.2% increase of accuracy (+25.5% relatively) in sentiment analysis, and -28.05 decrease of perplexity (-32.1% relatively) in language modeling.In addition, our probabilistic (α, ϵ, γ)-Robustness definition serves as an evaluation metric with higher levels of agreement with human annotations than typical performance-based metrics. Yu Yu 0004, Abdul Rafae Khan, Shahram Khadivi, Jia Xu 0004 |
EACL | 3 |
| 2023 | Energy-based Self-Training and Normalization for Unsupervised Domain AdaptationabstractWe propose an Unsupervised Domain Adaptation (UDA) method by making use of Energy-Based Learning (EBL) and demonstrate 1. EBL can be used to improve the instance selection for a self-training task on the unlabelled target domain, and 2. alignment and normalizing energy scores can learn domain-invariant representations. For the former, we show that an energy-based selection criterion can be used to model instance selections by mimicking the joint distribution between data and predictions in the target domain. As per learning domain invariant representations, we show that stable domain alignment can be achieved by a combined energy alignment and an energy normalization process. We implement our method in consistent with the vision-transformer (ViT) backbone and show that our proposed method can outperform state-of-the-art ViT based UDA methods on diverse benchmarks (DomainNet, Office-Home, and VISDA2017). Samitha Herath, Basura Fernando, Ehsan Abbasnejad, Munawar Hayat, Shahram Khadivi, Mehrtash Harandi, Seyed Hamid Rezatofighi, Gholamreza Haffari |
ICCV | 5 |
| 2023 | LenM: Improving Low-Resource Neural Machine Translation Using Target Length Modeling
Mohammad Mahdi Mahsuli, Shahram Khadivi, Mohammad Mehdi Homayounpour |
Neural Process. Lett. | 2 |
| 2022 | Can Data Diversity Enhance Learning Generalization?abstractThis paper introduces our Diversity Advanced Actor-Critic reinforcement learning (A2C) framework (DAAC) to improve the generalization and accuracy of Natural Language Processing (NLP). We show that the diversification of training samples alleviates overfitting and improves model generalization and accuracy. We quantify diversity on a set of samples using the max dispersion, convex hull volume, and graph entropy based on sentence embeddings in high-dimensional metric space. We also introduce A2C to select such a diversified training subset efficiently. Our experiments achieve up to +23.8 accuracy increase (38.0% relatively) in sentiment analysis, -44.7 perplexity decrease (37.9% relatively) in language modeling, and consistent improvements in named entity recognition over various domains. In particular, our method outperforms both domain adaptation and generalization baselines without using any target domain knowledge. Yu Yu 0004, Shahram Khadivi, Jia Xu 0004 |
COLING | 2 |
| 2020 | Kernel compositional embedding and its application in linguistic structured data classification
Hamed Ganji, Mohammad Mehdi Ebadzadeh, Shahram Khadivi |
Knowl. Based Syst. | 3 |
| 2020 | Matching Graph, a Method for Extracting Parallel Information from Comparable CorporaabstractComparable corpora are valuable alternatives for the expensive parallel corpora. They comprise informative parallel fragments that are useful resources for different natural language processing tasks. In this work, a generative model is proposed for efficient extraction of parallel fragments from a pair of comparable documents. The core of the proposed model is a graph called the Matching Graph. The ability of the Matching Graph to be trained on a small initial seed makes it a proper model for language pairs suffering from the scarce resource problem. Experiments show that the Matching Graph performs significantly better than other recently published models. According to the experiments on English-Persian and Arabic-Persian language pairs, the extracted parallel fragments can be used instead of parallel data for training statistical machine translation systems. Results reveal that the extracted fragments in the best case are able to retrieve about 90% of the information of a statistical machine translation system that is trained on a parallel corpus. Moreover, it is shown that using the extracted fragments as additional information for training statistical machine translation systems leads to an improvement of about 2% for English-Persian and about 1% for Arabic-Persian translation on BLEU score. Somayeh Bakhshaei, Reza Safabakhsh, Shahram Khadivi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2019 | Pivot-based Transfer Learning for Neural Machine Translation between Non-English LanguagesabstractYunsu Kim, Petre Petrov, Pavel Petrushkov, Shahram Khadivi, Hermann Ney. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Yunsu Kim 0001, Petre Petrov, Pavel Petrushkov, Shahram Khadivi, Hermann Ney |
EMNLP/IJCNLP (1) | 4 |
| 2019 | Extracting parallel fragments from comparable documents using a generative model
Somayeh Bakhshaei, Reza Safabakhsh, Shahram Khadivi |
Comput. Speech Lang. | 3 |
| 2019 | Support vector-based fuzzy classifier with adaptive kernel
Hamed Ganji, Shahram Khadivi, Mohammad Mehdi Ebadzadeh |
Neural Comput. Appl. | 2 |
| 2017 | Neural Machine Translation Leveraging Phrase-based Models in a Hybrid SearchabstractIn this paper, we introduce a hybrid search for attention-based neural machine translation (NMT).A target phrase learned with statistical MT models extends a hypothesis in the NMT beam search when the attention of the NMT model focuses on the source words translated by this phrase.Phrases added in this way are scored with the NMT model, but also with SMT features including phrase-level translation probabilities and a target language model.Experimental results on German→English news domain and English→Russian ecommerce domain translation tasks show that using phrase-based models in NMT search improves MT quality by up to 2.3% BLEU absolute as compared to a strong NMT baseline. Leonard Dahlmann, Evgeny Matusov, Pavel Petrushkov, Shahram Khadivi |
EMNLP | 4 |
| 2017 | Neural and Statistical Methods for Leveraging Meta-information in Machine Translation
Shahram Khadivi, Patrick Wilken, Leonard Dahlmann, Evgeny Matusov |
MTSummit (1) | 1 |
| 2016 | Phrase-boundary model for statistical machine translation
Shahram Salami, Mehrnoush Shamsfard, Shahram Khadivi |
Comput. Speech Lang. | 3 |
| 2015 | Improved search strategy for interactive predictions in computer-assisted translation
Fatemeh Azadi, Shahram Khadivi |
MTSummit | 2 |
| 2015 | A syntactically informed reordering model for statistical machine translationabstractWord reordering is one of the challengeable problems of machine translation. It is an important factor of quality and efficiency of machine translation systems. In this paper, we introduce a novel reordering model based on an innovative structure, named, phrasal dependency tree. The phrasal dependency tree is a modern syntactic structure which is based on dependency relationships between contiguous non-syntactic phrases. The proposed model integrates syntactical and statistical information in the context of log-linear model aimed at dealing with the reordering problems. It benefits from phrase dependencies, translation directions (orientations) and translation discontinuity between translated phrases. In comparison with well-known and popular reordering models such as distortion, lexicalised and hierarchical models, the experimental study demonstrates the superiority of our model in terms of translation quality. Performance is evaluated for Persian → English and English → German translation tasks using Tehran parallel corpus and WMT07 benchmarks, respectively. The results report 1.54/1.7 and 1.98/3.01 point improvements over the baseline in terms of BLEU/TER metrics on Persian → English and German → English translation tasks, respectively. On average our model retrieved a significant impact on precision with comparable recall value with respect to the lexicalised and distortion models. Saeed Farzi, Heshaam Faili, Shahram Khadivi |
J. Exp. Theor. Artif. Intell. | 3 |
| 2014 | An ant colony optimization method to detect communities in social networksabstractCommunity detection is an important task in social network analysis. It aims to partition the network into clusters so that interactions among members within a cluster are considerably more frequent than that across clusters. A typical instantiation is to maximize the modularity of clusters which is a NP-hard problem, and thus, heuristic and meta-heuristic algorithms are employed as approximation. We present a novel divisive algorithm based on ant colony optimization to detect hierarchical community structure by maximizing the modularity. Our algorithm splits the network into two local communities iteratively and incorporates both heuristic information and pheromone trails. Experimental results on a set of synthetic benchmarks and real-world networks verified that our algorithm is highly effective for hierarchical community structure detection. Saeed Haji Seyed Javadi, Shahram Khadivi, Mohammad Ebrahim Shiri, Jia Xu 0004 |
ASONAM | 2 |
| 2013 | Meta-level Statistical Machine Translation
Sajad Ebrahimi 0002, Kourosh Meshgi, Shahram Khadivi, Mohammad Ebrahim Shiri |
IJCNLP | 3 |
| 2012 | A Holistic Approach to Bilingual Sentence Fragment Extraction from Comparable Corpora
Mahdi Khademian, Kaveh Taghipour, Saab Mansour, Shahram Khadivi |
LREC | 4 |
| 2011 | Parallel Corpus Refinement as an Outlier Detection Algorithm
Kaveh Taghipour, Shahram Khadivi, Jia Xu 0004 |
MTSummit | 2 |
| 2009 | Recent advances in SRI'S IraqCommTM Iraqi Arabic-English speech-to-speech translation systemabstractWe summarize recent progress on SRI's IraqCommtrade Iraqi Arabic-English two-way speech-to-speech translation system. In the past year we made substantial developments in our speech recognition and machine translation technology, leading to significant improvements in both accuracy and speed of the IraqComm system. On the 2008 NIST-evaluation dataset our twoway speech-to-text (S2T) system achieved 6% to 8% absolute improvement in BLEU in both directions, compared to our previous year system. Murat Akbacak, Horacio Franco, Michael W. Frandsen, Sasa Hasan, Huda Jameel, Andreas Kathol, Shahram Khadivi, Arindam Mandal, Saab Mansour, Kristin Precoda, Colleen Richey, Dimitra Vergyri, Wen Wang 0001, Jing Zheng 0001 |
ICASSP | 7 |
| 2009 | Statistical Approaches to Computer-Assisted TranslationabstractCurrent machine translation (MT) systems are still not perfect. In practice, the output from these systems needs to be edited to correct errors. A way of increasing the productivity of the whole translation process (MT plus human work) is to incorporate the human correction activities within the translation process itself, thereby shifting the MT paradigm to that of computer-assisted translation. This model entails an iterative process in which the human translator activity is included in the loop: In each iteration, a prefix of the translation is validated (accepted or amended) by the human and the system computes its best (or n-best) translation suffix hypothesis to complete this prefix. A successful framework for MT is the so-called statistical (or pattern recognition) framework. Interestingly, within this framework, the adaptation of MT systems to the interactive scenario affects mainly the search process, allowing a great reuse of successful techniques and models. In this article, alignment templates, phrase-based models, and stochastic finite-state transducers are used to develop computer-assisted translation systems. These systems were assessed in a European project (TransType2) in two real tasks: The translation of printer manuals; manuals and the translation of the Bulletin of the European Union. In each task, the following three pairs of languages were involved (in both translation directions): English-Spanish, English-German, and English-French. Sergio Barrachina 0001, Oliver Bender, Francisco Casacuberta, Jorge Civera, Elsa Cubel, Shahram Khadivi, Antonio L. Lagarda, Hermann Ney, Jesús Tomás, Enrique Vidal 0001, Juan Miguel Vilar |
Comput. Linguistics | 6 |
| 2008 | Integration of Speech Recognition and Machine Translation in Computer-Assisted TranslationabstractParallel integration of automatic speech recognition (ASR) models and statistical machine translation (MT) models is an unexplored research area in comparison to the large amount of works done on integrating them in series, i.e., speech-to-speech translation. Parallel integration of these models is possible when we have access to the speech of a target language text and to its corresponding source language text, like a computer-assisted translation system. To our knowledge, only a few methods for integrating ASR models with MT models in parallel have been studied. In this paper, we systematically study a number of different translation models in the context of theN-best list rescoring. As an alternative to theN-best list rescoring, we use ASR word graphs in order to arrive at a tighter integration of ASR and MT models. The experiments are carried out on two tasks: English-to-German with an ASR vocabulary size of 17 K words, and Spanish-to-English with an ASR vocabulary of 58 K words. For the best method, the MT models reduce the ASR word error rate by a relative of 18% and 29% on the 17 K and the 58 K tasks, respectively. Shahram Khadivi, Hermann Ney |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | The RWTH Arabic-to-English spoken language translation systemabstractWe present the RWTH phrase-based statistical machine translation system designed for the translation of Arabic speech into English text. This system was used in the Global Autonomous Language Exploitation (GALE) Go/No-Go Translation Evaluation 2007. Using a two-pass approach, we first generate n-best translation candidates and then rerank these candidates using additional models. We give a short review of the decoder as well as of the models used in both passes. We stress the difficulties of spoken language translation, i.e. how to combine the recognition and translation systems and how to compensate for missing punctuation. In addition, we cover our work on domain adaptation for the applied language models. We present translation results for the official GALE 2006 evaluation set and the GALE 2007 development set. Oliver Bender, Evgeny Matusov, Stefan Hahn, Sasa Hasan, Shahram Khadivi, Hermann Ney |
ASRU | 5 |
| 2007 | A Sequence Alignment Model Based on the Averaged Perceptron
Dayne Freitag, Shahram Khadivi |
EMNLP-CoNLL | 2 |
| 2006 | Integration of Speech to Computer-Assisted Translation Using Finite-State Automata
Shahram Khadivi, Richard Zens, Hermann Ney |
ACL | 1 |
| 2006 | A Flexible Architecture for CAT Applications
Sasa Hasan, Shahram Khadivi, Richard Zens, Hermann Ney |
EAMT | 2 |
| 2005 | Automatic text dictation in computer-assisted translationabstractIn this paper, we study the incorporation of statistical machine translation models to automatic speech recognition models in the framework of computer-assisted translation. The system is given a source language text to be translated and it shows the source text to the human translator to translate it orally. The system captures the user speech which is the dictation of the target language sentence. Since the system has simultaneous access to the source language text and the speech signal of the target language text, it is possible to improve the speech recognition accuracy by incorporating the statistical machine translation models. We show that statistical translation models have a high impact on improving the speech recognition results. Using these models, we achieve a relative word error rate reduction of 17%. 1. Shahram Khadivi, András Zolnay, Hermann Ney |
INTERSPEECH | 1 |
| 2005 | Automatic Filtering of Bilingual Corpora for Statistical Machine Translation
Shahram Khadivi, Hermann Ney |
NLDB | 1 |