Heshaam Faili

dblp:40/9125 · also Heshaam Feili, Hesham Faili · DBLP profile ↗
← Back
44ranked-venue papers
4as first author
14since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 39 · 4 first-author · 14 since 2021Databases, data management, data science and information retrieval · 4Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 MDP-GRPO: Stabilized Group Relative Policy Optimization for Multi-Constraint Instruction Following
abstract
Mohammad Mahdi Salmani-Zarchi, Zahra Rahimi, Heshaam Faili, Mohammad Javad Dousti. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mohammad Mahdi Salmani-Zarchi, Zahra Rahimi, Heshaam Faili, Mohammad Javad Dousti
ACL (1)3
2026 PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark
abstract
Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine, particularly in low-resource languages, remains underexplored. In this work, we introduce PersianMedQA, a large-scale dataset of 20,785 expert-validated multiple-choice Persian medical questions from 14 years of Iranian national medical exams, spanning 23 medical specialties and designed to evaluate LLMs in both Persian and English. We benchmark 41 state-of-the-art models, including general-purpose, Persian, and medical LLMs, in zero-shot and chain-of-thought (CoT) settings. Our results show that closed-weight general models (e.g., GPT-4.1) consistently outperform all other categories, achieving 83.09% accuracy in Persian and 80.7% in English, while Persian LLMs such as Dorna underperform significantly (e.g., 34.9% in Persian), often struggling with both instruction-following and domain reasoning. We also analyze the impact of translation, showing that while English performance is generally higher, 3-10% of questions can only be answered correctly in Persian due to cultural and clinical contextual cues that are lost in translation. Finally, we demonstrate that model size alone is insufficient for robust performance without strong domain or language adaptation. PersianMedQA provides a foundation for evaluating bilingual and culturally grounded medical reasoning in LLMs. The dataset, along with a bilingual medical dictionary, is available: https://huggingface.co/datasets/MohammadJRanjbar/PersianMedQA .
Mohammad Javad Ranjbar Kalahroodi, Amirhossein Sheikholselami, Sepehr Karimi Arpanahi, Sepideh Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery
LREC5
2026 DeepQuestion: Systematic Generation of Real-World Challenges for Evaluating LLMs Performance
abstract
While Large Language Models (LLMs) achieve near-human performance on standard benchmarks, their capabilities often fail to generalize to complex, real-world problems. To bridge this gap, we introduce DeepQuestion, a scalable, automated framework that systematically elevates the cognitive complexity of existing datasets. Grounded in Bloom's taxonomy, DeepQuestion generates (1) scenario-based problems to test the application of knowledge in noisy, realistic contexts, and (2) instruction-based prompts that require models to create new questions from a given solution path, assessing synthesis and evaluation skills. Our extensive evaluation across ten leading open-source and proprietary models reveals a stark performance decline with accuracy dropping by up to 70% as tasks ascend the cognitive hierarchy. These findings underscore that current benchmarks overestimate true reasoning abilities and highlight the critical need for cognitively diverse evaluations to guide future LLM development.
Ali Khoramfar, Ali Ramezani, Mohammad Mahdi Mohajeri, Mohammad Javad Dousti, Majid Nili Ahmadabadi, Heshaam Faili
LREC6
2025 IRUEX: A Study on Large Language Models Problem-Solving Skills in Iran's University Entrance Exam
abstract
In this paper, we present the IRUEX dataset, a novel multiple-choice educational resource specifically designed to evaluate the performance of Large Language Models (LLMs) across seven distinct categories. The dataset contains 868 Iran university entrance exam questions (Konkour) and 36,485 additional questions. Each additional question is accompanied by detailed solutions, and the dataset also includes relevant high school textbooks, providing comprehensive study material. A key feature of IRUEX is its focus on underrepresented languages, particularly assessing problem-solving skills, language proficiency, and reasoning. Our evaluation shows that GPT-4o outperforms the other LLMs tested on the IRUEX dataset. Techniques such as few-shot learning and retrieval-augmented generation (RAG) display varied effects across different categories, highlighting their unique strengths in specific areas. Additionally, a comprehensive user study classifies the errors made by LLMs into ten problem-solving ability categories. The analysis highlights that calculations and linguistic knowledge, particularly in low-resource languages, remain significant weaknesses in current LLMs. IRUEX has the potential to serve as a benchmark for evaluating the reasoning capabilities of LLMs in non-English settings, providing a foundation for improving their performance in diverse languages and contexts
Hamed Khademi Khaledi, Heshaam Faili
COLING2
2025 Matina: A Large-Scale 73B Token Persian Text Corpus
abstract
Sara Bourbour Hosseinbeigi, Fatemeh Taherinezhad, Heshaam Faili, Hamed Baghbani, Fatemeh Nadi, Mostafa Amiri. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Sara Bourbour, Fatemeh Taherinezhad, Heshaam Faili, Hamed Baghbani, Fatemeh Nadi, Mostafa Amiri
NAACL (Long Papers)3
2025 Combining replay and LoRA for continual learning in natural language understanding
Zeinab Borhanifard, Heshaam Faili, Yadollah Yaghoobzadeh
Comput. Speech Lang.2
2024 Esposito: An English-Persian Scientific Parallel Corpus for Machine Translation
abstract
Neural machine translation requires large number of parallel sentences along with in-domain parallel data to attain best results. Nevertheless, no scientific parallel corpus for English-Persian language pair is available. In this paper, a parallel corpus called Esposito is introduced, which contains 3.5 million parallel sentences in the scientific domain for English-Persian language pair. In addition, we present a manually validated scientific test set that might serve as a baseline for future studies. We show that a system trained using Esposito along with other publicly available data improves the baseline on average by 7.6 and 8.4 BLEU scores for En->Fa and Fa->En directions, respectively. Additionally, domain analysis using the 5-gram KenLM model revealed notable distinctions between our parallel corpus and the existing generic parallel corpus. This dataset will be available to the public upon the acceptance of the paper.
Mersad Esalati, Mohammad Javad Dousti, Heshaam Faili
LREC/COLING3
2024 EPOQUE: An English-Persian Quality Estimation Dataset
abstract
Translation quality estimation (QE) is an important component in real-world machine translation applications. Unfortunately, human labeled QE datasets, which play an important role in developing and assessing QE models, are only available for limited language pairs. In this paper, we present the first English-Persian QE dataset, called EPOQUE, which has manually annotated direct assessment labels. EPOQUE contains 1000 sentences translated from English to Persian and annotated by three human annotators. It is publicly available, and thus can be used as a zero-shot test set, or for other scenarios in future work. We also evaluate and report the performance of two state-of-the-art QE models, i.e., Transquest and CometKiwi, as baselines on our dataset. Furthermore, our experiments show that using a small subset of the proposed dataset containing 300 sentences to fine-tune Transquest, can improve its performance by more that 8% in terms of the Pearson correlation with a held-out test set.
Mohammed Hossein Jafari Harandi, Fatemeh Azadi, Mohammad Javad Dousti, Heshaam Faili
LREC/COLING4
2024 Matching tasks to objectives: Fine-tuning and prompt-tuning strategies for encoder-decoder pre-trained language models
Ahmad Pouramini, Heshaam Faili
Appl. Intell.2
2024 Persian offensive language detection
Emad Kebriaei, Ali Homayouni, Roghayeh Faraji, Armita Razavi, Azadeh Shakery, Heshaam Faili, Yadollah Yaghoobzadeh
Mach. Learn.6
2022 PerCQA: Persian Community Question Answering Dataset
abstract
Community Question Answering (CQA) forums provide answers to many real-life questions. These forums are trendy among machine learning researchers due to their large size. Automatic answer selection, answer ranking, question retrieval, expert finding, and fact-checking are example learning tasks performed using CQA data. This paper presents PerCQA, the first Persian dataset for CQA. This dataset contains the questions and answers crawled from the most well-known Persian forum. After data acquisition, we provide rigorous annotation guidelines in an iterative process and then the annotation of question-answer pairs in SemEvalCQA format. PerCQA contains 989 questions and 21,915 annotated answers. We make PerCQA publicly available to encourage more research in Persian CQA. We also build strong benchmarks for the task of answer selection in PerCQA by using mono- and multi-lingual pre-trained language models.
Naghme Jamali, Yadollah Yaghoobzadeh, Heshaam Faili
LREC3
2022 Cross-lingual transfer learning for relation extraction using Universal Dependencies
Nasrin Taghizadeh, Heshaam Faili
Comput. Speech Lang.2
2021 Developing the Persian Wordnet of Verbs Using Supervised Learning
abstract
Nowadays, wordnets are extensively used as a major resource in natural language processing and information retrieval tasks. Therefore, the accuracy of wordnets has a direct influence on the performance of the involved applications. This paper presents a fully-automated method for extending a previously developed Persian wordnet to cover more comprehensive and accurate verbal entries. At first, by using a bilingual dictionary, some Persian verbs are linked to Princeton WordNet synsets. A feature set related to the semantic behavior of compound verbs as the majority of Persian verbs is proposed. This feature set is employed in a supervised classification system to select the proper links for inclusion in the wordnet. We also benefit from a pre-existing Persian wordnet, FarsNet, and a similarity-based method to produce a training set. This is the largest automatically developed Persian wordnet with more than 27,000 words, 28,000 PWN synsets and 67,000 word-sense pairs that substantially outperforms the previous Persian wordnet with about 16,000 words, 22,000 PWN synsets and 38,000 word-sense pairs.
Zahra Mousavi, Heshaam Faili
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2021 Cross-lingual Adaptation Using Universal Dependencies
abstract
We describe a cross-lingual adaptation method based on syntactic parse trees obtained from the Universal Dependencies (UD), which are consistent across languages, to develop classifiers in low-resource languages. The idea of UD parsing is to capture similarities as well as idiosyncrasies among typologically different languages. In this article, we show that models trained using UD parse trees for complex NLP tasks can characterize very different languages. We study two tasks of paraphrase identification and relation extraction as case studies. Based on UD parse trees, we develop several models using tree kernels and show that these models trained on the English dataset can correctly classify data of other languages, e.g., French, Farsi, and Arabic. The proposed approach opens up avenues for exploiting UD parsing in solving similar cross-lingual tasks, which is very useful for languages for which no labeled data is available.
Nasrin Taghizadeh, Heshaam Faili
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 MNCN: A Multilingual Ngram-Based Convolutional Network for Aspect Category Detection in Online Reviews
abstract
The advent of the Internet has caused a significant growth in the number of opinions expressed about products or services on e-commerce websites. Aspect category detection, which is one of the challenging subtasks of aspect-based sentiment analysis, deals with categorizing a given review sentence into a set of predefined categories. Most of the research efforts in this field are devoted to English language reviews, while there are a large number of reviews in other languages that are left unexplored. In this paper, we propose a multilingual method to perform aspect category detection on reviews in different languages, which makes use of a deep convolutional neural network with multilingual word embeddings. To the best of our knowledge, our method is the first attempt at performing aspect category detection on multiple languages simultaneously. Empirical results on the multilingual dataset provided by SemEval workshop demonstrate the effectiveness of the proposed method1.
Erfan Ghadery, Sajad Movahedi, Heshaam Faili, Azadeh Shakery
AAAI3
2019 LICD: A Language-Independent Approach for Aspect Category Detection
Erfan Ghadery, Sajad Movahedi, Masoud Jalili Sabet, Heshaam Faili, Azadeh Shakery
ECIR (1)4
2019 On the use of word embedding for cross language plagiarism detection
abstract
Cross language plagiarism is the unacknowledged reuse of text across language pairs. It occurs if a passage of text is translated from source language to target language and no proper citation is provided. Although various methods have been developed for detection of cross language plagiarism, less attention has been paid to measure and compare their performance, especially when tackling with different types of paraphrasing through translation. In this paper, we investigate various approaches to cross language plagiarism detection. Moreover, we present a novel approach to cross language plagiarism detection using word embedding methods and explore its performance against other state-of-the-art plagiarism detection algorithms. In order to evaluate the methods, we have constructed an English-Persian bilingual plagiarism detection corpus (referred to as HAMTA-CL) comprised of seven types of obfuscation. The results show that the word embedding approach outperforms the other approaches with respect to recall when encountering heavily paraphrased passages. On the other hand, translation based approach performs well when the precision is the main consideration of the cross language plagiarism detection system.
Habibollah Asghari, Omid Fatemi, Salar Mohtaj, Heshaam Faili, Paolo Rosso
Intell. Data Anal.4
2019 RACER: accurate and efficient classification based on rule aggregation approach
Javad Basiri, Fattaneh Taghiyareh, Heshaam Faili
Neural Comput. Appl.3
2019 A learning to rank approach for cross-language information retrieval exploiting multiple translation resources
abstract
Abstract Cross-language information retrieval (CLIR), finding information in one language in response to queries expressed in another language, has attracted much attention due to the explosive growth of multilingual information in the World Wide Web. One important issue in CLIR is how to apply monolingual information retrieval (IR) methods in cross-lingual environments. Recently, learning to rank (LTR) approach has been successfully employed in different IR tasks. In this paper, we use LTR for CLIR. In order to adapt monolingual LTR techniques in CLIR and pass the barrier of language difference, we map monolingual IR features to CLIR ones using translation information extracted from different translation resources. The performance of CLIR is highly dependent on the size and quality of available bilingual resources. Effective use of available resources is especially important in low-resource language pairs. In this paper, we further propose an LTR-based method for combining translation resources in CLIR. We have studied the effectiveness of the proposed approach using different translation resources. Our results also show that LTR can be used to successfully combine different translation resources to improve the CLIR performance. In the best scenario, the LTR-based combination method improves the performance of single-resource-based CLIR method by 6% in terms of Mean Average Precision.
Hosein Azarbonyad, Azadeh Shakery, Heshaam Faili
Nat. Lang. Eng.3
2019 Converting Dependency Structure Into Persian Phrase Structure
abstract
Treebank is one of the important and useful resources in natural language processing represented in two different annotated schemas: phrase and dependency structures. There are many works that convert a phrase structure into a dependency structure and vice versa. Most of them are based that exploit the handcrafted head percolation table and argument table in predefined deterministic ways. In this article, we propose a method to convert a dependency structure into a phrase structure by enriching a trainable model of former hybrid strategy approach. By adding a classifier to the algorithm and using postprocessing modification, the quality of conversion is increased. We evaluate our method in two different languages, English and Persian, and then analyze the errors. The results of our experiments show a 46.01% reduction of error rate in English and 76.50% for Persian compared to our baseline. We build a new phrase structure treebank by converting 10,000 sentences of Persian dependency treebank into corresponding phrase structures and correcting them manually.
Mohammad Hossein Dehghan, Heshaam Faili
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2018 A neural reordering model based on phrasal dependency tree for statistical machine translation
abstract
Machine translation is an important field of research and development. Word reordering is one of the main problems in machine translation. It is an important factor of quality and efficiency of machine translations and becomes more difficult when it deals with structurally divergent language pairs. To overcome this problem, we introduce a neural reordering model, using phrasal dependency trees which depict dependency relations among contiguous non-syntactic phrases. The model makes the use of reordering rules, which are automatically learned by a probabilistic neural network classifier from a reordered phrasal dependency tree bank. The proposed model combines the power of the lexical reordering and syntactic pre-ordering models by performing long-distance reorderings. The proposed reordering model is integrated into a standard phrase-based statistical machine translation system to translate input sentences. Our method is evaluated on syntactically divergent language-pairs, English → Persian and English → German using WMT07 benchmark. The results illustrate the superiority of the proposed method in terms of BLEU, TER and LRscore on both translation tasks. On average the proposed method retrieves a significant impact on precision and recall values respect to the hierarchical, lexicalized and distortion reordering models.
Saeed Farzi, Heshaam Faili, Sahar Kianian
Intell. Data Anal.2
2017 Multiple System Combination for PersoArabic-Latin Transliteration
Nima Hemmati, Heshaam Faili, Jalal Maleki
CICLing (2)2
2017 Dimension Projection Among Languages Based on Pseudo-Relevant Documents for Query Translation
Javid Dadashkarimi, Mahsa S. Shahshahani, Amirhossein Tebbifakhr, Heshaam Faili, Azadeh Shakery
ECIR4
2017 An expectation-maximization algorithm for query translation based on pseudo-relevant documents
Javid Dadashkarimi, Azadeh Shakery, Heshaam Faili, Hamed Zamani
Inf. Process. Manag.3
2017 A bootstrapping method for development of Treebank
abstract
Using statistical approaches beside the traditional methods of natural language processing could significantly improve both the quality and performance of several natural language processing (NLP) tasks. The effective usage of these approaches is subject to the availability of the informative, accurate and detailed corpora on which the learners are trained. This article introduces a bootstrapping method for developing annotated corpora based on a complex and rich linguistically motivated elementary structure called supertag. To this end, a hybrid method for supertagging is proposed that combines both of the generative and discriminative methods of supertagging. The method was applied on a subset of Wall Street Journal (WSJ) in order to annotate its sentences with a set of linguistically motivated elementary structures of the English XTAG grammar that is using a lexicalised tree-adjoining grammar formalism. The empirical results confirm that the bootstrapping method provides a satisfactory way for annotating the English sentences with the mentioned structures. The experiments show that the method could automatically annotate about 20% of WSJ with the accuracy of F-measure about 80% of which is particularly 12% higher than the F-measure of the XTAG Treebank automatically generated from the approach proposed by Basirat and Faili [(2013). Bridge the gap between statistical and hand-crafted grammars. Computer Speech and Language, 27, 1085–1104].
Farzaneh Zarei, Ali Basirat, Heshaam Faili, M. Mirain
J. Exp. Theor. Artif. Intell.3
2017 SWIM: Stepped Weighted Shell Decomposition Influence Maximization for Large-Scale Networks
abstract
A considerable amount of research has been devoted to the proposition of scalable algorithms for influence maximization. A number of such scalable algorithms exploit the community structure of the network. Besides the community structure, real-world social networks possess a different property, known as the layer structure. In this article, we propose a method based on the layer structure to maximize the influence in huge networks. Conducting experiments on a number of real-world networks, we will show that our method outperforms the state-of-the-art algorithms by its time complexity while having similar or slightly better final influence spread. Furthermore, unlike its predecessors, our method is able to show a high entanglement between structure and dynamics by giving insight on the reason why different networks have two contrasting behaviors in their saturation. By “saturation,” we mean a state during the seed selection process after which adjoining new nodes to the initial set will have a negligible effect on increasing the influence spread. We will demonstrate that how our method can predict the saturation dynamics in the networks. This prediction can be used to identify the network structures that are more vulnerable to the fast spread of the rumors.
Ali Vardasbi, Heshaam Faili, Masoud Asadpour
ACM Trans. Inf. Syst.2
2016 Persianp: A Persian Text Processing Toolbox
Mahdi Mohseni, Javad Ghofrani, Heshaam Faili
CICLing (1)3
2016 Improving Word Alignment of Rare Words with Word Embeddings
abstract
We address the problem of inducing word alignment for language pairs by developing an unsupervised model with the capability of getting applied to other generative alignment models. We approach the task by: i)proposing a new alignment model based on the IBM alignment model 1 that uses vector representation of words, and ii)examining the use of similar source words to overcome the problem of rare source words and improving the alignments. We apply our method to English-French corpora and run the experiments with different sizes of sentence pairs. Our results show competitive performance against the baseline and in some cases improve the results up to 6.9% in terms of precision.
Masoud Jalili Sabet, Heshaam Faili, Gholamreza Haffari
COLING2
2016 Sentence alignment using local and global information
Hamed Zamani, Heshaam Faili, Azadeh Shakery
Comput. Speech Lang.2
2016 Automatic Wordnet Development for Low-Resource Languages using Cross-Lingual WSD
abstract
‎Wordnets are an effective resource for natural language processing and information retrieval‎, ‎especially for semantic processing and meaning related tasks‎. ‎So far‎, ‎wordnets have been constructed for many languages‎. ‎However‎, ‎the automatic development of wordnets for low-resource languages has not been well studied‎. ‎In this paper‎, ‎an Expectation-Maximization algorithm is used to create high quality and large scale wordnets for poor-resource languages‎. ‎The proposed method benefits from possessing cross-lingual word sense disambiguation and develops a wordnet by only using a bi-lingual dictionary and a mono-lingual corpus‎. ‎The proposed method has been executed with Persian language and the resulting wordnet has been evaluated through several experiments‎. ‎The results show that the induced wordnet has a precision score of 90% and a recall score of 35%‎.
Nasrin Taghizadeh, Heshaam Faili
J. Artif. Intell. Res.2
2016 A statistical model for grammar mapping
abstract
Abstract The two main classes of grammars are (a)hand-crafted grammars, which are developed by language experts, and (b)data-driven grammars, which are extracted from annotated corpora. This paper introduces a statistical method for mapping the elementary structures of adata-driven grammaronto the elementary structures of ahand-crafted grammarin order to combine their advantages. The idea is employed in the context ofLexicalized Tree-Adjoining Grammars(LTAG) and tested on two LTAGs of English: the hand-crafted LTAG developed in the XTAG project, and the data-driven LTAG, which is automatically extracted from the Penn Treebank and used by the MICA parser. We propose a statistical model for mapping any elementary tree sequence of the MICA grammar onto a proper elementary tree sequence of the XTAG grammar. The model has been tested on three subsets of the WSJ corpus that have average lengths of 10, 16, and 18 words, respectively. The experimental results show that full-parse trees with averageF1-scores of 72.49, 64.80, and 62.30 points could be built from 94.97%, 96.01%, and 90.25% of the XTAG elementary tree sequences assigned to the subsets, respectively. Moreover, by reducing the amount ofsyntactic lexical ambiguityof sentences, the proposed model significantly improves the efficiency of parsing in the XTAG system.
Ali Basirat, Heshaam Faili, Joakim Nivre
Nat. Lang. Eng.2
2015 A swarm-inspired re-ranker system for statistical machine translation
Saeed Farzi, Heshaam Faili
Comput. Speech Lang.2
2015 Using decision tree to hybrid morphology generation of Persian verb for English-Persian translation
Alireza Mahmoudi, Heshaam Faili
Comput. Speech Lang.2
2015 A syntactically informed reordering model for statistical machine translation
abstract
Word reordering is one of the challengeable problems of machine translation. It is an important factor of quality and efficiency of machine translation systems. In this paper, we introduce a novel reordering model based on an innovative structure, named, phrasal dependency tree. The phrasal dependency tree is a modern syntactic structure which is based on dependency relationships between contiguous non-syntactic phrases. The proposed model integrates syntactical and statistical information in the context of log-linear model aimed at dealing with the reordering problems. It benefits from phrase dependencies, translation directions (orientations) and translation discontinuity between translated phrases. In comparison with well-known and popular reordering models such as distortion, lexicalised and hierarchical models, the experimental study demonstrates the superiority of our model in terms of translation quality. Performance is evaluated for Persian → English and English → German translation tasks using Tehran parallel corpus and WMT07 benchmarks, respectively. The results report 1.54/1.7 and 1.98/3.01 point improvements over the baseline in terms of BLEU/TER metrics on Persian → English and German → English translation tasks, respectively. On average our model retrieved a significant impact on precision with comparable recall value with respect to the lexicalised and distortion models.
Saeed Farzi, Heshaam Faili, Shahram Khadivi
J. Exp. Theor. Artif. Intell.2
2014 A Probabilistic Approach to Persian Ezafe Recognition
abstract
In this paper, we investigate the problem of Ezafe recognition in Persian language. Ezafe is an unstressed vowel that is usually not written, but is intelligently recognized and pronounced by human. Ezafe marker can be placed into noun phrases, adjective phrases and some prepositional phrases linking the head and modifiers. Ezafe recognition in Persian is indeed a homograph disambiguation problem, which is a useful task for some language applications in Persian like TTS. In this paper, Part of Speech tags augmented by Ezafe marker (POSE) have been used to train a probabilistic model for Ezafe recognition. In order to build this model, a ten million word tagged corpus was used for training the system. For building the probabilistic model, three different approaches were used; Maximum Entropy POSE tagger, Conditional Random Fields (CRF) POSE tagger and also a statistical machine translation approach based on parallel corpus. It is shown that comparing to previous works, the use of CRF POSE tagger can achieve outstanding results.
Habibollah Asghari, Jalal Maleki, Heshaam Faili
EACL3
2014 Semi-supervised word polarity identification in resource-lean languages
Iman Dehdarbehbahani, Azadeh Shakery, Heshaam Faili
Neural Networks3
2013 Bridge the gap between statistical and hand-crafted grammars
Ali Basirat, Heshaam Faili
Comput. Speech Lang.2
2013 Grammatical and context-sensitive error correction using a statistical machine translation framework
abstract
SUMMARY Producing electronic rather than paper documents has considerable benefits such as easier organizing and data management. Therefore, existence of automatic writing assistance tools such as spell and grammar checker/correctors can increase the quality of electronic texts by removing noise and correcting the erroneous sentences. Different kinds of errors in a text can be categorized into spelling, grammatical and real‐word errors. In this article, we present a language‐independent approach based on a statistical machine translation framework to develop a proofreading tool, which detects grammatical errors as well as context‐sensitive spelling mistakes (real‐word errors). A hybrid model for grammar checking is suggested by combining the mentioned approach with an existing rule‐based grammar checker. Experimental results on both English and Persian languages indicate that the proposed statistical method and the rule‐based grammar checker are complementary in detecting and correcting syntactic errors. The results of the hybrid grammar checker, applied to some English texts, show an improvement of about 24% with respect to the recall metric with almost similar value for precision. Experiments on real‐world data set show that state‐of‐the‐art results are achieved for grammar checking and context‐sensitive spell checking for Persian language. Copyright © 2012 John Wiley & Sons, Ltd.
Nava Ehsan, Heshaam Faili
Softw. Pract. Exp.2
2012 Applying Sentiment and Social Network Analysis in User Modeling
Mohammadreza Shams, Mohammad Taghi Saffar, Azadeh Shakery, Heshaam Faili
CICLing (1)4
2011 An Unsupervised Approach for Linking Automatically Extracted and Manually Crafted LTAGs
Heshaam Faili, Ali Basirat
CICLing (1)1
2011 TEP: Tehran English-Persian Parallel Corpus
Mohammad Taher Pilehvar, Heshaam Faili, Abdol Hamid Pilevar
CICLing (2)2
2006 Hisory-Based Inside-Outside Algorithm
Heshaam Faili, Gholamreza Ghassem-Sani
ECAI1
2006 Unsupervised grammar induction using history based approach
Heshaam Faili, Gholamreza Ghassem-Sani
Comput. Speech Lang.1
2004 An Application of Lexicalized Grammars in English-Persian Translation
Heshaam Faili, Gholamreza Ghassem-Sani
ECAI1