Azadeh Shakery

dblp:50/2756 · DBLP profile ↗
← Back
60ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-1799-8340ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 40 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 24 · 1 first-author · 7 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 PersianMedQA: Evaluating Large Language Models on a Persian-English Bilingual Medical Question Answering Benchmark
abstract
Large Language Models (LLMs) have achieved remarkable performance on a wide range of Natural Language Processing (NLP) benchmarks, often surpassing human-level accuracy. However, their reliability in high-stakes domains such as medicine, particularly in low-resource languages, remains underexplored. In this work, we introduce PersianMedQA, a large-scale dataset of 20,785 expert-validated multiple-choice Persian medical questions from 14 years of Iranian national medical exams, spanning 23 medical specialties and designed to evaluate LLMs in both Persian and English. We benchmark 41 state-of-the-art models, including general-purpose, Persian, and medical LLMs, in zero-shot and chain-of-thought (CoT) settings. Our results show that closed-weight general models (e.g., GPT-4.1) consistently outperform all other categories, achieving 83.09% accuracy in Persian and 80.7% in English, while Persian LLMs such as Dorna underperform significantly (e.g., 34.9% in Persian), often struggling with both instruction-following and domain reasoning. We also analyze the impact of translation, showing that while English performance is generally higher, 3-10% of questions can only be answered correctly in Persian due to cultural and clinical contextual cues that are lost in translation. Finally, we demonstrate that model size alone is insufficient for robust performance without strong domain or language adaptation. PersianMedQA provides a foundation for evaluating bilingual and culturally grounded medical reasoning in LLMs. The dataset, along with a bilingual medical dictionary, is available: https://huggingface.co/datasets/MohammadJRanjbar/PersianMedQA .
Mohammad Javad Ranjbar Kalahroodi, Amirhossein Sheikholselami, Sepehr Karimi Arpanahi, Sepideh Ranjbar Kalahroodi, Heshaam Faili, Azadeh Shakery
LREC6
2025 PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian
abstract
Erfan Moosavi Monazzah, Vahid Rahimzadeh, Yadollah Yaghoobzadeh, Azadeh Shakery, Mohammad Taher Pilehvar. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Erfan Moosavi Monazzah, Vahid Rahimzadeh, Yadollah Yaghoobzadeh, Azadeh Shakery, Mohammad Taher Pilehvar
NAACL (Long Papers)4
2024 Entity-centric multi-domain transformer for improving generalization in fake news detection
Parisa Bazmi, Masoud Asadpour, Azadeh Shakery, Abbas Maazallahi
Inf. Process. Manag.3
2024 Persian offensive language detection
Emad Kebriaei, Ali Homayouni, Roghayeh Faraji, Armita Razavi, Azadeh Shakery, Heshaam Faili, Yadollah Yaghoobzadeh
Mach. Learn.5
2023 $\Lambda$-DARTS: Mitigating Performance Collapse by Harmonizing Operation Selection among Cells
Sajad Movahedi, Melika Adabinejad, Ayyoob Imani, Arezou Keshavarz, Mostafa Dehghani 0001, Azadeh Shakery, Babak Nadjar Araabi
ICLR6
2023 Multi-view co-attention network for fake news detection by modeling topic-specific user and news source credibility
Parisa Bazmi, Masoud Asadpour, Azadeh Shakery
Inf. Process. Manag.3
2023 A deep learning-based expert finding method to retrieve agile software teams from CQAs
Peyman Rostami, Azadeh Shakery
Inf. Process. Manag.2
2023 Enhancement of Twitter event detection using news streams
abstract
Abstract A new framework for improving event detection is proposed that employs joint information in news media content and social networks, such as Twitter, to leverage detailed coverage of news media and the timeliness of social media. Specifically, a short text clustering method is employed to detect events from tweets, then the language model representations of the detected events are expanded using another set of events obtained from news articles published simultaneously. The expanded representations of events are employed as a new initialization of the clustering method to run another iteration and consequently enhance the event detection results. The proposed framework is evaluated using two datasets: a tweet dataset with event labels and a news dataset containing news articles published during the same time interval as the tweets. Experimental results show that the proposed framework improves the event detection results in terms of F 1 measure compared to the results obtained from tweets only.
Samaneh Karimi, Azadeh Shakery, Rakesh M. Verma
Nat. Lang. Eng.2
2022 A Learning to rank framework based on cross-lingual loss function for cross-lingual information retrieval
Elham Ghanbari, Azadeh Shakery
Appl. Intell.2
2021 ARMAN: Pre-training with Semantically Selecting and Reordering of Sentences for Persian Abstractive Summarization
abstract
Abstractive text summarization is one of the areas influenced by the emergence of pre-trained language models.Current pre-training works in abstractive summarization give more points to the summaries with more words in common with the main text and pay less attention to the semantic similarity between generated sentences and the original document.We propose ARMAN, a Transformer-based encoderdecoder model pre-trained with three novel objectives to address this issue.In ARMAN, salient sentences from a document are selected according to a modified semantic score to be masked and form a pseudo summary.To summarize more accurately and similar to human writing patterns, we applied modified sentence reordering.We evaluated our proposed models on six downstream Persian summarization tasks.Experimental results show that our proposed model achieves state-of-the-art performance on all six summarization tasks measured by ROUGE and BERTScore.Our models also outperform prior works in textual entailment, question paraphrasing, and multiple choice question answering.Finally, we established a human evaluation and show that using the semantic score significantly improves summarization results.
Alireza Salemi, Emad Kebriaei, Ghazal Neisi Minaei, Azadeh Shakery
EMNLP (1)4
2020 Distilling Knowledge for Fast Retrieval-based Chat-bots
abstract
Response retrieval is a subset of neural ranking in which a model selects a suitable response from a set of candidates given a conversation history. Retrieval-based chat-bots are typically employed in information seeking conversational systems such as customer support agents. To make pairwise comparisons between a conversation history and a candidate response, two approaches are common: cross-encoders performing full self-attention over the pair and bi-encoders encoding the pair separately. The former gives better prediction quality but is too slow for practical use. In this paper, we propose a new cross-encoder architecture and transfer knowledge from this model to a bi-encoder model using distillation. This effectively boosts bi-encoder performance at no cost during inference time. We perform a detailed analysis of this approach on three response retrieval datasets.
Amir Vakili, Kamyar Ghajar, Azadeh Shakery
SIGIR3
2020 Swash: A collective personal name matching framework
Mohsen Raeesi, Masoud Asadpour, Azadeh Shakery
Expert Syst. Appl.3
2020 An axiomatic approach to corpus-based cross-language information retrieval
Razieh Rahimi, Ali Montazeralghaem, Azadeh Shakery
Inf. Retr. J.3
2019 MNCN: A Multilingual Ngram-Based Convolutional Network for Aspect Category Detection in Online Reviews
abstract
The advent of the Internet has caused a significant growth in the number of opinions expressed about products or services on e-commerce websites. Aspect category detection, which is one of the challenging subtasks of aspect-based sentiment analysis, deals with categorizing a given review sentence into a set of predefined categories. Most of the research efforts in this field are devoted to English language reviews, while there are a large number of reviews in other languages that are left unexplored. In this paper, we propose a multilingual method to perform aspect category detection on reviews in different languages, which makes use of a deep convolutional neural network with multilingual word embeddings. To the best of our knowledge, our method is the first attempt at performing aspect category detection on multiple languages simultaneously. Empirical results on the multilingual dataset provided by SemEval workshop demonstrate the effectiveness of the proposed method1.
Erfan Ghadery, Sajad Movahedi, Heshaam Faili, Azadeh Shakery
AAAI4
2019 LICD: A Language-Independent Approach for Aspect Category Detection
Erfan Ghadery, Sajad Movahedi, Masoud Jalili Sabet, Heshaam Faili, Azadeh Shakery
ECIR (1)5
2019 An Axiomatic Study of Query Terms Order in Ad-Hoc Retrieval
Ayyoob Imani, Amir Vakili, Ali Montazeralghaem, Azadeh Shakery
ECIR (2)4
2019 Deep Neural Networks for Query Expansion Using Word Embeddings
Ayyoob Imani, Amir Vakili, Ali Montazeralghaem, Azadeh Shakery
ECIR (2)4
2019 ERR.Rank: An algorithm based on learning to rank for direct optimization of Expected Reciprocal Rank
Elham Ghanbari, Azadeh Shakery
Appl. Intell.2
2019 Query-dependent learning to rank for cross-lingual information retrieval
Elham Ghanbari, Azadeh Shakery
Knowl. Inf. Syst.2
2019 A learning to rank approach for cross-language information retrieval exploiting multiple translation resources
abstract
Abstract Cross-language information retrieval (CLIR), finding information in one language in response to queries expressed in another language, has attracted much attention due to the explosive growth of multilingual information in the World Wide Web. One important issue in CLIR is how to apply monolingual information retrieval (IR) methods in cross-lingual environments. Recently, learning to rank (LTR) approach has been successfully employed in different IR tasks. In this paper, we use LTR for CLIR. In order to adapt monolingual LTR techniques in CLIR and pass the barrier of language difference, we map monolingual IR features to CLIR ones using translation information extracted from different translation resources. The performance of CLIR is highly dependent on the size and quality of available bilingual resources. Effective use of available resources is especially important in low-resource language pairs. In this paper, we further propose an LTR-based method for combining translation resources in CLIR. We have studied the effectiveness of the proposed approach using different translation resources. Our results also show that LTR can be used to successfully combine different translation resources to improve the CLIR performance. In the best scenario, the LTR-based combination method improves the performance of single-resource-based CLIR method by 6% in terms of Mean Average Precision.
Hosein Azarbonyad, Azadeh Shakery, Heshaam Faili
Nat. Lang. Eng.2
2018 Towards a Unified Supervised Approach for Ranking Triples of Type-Like Relations
Mahsa S. Shahshahani, Faegheh Hasibi, Hamed Zamani, Azadeh Shakery
ECIR4
2018 Theoretical Analysis of Interdependent Constraints in Pseudo-Relevance Feedback
abstract
Axiomatic analysis is a well-defined theoretical framework for analytical evaluation of information retrieval models. The current studies in axiomatic analysis implicitly assume that the constraints (axioms) are independent. In this paper, we revisit this assumption and hypothesize that there might be interdependence relationships between the existing constraints. As a preliminary study, we focus on the pseudo-relevance feedback (PRF) models that have been theoretically studied using the axiomatic analysis approach. In this paper, we introduce two novel interdependent PRF constraints which emphasize on the effect of existing constraints on each other. We further modify two state-of-the-art PRF models, log-logistic and relevance models, in order to satisfy the proposed constraints. Experiments on three TREC newswire and web collections demonstrate that the proposed modifications significantly outperform the baselines, in all cases.
Ali Montazeralghaem, Hamed Zamani, Azadeh Shakery
SIGIR3
2018 Expert finding by the Dempster-Shafer theory for evidence combination
abstract
Abstract The expertise of human experts can be formally extracted from their written documents, research projects, and everyday activities. The process whereby experts are recognized according to their activities is called expert finding. In this paper, we propose an approach to identify the experts in a given field according to the content of 3 easily accessible sources of information: (a) “Publications,” (b) “Social interactions,” and (c) “Scientometric information.” We employed the Dempster‐Shafer theory to combine the results obtained from individual sources to find a final unified ranking. In Digital Bibliography & Library Project standard data, it is shown that the Dempster‐Shafer combination creates a desired synergy between 2 bodies of knowledge, which improves the precision of the top‐ranked results.
Nafiseh Torkzadeh Mahani, Mostafa Dehghani 0001, Maryam S. Mirian, Azadeh Shakery, Khalil Taheri
Expert Syst. J. Knowl. Eng.4
2018 PERSON: Personalized information retrieval evaluation based on citation networks
Shayan A. Tabrizi, Azadeh Shakery, Hamed Zamani, Mohammad Ali Tavallaei 0002
Inf. Process. Manag.2
2018 A language model-based framework for multi-publisher content-based recommender systems
Hamed Zamani, Azadeh Shakery
Inf. Retr. J.2
2018 Bottom-up sequential anonymization in the presence of adversary knowledge
Fatemeh Amiri, Nasser Yazdani, Azadeh Shakery
Inf. Sci.3
2017 Iterative Estimation of Document Relevance Score for Pseudo-Relevance Feedback
Mozhdeh Ariannezhad, Ali Montazeralghaem, Hamed Zamani, Azadeh Shakery
ECIR4
2017 Dimension Projection Among Languages Based on Pseudo-Relevant Documents for Query Translation
Javid Dadashkarimi, Mahsa S. Shahshahani, Amirhossein Tebbifakhr, Heshaam Faili, Azadeh Shakery
ECIR5
2017 Negative Feedback in the Language Modeling Framework for Text Recommendation
Hossein Rahmatizadeh Zagheli, Mozhdeh Ariannezhad, Azadeh Shakery
ECIR3
2017 A Semantic-Aware Profile Updating Model for Text Recommendation
abstract
Content-based recommender systems (CBRSs) rely on user-item similarities that are calculated between user profiles and item representations. Appropriate representation of each user profile based on the user's past preferences can have a great impact on user's satisfaction in CBRSs. In this paper, we focus on text recommendation and propose a novel profile updating model based on previously recommended items as well as semantic similarity of terms calculated using distributed representation of words. We evaluate our model using two standard text recommendation datasets: TREC-9 Filtering Track and CLEF 2008-09 INFILE Track collections. Our experiments investigate the importance of both past recommended items and semantic similarities in recommendation performance. The proposed profile updating method significantly outperforms the baselines, which confirms the importance of incorporating semantic similarities in the profile updating task.
Hossein Rahmatizadeh Zagheli, Hamed Zamani, Azadeh Shakery
RecSys3
2017 Improving Retrieval Performance for Verbose Queries via Axiomatic Analysis of Term Discrimination Heuristic
abstract
Number of terms in a query is a query-specific constant that is typically ignored in retrieval functions. However, previous studies have shown that the performance of retrieval models varies for different query lengths, and it usually degrades when query length increases. A possible reason for this issue can be the extraneous terms in longer queries that makes it a challenge for the retrieval models to distinguish between the key and complementary concepts of the query. As a signal to understand the importance of a term, inverse document frequency (IDF) can be used to discriminate query terms. In this paper, we propose a constraint to model the interaction between query length and IDF. Our theoretical analysis shows that current state-of-the-art retrieval models, such as BM25, do not satisfy the proposed constraint. We further analyze the BM25 model and suggest a modification to adapt BM25 so that it adheres to the new constraint. Our experiments on three TREC collections demonstrate that the proposed modification outperforms the baselines, especially for verbose queries.
Mozhdeh Ariannezhad, Ali Montazeralghaem, Hamed Zamani, Azadeh Shakery
SIGIR4
2017 Term Proximity Constraints for Pseudo-Relevance Feedback
abstract
Pseudo-relevance feedback (PRF) refers to a query expansion strategy based on top-retrieved documents, which has been shown to be highly effective in many retrieval models. Previous work has introduced a set of constraints (axioms) that should be satisfied by any PRF model. In this paper, we propose three additional constraints based on the proximity of feedback terms to the query terms in the feedback documents. As a case study, we consider the log-logistic model, a state-of-the-art PRF model that has been proven to be a successful method in satisfying the existing PRF constraints, and show that it does not satisfy the proposed constraints. We further modify the log-logistic model based on the proposed proximity-based constraints. Experiments on four TREC collections demonstrate the effectiveness of the proposed constraints. Our modification the log-logistic model leads to significant and substantial (up to 15%) improvements. Furthermore, we show that the proposed proximity-based function outperforms the well-known Gaussian kernel which does not satisfy all the proposed constraints.
Ali Montazeralghaem, Hamed Zamani, Azadeh Shakery
SIGIR3
2017 Online Learning to Rank for Cross-Language Information Retrieval
abstract
Online learning to rank for information retrieval has shown great promise in optimization of Web search results based on user interactions. However, online learning to rank has been used only in the monolingual setting where queries and documents are in the same language. In this work, we present the first empirical study of optimizing a model for Cross-Language Information Retrieval (CLIR) based on implicit feedback inferred from user interactions. We show that ranking models for CLIR with acceptable performance can be learned in an online setting although ranking features are noisy because of the language mismatch.
Razieh Rahimi, Azadeh Shakery
SIGIR2
2017 An expectation-maximization algorithm for query translation based on pseudo-relevant documents
Javid Dadashkarimi, Azadeh Shakery, Heshaam Faili, Hamed Zamani
Inf. Process. Manag.2
2016 Pseudo-Relevance Feedback Based on Matrix Factorization
abstract
In information retrieval, pseudo-relevance feedback (PRF) refers to a strategy for updating the query model using the top retrieved documents. PRF has been proven to be highly effective in improving the retrieval performance. In this paper, we look at the PRF task as a recommendation problem: the goal is to recommend a number of terms for a given query along with weights, such that the final weights of terms in the updated query model better reflect the terms' contributions in the query. To do so, we propose RFMF, a PRF framework based on matrix factorization which is a state-of-the-art technique in collaborative recommender systems. Our purpose is to predict the weight of terms that have not appeared in the query and matrix factorization techniques are used to predict these weights. In RFMF, we first create a matrix whose elements are computed using a weight function that shows how much a term discriminates the query or the top retrieved documents from the collection. Then, we re-estimate the created matrix using a matrix factorization technique. Finally, the query model is updated using the re-estimated matrix. RFMF is a general framework that can be employed with any retrieval model. In this paper, we implement this framework for two widely used document retrieval frameworks: language modeling and the vector space model. Extensive experiments over several TREC collections demonstrate that the RFMF framework significantly outperforms competitive baselines. These results indicate the potential of using other recommendation techniques in this task.
Hamed Zamani, Javid Dadashkarimi, Azadeh Shakery, W. Bruce Croft
CIKM3
2016 Learning to Weight Translations using Ordinal Linear Regression and Query-generated Training Data for Ad-hoc Retrieval with Long Queries
abstract
Ordinal regression which is known with learning to rank has long been used in information retrieval (IR). Learning to rank algorithms, have been tailored in document ranking, information filtering, and building large aligned corpora successfully. In this paper, we propose to use this algorithm for query modeling in cross-language environments. To this end, first we build a query-generated training data using pseudo-relevant documents to the query and all translation candidates. The pseudo-relevant documents are obtained by top-ranked documents in response to a translation of the original query. The class of each candidate in the training data is determined based on presence/absence of the candidate in the pseudo-relevant documents. We learn an ordinal regression model to score the candidates based on their relevance to the context of the query, and after that, we construct a query-dependent translation model using a softmax function. Finally, we re-weight the query based on the obtained model. Experimental results on French, German, Spanish, and Italian CLEF collections demonstrate that the proposed method achieves better results compared to state-of-the-art cross-language information retrieval methods, particularly in long queries with large training data.
Javid Dadashkarimi, Masoud Jalili Sabet, Azadeh Shakery
COLING3
2016 Using a Dictionary and n-gram Alignment to Improve Fine-grained Cross-Language Plagiarism Detection
abstract
The Web offers fast and easy access to a wide range of documents in various languages, and translation and editing tools provide the means to create derivative documents fairly easily. This leads to the need to develop effective tools for detecting cross-language plagiarism. Given a suspicious document, cross-language plagiarism detection comprises two main subtasks: retrieving documents that are candidate sources for that document and analyzing those candidates one by one to determine their similarity to the suspicious document. In this paper we focus on the second subtask and introduce a novel approach for assessing cross-language similarity between texts for detecting plagiarized cases. Our proposed approach has two main steps: a vector-based retrieval framework that focuses on high recall, followed by a more precise similarity analysis based on dynamic text alignment. Experiments show that our method outperforms the methods of the best results in PAN-2012 and PAN-2014 in terms of plagdet score. We also show that aligning n-gram units, instead of aligning complete sentences, improves the accuracy of detecting plagiarism.
Nava Ehsan, Frank Wm. Tompa, Azadeh Shakery
DocEng3
2016 Cross Domain User Engagement Evaluation
Ali Montazeralghaem, Hamed Zamani, Azadeh Shakery
ECIR3
2016 Axiomatic Analysis for Improving the Log-Logistic Feedback Model
abstract
Pseudo-relevance feedback (PRF) has been proven to be an effective query expansion strategy to improve retrieval performance. Several PRF methods have so far been proposed for many retrieval models. Recent theoretical studies of PRF methods show that most of the PRF methods do not satisfy all necessary constraints. Among all, the log-logistic model has been shown to be an effective method that satisfies most of the PRF constraints. In this paper, we first introduce two new PRF constraints. We further analyze the log-logistic feedback model and show that it does not satisfy these two constraints as well as the previously proposed "relevance effect" constraint. We then modify the log-logistic formulation to satisfy all these constraints. Experiments on three TREC newswire and web collections demonstrate that the proposed modification significantly outperforms the original log-logistic model, in all collections.
Ali Montazeralghaem, Hamed Zamani, Azadeh Shakery
SIGIR3
2016 Sentence alignment using local and global information
Hamed Zamani, Heshaam Faili, Azadeh Shakery
Comput. Speech Lang.3
2016 Candidate document retrieval for cross-lingual plagiarism detection using two-level proximity information
Nava Ehsan, Azadeh Shakery
Inf. Process. Manag.2
2016 Extracting translations from comparable corpora for Cross-Language Information Retrieval using the language modeling framework
Razieh Rahimi, Azadeh Shakery, Irwin King
Inf. Process. Manag.2
2016 Hierarchical anonymization algorithms against background knowledge attack in data releasing
Fatemeh Amiri, Nasser Yazdani, Azadeh Shakery, Amir H. Chinaei
Knowl. Based Syst.3
2016 Alecsa: Attentive Learning for Email Categorization using Structural Aspects
Mostafa Dehghani 0001, Azadeh Shakery, Maryam S. Mirian
Knowl. Based Syst.2
2016 Building a multi-domain comparable corpus using a learning to rank method
abstract
Abstract Comparable corpora are key translation resources for both languages and domains with limited linguistic resources. The existing approaches for building comparable corpora are mostly based on ranking candidate documents in the target language for each source document using a cross-lingual retrieval model. These approaches also exploit other evidence of document similarity, such as proper names and publication dates, to build more reliable alignments. However, the importance of each evidence in the scores of candidate target documents is determined heuristically. In this paper, we employ a learning to rank method for ranking candidate target documents with respect to each source document. The ranking model is constructed by defining each evidence for similarity of bilingual documents as a feature whose weight is learned automatically. Learning feature weights can significantly improve the quality of alignments, because the reliability of features depends on the characteristics of both source and target languages of a comparable corpus. We also propose a method to generate appropriate training data for the task of building comparable corpora. We employed the proposed learning-based approach to build a multi-domain English–Persian comparable corpus which covers twelve different domains obtained from Open Directory Project. Experimental results show that the created alignments have high degrees of comparability. Comparison with existing approaches for building comparable corpora shows that our learning-based approach improves both quality and coverage of alignments.
Razieh Rahimi, Azadeh Shakery, Javid Dadashkarimi, Mozhdeh Ariannezhad, Mostafa Dehghani 0001, Hossein Nasr Esfahani
Nat. Lang. Eng.2
2015 Adaptive User Engagement Evaluation via Multi-task Learning
abstract
User engagement evaluation task in social networks has recently attracted considerable attention due to its applications in recommender systems. In this task, the posts containing users' opinions about items, e.g., the tweets containing the users' ratings about movies in the IMDb website, are studied. In this paper, we try to make use of tweets from different web applications to improve the user engagement evaluation performance. To this aim, we propose an adaptive method based on multi-task learning. Since in this paper we study the problem of detecting tweets with positive engagement which is a highly imbalanced classification problem, we modify the loss function of multi-task learning algorithms to cope with the imbalanced data. Our evaluations over a dataset including the tweets of four diverse and popular data sources, i.e., IMDb, YouTube, Goodreads, and Pandora, demonstrate the effectiveness of the proposed method. Our findings suggest that transferring knowledge between data sources can improve the user engagement evaluation performance.
Hamed Zamani, Pooya Moradi, Azadeh Shakery
SIGIR3
2015 Multilingual information retrieval in the language modeling framework
Razieh Rahimi, Azadeh Shakery, Irwin King
Inf. Retr. J.2
2014 Axiomatic Analysis of Cross-Language Information Retrieval
abstract
A major challenge in Cross-Language Information Retrieval (CLIR) is the adoption of translation knowledge in retrieval models, as it affects the term weighting which is known to highly impact the retrieval performance. In this paper, we present an analytical study of using translation knowledge in CLIR. In particular, by adopting axiomatic analysis framework, we formulate the impacts of translation knowledge on document ranking as constraints that any cross-language retrieval model should satisfy. We then consider the state-of-the-art CLIR methods and check whether they satisfy these constraints. Finally, we show through empirical evaluation that violating one of the constraints harms the retrieval performance significantly which calls for further investigation.
Razieh Rahimi, Azadeh Shakery, Irwin King
CIKM2
2014 Mining a Persian-English comparable corpus for cross-language information retrieval
abstract
Knowledge acquisition and bilingual terminology extraction from multilingual corpora are challenging tasks for cross-language information retrieval. In this study, we propose a novel method for mining high quality translation knowledge from our constructed Persian–English comparable corpus, University of Tehran Persian–English Comparable Corpus (UTPECC). We extract translation knowledge based on Term Association Network (TAN) constructed from term co-occurrences in same language as well as term associations in different languages. We further propose a post-processing step to do term translation validity check by detecting the mistranslated terms as outliers. Evaluation results on two different data sets show that translating queries using UTPECC and using the proposed methods significantly outperform simple dictionary-based methods. Moreover, the experimental results show that our methods are especially effective in translating Out-Of-Vocabulary terms and also expanding query words based on their associated terms.
Homa Baradaran Hashemi, Azadeh Shakery
Inf. Process. Manag.2
2014 Semi-supervised word polarity identification in resource-lean languages
Iman Dehdarbehbahani, Azadeh Shakery, Heshaam Faili
Neural Networks2
2013 A Language Modeling Approach for Extracting Translation Knowledge from Comparable Corpora
Razieh Rahimi, Azadeh Shakery
ECIR2
2013 Leveraging comparable corpora for cross-lingual information retrieval in resource-lean language pairs
Azadeh Shakery, ChengXiang Zhai
Inf. Retr.1
2012 An Evolutionary-Based Method for Reconstructing Conversation Threads in Email Corpora
abstract
Email is a type of Web data which is produced in enormous quantities. It is beneficial to detect conversation threads contained in the email corpora for various applications, including discussion search, expert finding and even email clustering and classification. Conversation thread in email corpora can be defined as a cluster of exchanged emails among the same group of people by reply or forwarding on the same topic. According to this definition, we can define parent-child relation between emails, so email conversation threads seem to demonstrate tree structure. This paper presents a new approach based on genetic programming for reconstruction of conversation threads in emails data. This approach considers finding email conversation threads as an optimization problem, and exploits genetic programming to search intelligently in the space of possible solutions. Rather than several studies that have been conducted on this problem, this work concentrates on detecting accurate structure of conversation threads in high recall. This paper provides a comprehensive evaluation on the BC3 data set. Preliminary results suggest that our method provides acceptable precision and higher recall than existing methods.
Mostafa Dehghani 0001, Masoud Asadpour, Azadeh Shakery
ASONAM3
2012 Applying Sentiment and Social Network Analysis in User Modeling
Mohammadreza Shams, Mohammad Taghi Saffar, Azadeh Shakery, Heshaam Faili
CICLing (1)3
2011 Mutual information-based feature selection for intrusion detection systems
Fatemeh Amiri, Mohammad Mahdi Rezaei Yousefi, Caro Lucas, Azadeh Shakery, Nasser Yazdani
J. Netw. Comput. Appl.4
2009 Beyond hyperlinks: organizing information footprints in search logs to support effective browsing
abstract
While current search engines serve known-item search such as homepage finding very well, they generally cannot support exploratory search effectively. In exploratory search, users do not know their information needs precisely and also often lack the needed knowledge to formulate effective queries, thus querying alone, as supported by the current search engines, is insufficient, and browsing into related information would be very useful. Currently, browsing is mostly done by following hyperlinks embedded on Web pages. In this paper, we propose to leverage search logs to allow a user to browse beyond hyperlinks with a multi-resolution topic map constructed based on search logs. Specifically, we treat search logs as "footprints" left by previous users in the information space and build a multi-resolution topic map to semantically capture and organize them in multiple granularities. Such a topic map can support a user to zoom in, zoom out, and navigate horizontally over the information space, and thus provide flexible and effective browsing capabilities for end users. To test the effectiveness of the proposed methods of supporting browsing, we rely on real search logs and a commercial search engine to implement our proposed methods. Our experimental results show that the proposed topic map is effective to support browsing beyond hyperlinks.
Xuanhui Wang, Azadeh Shakery, ChengXiang Zhai
CIKM3
2008 Smoothing document language models with probabilistic term count propagation
Azadeh Shakery, ChengXiang Zhai
Inf. Retr.1
2008 DirichletRank: Solving the zero-one gap problem of PageRank
abstract
Link-based ranking algorithms are among the most important techniques to improve web search. In particular, the PageRank algorithm has been successfully used in the Google search engine and has been attracting much attention recently. However, we find that PageRank has a “zero-one gap” problem which, to the best of our knowledge, has not been addressed in any previous work. This problem can be potentially exploited to spam PageRank results and make the state-of-the-art link-based antispamming techniques ineffective. The zero-one gap problem arises as a result of the current ad hoc way of computing transition probabilities in the random surfing model. We therefore propose a novel DirichletRank algorithm which calculates these probabilities using Bayesian estimation with a Dirichlet prior. DirichletRank is a variant of PageRank, but does not have the problem of zero-one gap and can be analytically shown substantially more resistant to some link spams than PageRank. Experiment results on TREC data show that DirichletRank can achieve better retrieval accuracy than PageRank due to its more reasonable allocation of transition probabilities. More importantly, experiments on the TREC dataset and another real web dataset from the Webgraph project show that, compared with the original PageRank, DirichletRank is more stable under link perturbation and is significantly more robust against both manually identified web spams and several simulated link spams. DirichletRank can be computed as efficiently as PageRank, and thus is scalable to large-scale web applications.
Xuanhui Wang, Tao Tao 0003, Jian-Tao Sun, Azadeh Shakery, ChengXiang Zhai
ACM Trans. Inf. Syst.4
2006 A probabilistic relevance propagation model for hypertext retrieval
abstract
A major challenge in developing models for hypertext retrieval is to effectively combine content information with the link structure available in hypertext collections. Although several link-based ranking methods have been developed to improve retrieval results, none of them can fully exploit the discrimination power of contents as well as fully exploit all useful link structures. In this paper, we propose a general relevance propagation framework for combining content and link information. The framework gives a probabilistic score to each document defined based on a probabilistic surfing model. Two main characteristics of our framework are our probabilistic view on the relevance propagation model and propagation through multiple sets of neighbors. We compare eight different models derived from the probabilistic relevance propagation framework on two standard TREC Web test collections. Our results show that all the eight relevance propagation models can outperform the baseline content only ranking method for a wide range of parameter values, indicating that the relevance propagation framework provides a general, effective and robust way of exploiting link information. Our experiments also show that using multiple neighbor sets outperforms using just one type of neighbors significantly and taking a probabilistic view of propagation provides guidance on setting propagation parameters.
Azadeh Shakery, ChengXiang Zhai
CIKM1
2005 Dirichlet PageRank
abstract
PageRank has been known to be a successful algorithm in ranking web sources. In order to avoid the rank sink problem, PageRank assumes that a surfer, being in a page, jumps to a random page with a certain probability. In the standard PageRank algorithm, the jumping probabilities are assumed to be the same for all the pages, regardless of the page properties. This is not the case in the real world, since presumably a surfer would more likely follow the out-links of a high-quality hub page than follow the links of a low-quality one. In this poster, we propose a novel algorithm "Dirichlet PageRank" to address this problem by adapting exible jumping probabilities based on the number of out-links in a page. Empirical results on TREC data show that our method outperforms the standard PageRank algorithm.
Xuanhui Wang, Azadeh Shakery, Tao Tao 0003
SIGIR2