EDBT 2026 Demo / reviewers in the wild / expert
Saeedeh Momtazi
dblp:51/3441
· DBLP profile ↗
34ranked-venue papers
13as first author
17since 2021 · last 2026
0000-0002-8110-1342ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 8 first-author · 9 since 2021Databases, data management, data science and information retrieval · 11 · 7 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language ModelsabstractHallucination is a persistent issue affecting all large language Models (LLMs), particularly within low-resource languages such as Persian. PerHalluEval (Persian Hallucination Evaluation) is the first dynamic hallucination evaluation benchmark tailored for the Persian language. Our benchmark leverages a three-stage LLM-driven pipeline, augmented with human validation, to generate plausible answers and summaries regarding QA and summarization tasks, focusing on detecting extrinsic and intrinsic hallucinations. Moreover, we used the log probabilities of generated tokens to select the most believable hallucinated instances. In addition, we engaged human annotators to highlight Persian-specific contexts in the QA dataset in order to evaluate LLMs' performance on content specifically related to Persian culture. Our evaluation of 12 LLMs, including open- and closed-source models using PerHalluEval, revealed that the models generally struggle in detecting hallucinated Persian text. We showed that providing external knowledge, i.e., the original document for the summarization task, could mitigate hallucination partially. Furthermore, there was no significant difference in terms of hallucination when comparing LLMs specifically trained for Persian with others. Kimia Hosseini, Shayan Bali, Zahra Zanjani, Saeedeh Momtazi |
LREC | 5 |
| 2025 | E2TP: Element to tuple prompting improves aspect sentiment tuple prediction
Mohammad Ghiasvand Mohammadkhani, Niloofar Ranjbar, Saeedeh Momtazi |
Neural Networks | 3 |
| 2024 | Explaining recommendation system using counterfactual textual explanations
Niloofar Ranjbar, Saeedeh Momtazi, MohammadMehdi Homayoonpour |
Mach. Learn. | 2 |
| 2024 | A semantic modular framework for events topic modeling in social media
Arya Hadizadeh Moghaddam, Saeedeh Momtazi |
Multim. Tools Appl. | 2 |
| 2024 | Multi sentence description of complex manipulation action videosabstractAbstract Automatic video description necessitates generating natural language statements that encapsulate the actions, events, and objects within a video. An essential human capability in describing videos is to vary the level of detail, a feature that existing automatic video description methods, which typically generate single, fixed-level detail sentences, often overlook. This work delves into video descriptions of manipulation actions, where varying levels of detail are crucial to conveying information about the hierarchical structure of actions, also pertinent to contemporary robot learning techniques. We initially propose two frameworks: a hybrid statistical model and an end-to-end approach. The hybrid method, requiring significantly less data, statistically models uncertainties within video clips. Conversely, the end-to-end method, more data-intensive, establishes a direct link between the visual encoder and the language decoder, bypassing any statistical processing. Furthermore, we introduce an Integrated Method, aiming to amalgamate the benefits of both the hybrid statistical and end-to-end approaches, enhancing the adaptability and depth of video descriptions across different data availability scenarios. All three frameworks utilize LSTM stacks to facilitate description granularity, allowing videos to be depicted through either succinct single sentences or elaborate multi-sentence narratives. Quantitative results demonstrate that these methods produce more realistic descriptions than other competing approaches. Fatemeh Ziaeetabar, Reza Safabakhsh, Saeedeh Momtazi, Minija Tamosiunaite, Florentin Wörgötter |
Mach. Vis. Appl. | 3 |
| 2024 | SNRBERT: session-based news recommender using BERT
Ali Azizi, Saeedeh Momtazi |
User Model. User Adapt. Interact. | 2 |
| 2023 | Generative adversarial network for sentiment-based stock predictionabstractAbstract Financial markets received more attention due to technological advancements, such as Artificial Intelligence (AI). In addition to the price index, traders and investors constantly monitor stock news on social media. Therefore, predicting the market by analyzing public opinions is an important issue. In this research, we propose three models based on Generative Adversarial Network (GAN), namely Price‐GAN, Price‐Sentiment‐GAN, and Price‐Sentiment‐WGAN. The first model uses only optimized price features, and the two other models use sentiment features collected from social media as well as optimized price features. All the proposed GAN models include Long Short‐Term Memory (LSTM) as generators and Convolution Neural Networks (CNN) as discriminators. To evaluate the proposed models, two different social media datasets in English and Persian are used. Our proposed models predict the close stock price for 15 English and 5 Persian stocks. All of the proposed GAN models outperform the state‐of‐the‐art models by enhancing the performance of the English dataset by 2.44% and the Persian dataset by 12.11%. Sepehr Asgarian, Rouzbeh Ghasemi, Saeedeh Momtazi |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | PQuAD: A Persian question answering dataset
Kasra Darvishi, Newsha Shahbodaghkhan, Zahra Abbasiantaeb, Saeedeh Momtazi |
Comput. Speech Lang. | 4 |
| 2023 | Deep neural ranking model using distributed smoothing
Zahra Pourbahman, Saeedeh Momtazi, Alireza Bagheri |
Expert Syst. Appl. | 2 |
| 2023 | How a Deep Contextualized Representation and Attention Mechanism Justifies Explainable Cross-Lingual Sentiment AnalysisabstractThe number of applications in sentiment analysis is growing daily, and research in this field is increasing. Despite the rapid growth of data sources in English, low-resource languages suffer from a lack of data for accurate training models. Moreover, users cannot trust such systems without explaining the output. In this study, we propose a cross-lingual deep neural model to improve the accuracy of sentiment analysis for low-resource languages while providing an explainable description of the predictions. The proposed model contains a word representation model where we use XLM-RoBERTa, a pre-trained contextualized transformer-based cross-lingual language model, and a long short-term memory network together with an attention mechanism that helps improve the explainability of the model and detect the informative words that impact text polarity. Our experiments show the superiority of the proposed model compared to the state-of-the-art mono-lingual techniques and cross-lingual models. The results show 0.55% improvement compared to the cross-lingual sentiment analysis proposed by Ghasemi et al. and 15.08% improvement compared to the mono-lingual contextualized sentiment analysis. Moreover, we achieve 0.54% further improvement when using attention mechanisms for enhancing the model with explainability. Rouzbeh Ghasemi, Saeedeh Momtazi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Sequential credit card fraud detection: A joint deep neural network and probabilistic graphical model approachabstractAbstract With the wide usage of e‐banking in recent years, and by increased opportunities for fraudsters subsequently, we are witnessing a loss of billions of Euros worldwide due to credit card fraud every year. Therefore, credit card fraud detection has become a critical necessity for financial institutions. Several studies have used machine learning techniques for proposing a method to address the problem. However, most of them did not take into account the sequential nature of transactional data. In this paper, we proposed a novel credit card fraud detection model using sequence labelling based on both deep neural networks and probabilistic graphical models (PGM). Then by using two real‐world datasets, we compared our model with the baseline model and examined how considering hidden sequential dependencies among transactions and also among predicted labels can improve the results. Moreover, we introduce a novel undersampling algorithm, which helps to maintain the sequential patterns of data during the random undersampling process. Our experiments demonstrate that this algorithm achieves promising results compared to the state‐of‐the‐art methods in oversampling and undersampling. Javad Forough, Saeedeh Momtazi |
Expert Syst. J. Knowl. Eng. | 2 |
| 2022 | Entity-aware answer sentence selection for question answering with transformer-based language models
Zahra Abbasiantaeb, Saeedeh Momtazi |
J. Intell. Inf. Syst. | 2 |
| 2022 | Persian Fake News Detection: Neural Representation and Classification at Word and Text LevelsabstractNowadays, broadcasting news on social media and websites has grown at a swifter pace, which has had negative impacts on both the general public and governments; hence, this has urged us to build a fake news detection system. Contextualized word embeddings have achieved great success in recent years due to their power to embed both syntactic and semantic features of textual contents. In this article, we aim to address the problem of the lack of fake news datasets in Persian by introducing a new dataset crawled from different news agencies, and propose two deep models based on the Bidirectional Encoder Representations from Transformers model (BERT), which is a deep contextualized pre-trained model for extracting valuable features. In our proposed models, we benefit from two different settings of BERT, namely pool-based representation, which provides a representation for the whole document, and sequence representation, which provides a representation for each token of the document. In the former one, we connect a Single Layer Perceptron (SLP) to the BERT to use the embedding directly for detecting fake news. The latter one uses Convolutional Neural Network (CNN) after the BERT’s embedding layer to extract extra features based on the collocation of words in a corpus. Furthermore, we present the TAJ dataset, which is a new Persian fake news dataset crawled from news agencies’ websites. We evaluate our proposed models on the newly provided TAJ dataset as well as the two different Persian rumor datasets as baselines. The results indicate the effectiveness of using deep contextualized embedding approaches for the fake news detection task. We also show that both BERT-SLP and BERT-CNN models achieve superior performance to the previous baselines and traditional machine learning models, with 15.58% and 17.1% improvement compared to the reported results by Zamani et al. [ 30 ], and 11.29% and 11.18% improvement compared to the reported results by Jahanbakhsh-Nagadeh et al. [ 9 ]. Mohammadreza Samadi, Maryam Mousavian, Saeedeh Momtazi |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2022 | GuidedWalk
Mohsen Fazaeli, Saeedeh Momtazi |
World Wide Web | 2 |
| 2021 | Convolutional neural network with margin loss for fake news detection
Mohammad Hadi Goldani, Reza Safabakhsh, Saeedeh Momtazi |
Inf. Process. Manag. | 3 |
| 2021 | Deep contextualized text representation and learning for fake news detection
Mohammadreza Samadi, Maryam Mousavian, Saeedeh Momtazi |
Inf. Process. Manag. | 3 |
| 2021 | User Embedding for Expert Finding in Community Question AnsweringabstractThe number of users who have the appropriate knowledge to answer asked questions in community question answering is lower than those who ask questions. Therefore, finding expert users who can answer the questions is very crucial and useful. In this article, we propose a framework to find experts for given questions and assign them the related questions. The proposed model benefits from users’ relations in a community along with the lexical and semantic similarities between new question and existing answers. Node embedding is applied to the community graph to find similar users. Our experiments on four different Stack Exchange datasets show that adding community relations improves the performance of expert finding models. Negin Ghasemi, Ramin Fatourechi, Saeedeh Momtazi |
ACM Trans. Knowl. Discov. Data | 3 |
| 2020 | Cross-lingual embedding for cross-lingual question retrieval in low-resource community question answering
Shahrzad HajiAminShirazi, Saeedeh Momtazi |
Mach. Transl. | 2 |
| 2019 | Corrigendum to Unsupervised Latent Dirichlet Allocation for supervised question classification. [Information Processing & Management, 54(3), 380-393]
Saeedeh Momtazi, Iryna Gurevych |
Inf. Process. Manag. | 1 |
| 2018 | Unsupervised Latent Dirichlet Allocation for supervised question classification
Saeedeh Momtazi |
Inf. Process. Manag. | 1 |
| 2015 | Who wants a computer to be a millionaire?
Saeedeh Momtazi, Felix Naumann |
Inf. Process. Lett. | 1 |
| 2015 | Bridging the vocabulary gap between questions and answer sentences
Saeedeh Momtazi, Dietrich Klakow |
Inf. Process. Manag. | 1 |
| 2012 | Mobile texting: can post-ASR correction solve the issues? an experimental study on gain vs. costsabstractThe next big step in embedded, mobile speech recognition will be to allow completely free input as it is needed for messaging like SMS or email. However, unconstrained dictation remains error-prone, especially when the environment is noisy. In this paper, we compare different methods for improving a given free-text dictation system used to enter textbased messages in embedded mobile scenarios, where distraction, interaction cost, and hardware limitations enforce strict constraints over traditional scenarios. We present a corpus-based evaluation, measuring the trade-off between improvement of the word error rate versus the interaction steps that are required under various parameters. Results show that by post-processing the output of a "black box" speech recognizer (e.g. a web-based speech recognition service), a reduction of word error rate by 55% (10.3% abs.) can be obtained. For further error reduction, however, a richer representation of the original hypotheses (e.g. lattice) is necessary. Michael Feld, Saeedeh Momtazi, Farina Freigang, Dietrich Klakow, Christian Müller 0014 |
IUI | 2 |
| 2012 | Fine-grained German Sentiment Analysis on Social Media
Saeedeh Momtazi |
LREC | 1 |
| 2011 | Trained trigger language model for sentence retrieval in QA: bridging the vocabulary gapabstractWe propose a novel language model for sentence retrieval in Question Answering (QA) systems called trained trigger language model. This model addresses the word mismatch problem in information retrieval. The proposed model captures pairs of trigger and target words while training on a large corpus. The word pairs are extracted based on both unsupervised and supervised approaches while different notions of triggering are used. In addition, we study the impact of corpus size and domain for a supervised model. All notions of the trained trigger model are finally used in a language model-based sentence retrieval framework. Our experiments on TREC QA collection verify that the proposed model significantly improves the sentence retrieval performance compared to the state-of-the-art translation model and class model which address the same problem. Saeedeh Momtazi, Dietrich Klakow |
CIKM | 1 |
| 2010 | Within and across sentence boundary language modelabstractIn this paper, we propose two different language modeling approaches, namely skip trigram and across sentence boundary, to capture the long range dependencies. The skip trigram model is able to cover more predecessor words of the present word compared to the normal trigram while the same memory space is required. The across sentence boundary model uses the word distribution of the previous sentences to calculate the unigram probability which is applied as the emission probability in the word and the class model frameworks. Our experiments on the Penn Treebank [1] show that each of our proposed models and also their combination significantly outperform the baseline for both the word and the class models and their linear interpolation. The linear interpolation of the word and the class models with the proposed skip trigram and across sentence boundary models achieves 118.4 perplexity while the best state-of-the-art language model has a perplexity of 137.2 on the same dataset. 1. Saeedeh Momtazi, Friedrich Faubel, Dietrich Klakow |
INTERSPEECH | 1 |
| 2010 | A Comparative Study of Word Co-occurrence for Term Clustering in Language Model-based Sentence Retrieval
Saeedeh Momtazi, Sanjeev Khudanpur, Dietrich Klakow |
HLT-NAACL | 1 |
| 2010 | Hierarchical pitman-yor language model for information retrievalabstractIn this paper, we propose a new application of Bayesian language model based on Pitman-Yor process for information retrieval. This model is a generalization of the Dirichlet distribution. The Pitman-Yor process creates a power-law distribution which is one of the statistical properties of word frequency in natural language. Our experiments on Robust04 indicate that this model improves the document retrieval performance compared to the commonly used Dirichlet prior and absolute discounting smoothing techniques. Saeedeh Momtazi, Dietrich Klakow |
SIGIR | 1 |
| 2009 | A word clustering approach for language model-based sentence retrieval in question answering systemsabstractIn this paper we propose a term clustering approach to improve the performance of sentence retrieval in Question Answering (QA) systems. As the search in question answering is conducted over smaller segments of data than in a document retrieval task, the problems of data sparsity and exact matching become more critical. In this paper we propose Language Modeling (LM) techniques to overcome such problems and improve the sentence retrieval performance. Saeedeh Momtazi, Dietrich Klakow |
CIKM | 1 |
| 2009 | A Combined Query Expansion Technique for Retrieving Opinions from BlogsabstractIn this paper, we discuss the the role of the retrieval component in an TREC style opinion question answering system. Since blog retrieval differs from traditional ad-hoc document retrieval, we need to work on dedicated retrieval methods. In particular we focus on a new query expansion technique to retrieve people's opinions from blog posts. We propose a combined approach for expanding queries while considering two aspects: finding more relevant data, and finding more opinionative data. We introduce a method to select opinion bearing terms for query expansion based on a chi-squared test and use this new query expansion to combine it in a liner weighting scheme with the original query terms and relevant feedback terms from Web. We report our experiments on the TREC 2006 and TREC 2007 queries from the blog retrieval track. The results show that the methods investigated here enhanced mean average precision of document retrieval from 17.91% to 25.20% on TREC 2006 and from 22.28% to 32.61% on TREC 2007 queries. Saeedeh Momtazi, Stefan Kazalski, Dietrich Klakow |
ISDA | 1 |
| 2009 | A Possibilistic Approach for Building Statistical Language ModelsabstractClass-based n-gram language models are those most frequently-used in continuous speech recognition systems, especially for languages for which no richly annotated corpora are available. Various word clustering algorithms have been proposed to build such class-based models. In this work, we discuss the superiority of soft approaches to class construction, whereby each word can be assigned to more than one class. We also propose a new method for possibilistic word clustering. The possibilistic C-mean algorithm is used as our clustering method. Various parameters of this algorithm are investigated; e.g., centroid initialization, distance measure, and words' feature vector. In the experiments reported here, this algorithm is applied to the 20,000 most frequent Persian words, and the language model built with the clusters created in this fashion is evaluated based on its perplexity and the accuracy of a continuous speech recognition system. Our results indicate a 10% reduction in perplexity and a 4% reduction in word error rate. Saeedeh Momtazi, Hossein Sameti |
ISDA | 1 |
| 2009 | An Overview on the Existing Language Models for Prediction Systems as Writing Assistant ToolsabstractThe prediction task in national language processing means to guess the missing letter, word, phrase, or sentence that likely follow in a given segment of a text. Since 1980s many systems with different methods were developed for different languages. In this paper an overview of the existing prediction methods that have been used for more than two decades are described and a general classification of the approaches is presented. The three main categories of the classification are statistical modeling, knowledge-based modeling, and heuristic modeling (adaptive). Masood Ghayoomi, Saeedeh Momtazi |
SMC | 2 |
| 2008 | A New Word Clustering Method for Building N-Gram Language Models in Continuous Speech Recognition Systems
Mohammad Bahrani, Hossein Sameti, Nazila Hafezi, Saeedeh Momtazi |
IEA/AIE | 4 |
| 2008 | Solving Stochastic Path Problem: Particle Swarm Optimization Approach
Saeedeh Momtazi, Somayeh Kafi, Hamid Beigy |
IEA/AIE | 1 |