VLDB 2026 Research / reviewers in the wild / expert
Diana Inkpen
dblp:i/DianaInkpen · also Diana Zaiu Inkpen
· DBLP profile ↗
58ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-0202-2444ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 43 · 8 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Do Language Models Know Theo Has a Wife? Investigating the Proviso Problem
Tara Azin, Daniel Dumitrescu, Diana Inkpen, Raj Singh |
LREC | 3 |
| 2026 | SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant Hostility
Xuanyu Su, Diana Inkpen, Nathalie Japkowicz |
WWW | 2 |
| 2024 | Detecting Multiple Mental Health Disorders with Large Language ModelsabstractArtificial Intelligence systems have become useful in many sectors, including healthcare. In this paper, we focus on detecting signs of mental health disorders from social media texts. This can be used to direct patients to consult a healthcare professional, while on waiting lists, or for post-monitoring. The performance of the current algorithms for detecting multiple disorders is limited. We propose a new method to increase performance by including target-domain knowledge for each type of mental health disorder (for nine disorders). We used this information to select the best training examples to include in our carefully engineered prompts for performing few-shot learning. We show that the results improved when compared to zero-shot learning based on Large Language Models and when compared to state-of-the-art results on the same test set. Mehul Nanda, Diana Inkpen, Aroldo Dargel |
IV | 2 |
| 2024 | Towards Interpretable Emotion Classification: Evaluating LIME, SHAP, and Generative AI for Decision ExplanationsabstractThis paper explores the classification of multi-label emotions utilizing fine-tuned RoBERTa base and zero-shot GPT4 models, with experiments conducted on the SemEval 2018 E-c dataset encompassing 11 emotions, where more than one label is allowed for a text. Employing SHAP and LIME for RoBERTa explanations and generative AI for GPT4, we assess the sufficiency of explanations using the BERT score metric. We show the explanations generated by LIME and SHAP visually using different plots. The BERT score indicates that generative AI produces better explanations than the statistical models, providing deeper insights into emotion selection, with a BERT score of 59.66% compared to SHAP-RoBERTa's 54.17% and LIME-RoBERTa's 53.22%. This shows the potential of generative AI in revealing the reasoning behind decisions within complex emotional contexts. Though the performance is superior, we also discuss the limitations of these models that hinder wide-scale adoption. Muhammad Hammad Fahim Siddiqui, Diana Inkpen, Alexander F. Gelbukh |
IV | 2 |
| 2024 | SSL-GAN-RoBERTa: A robust semi-supervised model for detecting Anti-Asian COVID-19 hate speech on social mediaabstractAbstract Anti-Asian speech during the COVID-19 pandemic has been a serious problem with severe consequences. A hate speech wave swept social media platforms. The timely detection of Anti-Asian COVID-19-related hate speech is of utmost importance, not only to allow the application of preventive mechanisms but also to anticipate and possibly prevent other similar discriminatory situations. In this paper, we address the problem of detecting Anti-Asian COVID-19-related hate speech from social media data. Previous approaches that tackled this problem used a transformer-based model, BERT/RoBERTa, trained on the homologous annotated dataset and achieved good performance on this task. However, this requires extensive and annotated datasets with a strong connection to the topic. Both goals are difficult to meet without employing reliable, vast, and costly resources. In this paper, we propose a robust semi-supervised model, SSL-GAN-RoBERTa, that learns from a limited heterogeneous dataset and whose performance is further enhanced by using vast amounts of unlabeled data from another related domain. Compared with the RoBERTa baseline model, the experimental results show that the model has substantial performance gains in terms of Accuracy and Macro-F1 score in different scenarios that use data from different domains. Our proposed model achieves state-of-the-art performance results while efficiently using unlabeled data, showing promising applicability to other complex classification tasks where large amounts of labeled examples are difficult to obtain. Xuanyu Su, Paula Branco, Diana Inkpen |
Nat. Lang. Eng. | 4 |
| 2022 | Co-Regularized Adversarial Learning for Multi-Domain Text ClassificationabstractMulti-domain text classification (MDTC) aims to leverage all available resources from multiple domains to learn a predictive model that can generalize well on these domains. Recently, many MDTC methods adopt adversarial learning, shared-private paradigm, and entropy minimization to yield state-of-the-art results. However, these approaches face three issues: (1) Minimizing domain divergence can not fully guarantee the success of domain alignment; (2) Aligning marginal feature distributions can not fully guarantee the discriminability of the learned features; (3) Standard entropy minimization may make the predictions on unlabeled data over-confident, deteriorating the discriminability of the learned features. In order to address the above issues, we propose a co-regularized adversarial learning (CRAL) mechanism for MDTC. This approach constructs two diverse shared latent spaces, performs domain alignment in each of them, and punishes the disagreements of these two alignments with respect to the predictions on unlabeled data. Moreover, virtual adversarial training (VAT) with entropy minimization is incorporated to impose consistency regularization to the CRAL method. Experiments show that our model outperforms state-of-the-art methods on two MDTC benchmarks. Yuan Wu 0002, Diana Inkpen, Ahmed El-Roby |
AISTATS | 2 |
| 2022 | Maximum Batch Frobenius Norm for Multi-Domain Text ClassificationabstractMulti-domain text classification (MDTC) has obtained remarkable achievements due to the advent of deep learning. Recently, many endeavors are devoted to applying adversarial learning to extract domain-invariant features to yield state-of-the-art results. However, these methods still face one challenge: transforming original features to be domain-invariant distorts the distributions of the original features, degrading the discriminability of the learned features. To address this issue, we first investigate the structure of the batch classification output matrix and theoretically justify that the discriminability of the learned features has a positive correlation with the Frobenius norm of the batch output matrix. Based on this finding, we propose a maximum batch Frobenius norm (MBF) method to boost the feature discriminability for MDTC. Experiments on two MDTC benchmarks show that our MBF approach can effectively advance state-of-the-art performance. Yuan Wu 0002, Diana Inkpen, Ahmed El-Roby |
ICASSP | 2 |
| 2022 | Evaluation of Deep Learning Context-Sensitive Visualization ModelsabstractThe introduction of Transformer neural networks has changed the landscape of Natural Language Processing (NLP) during the recent years. These models are very complex, and therefore hard to debug and explain. In this context, visual explanation became an attractive approach. The visualization of the path that leads to certain outputs of a model is at the core of visual explanation, as this illuminates the features or parts of the model that may need to be changed to achieve the desired results. In particular, one goal of a NLP visual explanation is to highlight the most significant parts of the text that have the greatest impact on the model output. Several visual explanation methods for NLP models were recently proposed. A major challenge is how to compare the performances of such methods since we cannot simply use the usual classification accuracy measures to evaluate the quality of visualizations. We need good metrics and rigorous criteria to measure how useful the extracted knowledge is for explaining the models. In addition, we want to visualize the differences between the knowledge extracted by different models, in order to be able to rank them. In this paper, we investigate how to evaluate explanations/visualizations resulted from machine learning models for text classification. The goal is not to improve the accuracy of a particular NLP classifier, but to assess the quality of the visualizations that explain its decisions. We describe several methods for evaluating the quality of NLP visualizations, including both automated techniques based on quantifiable measures and subjective techniques based on human judgements. Andrew Dunn, Diana Inkpen, Razvan Andonie |
IV | 2 |
| 2021 | On the Softmax Bottleneck of Recurrent Language Models
Dwarak Govind Parthiban, Yongyi Mao, Diana Inkpen |
AAAI | 3 |
| 2021 | Mixup Regularized Adversarial Networks for Multi-Domain Text ClassificationabstractUsing the shared-private paradigm and adversarial training can significantly improve the performance of multi-domain text classification (MDTC) models. However, there are two issues for the existing methods: First, instances from the multiple domains are not sufficient for domain-invariant feature extraction. Second, aligning on the marginal distributions may lead to a fatal mismatch. In this paper, we propose mixup regularized adversarial networks (MRANs) to address these two issues. More specifically, the domain and category mixup regularizations are introduced to enrich the intrinsic features in the shared latent space and enforce consistent predictions in-between training instances such that the learned features can be more domain-invariant and discriminative. We conduct experiments on two benchmarks: The Amazon review dataset and the FDU-MTL dataset. Our approach on these two datasets yields average accuracies of 87.64% and 89.0% respectively, outperforming all relevant baselines. Yuan Wu 0002, Diana Inkpen, Ahmed El-Roby |
ICASSP | 2 |
| 2021 | Context-Sensitive Visualization of Deep Learning Natural Language Processing ModelsabstractThe introduction of Transformer neural networks has changed the landscape of Natural Language Processing (NLP) during the last years. So far, none of the visualization systems has yet managed to examine all the facets of the Transformers. This gave us the motivation of the current work. We propose a novel NLP Transformer context-sensitive visualization method that leverages existing NLP tools to find the most significant groups of tokens (words) that have the greatest effect on the output, thus preserving some context from the original text. The original contribution is a context-aware visualization method of the most influential word combinations with respect to a classifier. This context-sensitive approach leads to heatmaps that include more of the relevant information pertaining to the classification, as well as more accurately highlighting the most important words from the input text. The proposed method uses a dependency parser, a BERT model, and the leave-n-out technique. Experimental results suggest that improved visualizations increase the understanding of the model, and help design models that perform closer to the human level of understanding for these problems. Andrew Dunn, Diana Inkpen, Razvan Andonie |
IV | 2 |
| 2021 | Towards Unifying the Explainability Evaluation Methods for NLP
Diana Lucaci, Diana Inkpen |
NLPCC (2) | 2 |
| 2020 | Dual Mixup Regularized Learning for Adversarial Domain Adaptation
Yuan Wu 0002, Diana Inkpen, Ahmed El-Roby |
ECCV (29) | 2 |
| 2019 | Exploring deep neural networks for multitarget stance detectionabstractAbstract Detecting subjectivity expressed toward concerned targets is an interesting problem and has received intensive study. Previous work often treated each target independently, ignoring the potential (sometimes very strong) dependency that could exist among targets (eg, the subjectivity expressed toward two products or two political candidates in an election). In this paper, we relieve such an independence assumption in order to jointly model the subjectivity expressed toward multiple targets. We propose and show that an attention‐based encoder‐decoder framework is very effective for this problem, outperforming several alternatives that jointly learn dependent subjectivity through cascading classification or multitask learning, as well as models that independently predict subjectivity toward individual targets. Parinaz Sobhani, Diana Inkpen, Xiaodan Zhu 0001 |
Comput. Intell. | 2 |
| 2018 | Neural Natural Language Inference Models Enhanced with External KnowledgeabstractModeling natural language inference is a very challenging task.With the availability of large annotated data, it has recently become feasible to train complex models such as neural-network-based inference models, which have shown to achieve the state-of-the-art performance.Although there exist relatively large annotated data, can machines learn all knowledge needed to perform natural language inference (NLI) from these data?If not, how can neural-network-based NLI models benefit from external knowledge and how to build NLI models to leverage it?In this paper, we enrich the state-of-the-art neural natural language inference models with external knowledge.We demonstrate that the proposed models improve neural NLI models to achieve the state-of-the-art performance on the SNLI and MultiNLI datasets. Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Diana Inkpen, Si Wei |
ACL (1) | 4 |
| 2018 | Authorship Identification for Literary Book RecommendationsabstractBook recommender systems can help promote the practice of reading for pleasure, which has been declining in recent years. One factor that influences reading preferences is writing style. We propose a system that recommends books after learning their authors’ style. To our knowledge, this is the first work that applies the information learned by an author-identification model to book recommendations. We evaluated the system according to a top-k recommendation scenario. Our system gives better accuracy when compared with many state-of-the-art methods. We also conducted a qualitative analysis by checking if similar books/authors were annotated similarly by experts. Haifa Alharthi, Diana Inkpen, Stan Szpakowicz |
COLING | 2 |
| 2018 | Environmental and Geo-Spatial Data Analytics (EnGeoData'2018)abstractThe following topics are dealt with: learning (artificial intelligence); data analysis; social networking (online); pattern classification; regression analysis; data mining; Internet; neural nets; graph theory; trees (mathematics). Maguelonne Teisseire, Mathieu Roche, Diana Inkpen |
DSAA | 3 |
| 2018 | EDITORIALabstractIt is a great honor for me to be appointed as editor-in-chief of Computational Intelligence, as of July 1, 2018. I want to give special thanks to Evangelos Milios and Ali Ghorbani for being editors-in-chief of the journal for 13 years. They did a great job in shaping the direction of the journal and making it a venue for papers in the most important areas of Artificial Intelligence (AI), especially the emerging sub-areas that did not have other primary venues for the publication of their results. I also thank Hamid Nourashraf for his continuous help as the editorial assistant of the journal. The journal was created in 1985, with the support of the National Research Council (NRC) of Canada. It has become a truly international journal throughout the years, with a strong and diverse editorial board. Since 2007, the journal has been published by John Wiley and Sons Inc. In Volume 1 Issue 1 in 1985, the first editors-in-chief, Gordon McCall and Nick Cercone, wrote that “Research from the very large and expanding field of AI has been increasingly reported in conference proceedings, technical reports, and by personal communication,” justifying the need for another AI journal to provide the scientific community with more publication venues. This is even more the case now, when the field of AI expanded at a much larger scale with the advent of deep learning and other techniques based on learning from data and from knowledge bases. The recent improvements in image processing, speech recognition, natural language processing, and other areas make a difference in AI research and its practical applications. We see more and more companies, from start-ups to large corporations, developing commercial applications using these techniques. I have been passionate about AI since I was an undergraduate student in Computer Science at the Technical University of Cluj-Napoca, Romania. I started my research in Natural Language Processing and Machine Learning during my M.Sc. degree at the same university and, after that, during my PhD studies at the University of Toronto, Canada. Now I am a Professor of Computer Science at the University of Ottawa, Canada. I joined the faculty there 15 years ago and I started many research projects in AI together with my graduate students and collaborators. I am very happy to witness the unprecedented progress in the field of AI since I first started to work in this area and to be able to contribute to it. For my recent research activities, I was honored to be an invited speaker for several AI conferences, such as the International Conference on Pattern Recognition and Artificial Intelligence (ICPRAI 2018, Montreal, QC, May 2018), the Applied Natural Language Processing track at the 29th Florida Artificial Intelligence Research Society Conference (FLAIRS 2016, Key Largo, FL, May 2016), the 28th Canadian Conference on Artificial Intelligence (AI 2015, Halifax, NS, June 2015), and the International Symposium on Information Management and Big Data (SimBig 2015, Cuzco, Peru, September 2015). Recently, I published a book titled Natural Language Processing for Social Media (Morgan and Claypool Publishers, Synthesis Lectures on Human Language Technologies). The second edition published in December 2017. My goal is to continue to increase the international reputation of the Computational Intelligence journal and to maintain the high-quality standard of the published papers. We have a new structure of the editorial board: We have three new senior editors (area editors) on board, in addition to a renewed list of associate editors, with a variety of expertise in the latest AI techniques. The three area editors will represent the three main areas covered by the journal: Data Mining; Machine Learning & Applications, and Reasoning in Artificial Intelligence. We plan to streamline the reviewing process more, to ensure timely processing and publication. I invite you all to submit your exciting research results to Computational Intelligence and I hope you will find it a valuable platform to exchange ideas and to contribute to advancing the state-of-the art in AI. I also encourage the submission of surveys of specific sub-areas and the editing of focused special issues. Diana Inkpen |
Comput. Intell. | 1 |
| 2018 | Introduction to the Special Issue on Language in Social Media: Exploiting Discourse and Other Contextual InformationabstractSocial media content is changing the way people interact with each other and share information, personal messages, and opinions about situations, objects, and past experiences. Most social media texts are short online conversational posts or comments that do not contain enough information for natural language processing (NLP) tools, as they are often accompanied by non-linguistic contextual information, including meta-data (e.g., the user’s profile, the social network of the user, and their interactions with other users). Exploiting such different types of context and their interactions makes the automatic processing of social media texts a challenging research task. Indeed, simply applying traditional text mining tools is clearly sub-optimal, as, typically, these tools take into account neither the interactive dimension nor the particular nature of this data, which shares properties with both spoken and written language. This special issue contributes to a deeper understanding of the role of these interactions to process social media data from a new perspective in discourse interpretation. This introduction first provides the necessary background to understand what context is from both the linguistic and computational linguistic perspectives, then presents the most recent context-based approaches to NLP for social media. We conclude with an overview of the papers accepted in this special issue, highlighting what we believe are the future directions in processing social media texts. Farah Benamara, Diana Inkpen, Maite Taboada |
Comput. Linguistics | 2 |
| 2018 | A survey of book recommender systems
Haifa Alharthi, Diana Inkpen, Stan Szpakowicz |
J. Intell. Inf. Syst. | 2 |
| 2017 | Enhanced LSTM for Natural Language InferenceabstractReasoning and inference are central to human and artificial intelligence.Modeling inference in human language is very challenging.With the availability of large annotated data (Bowman et al., 2015), it has recently become feasible to train neural network based inference models, which have shown to be very effective.In this paper, we present a new state-of-the-art result, achieving the accuracy of 88.6% on the Stanford Natural Language Inference Dataset.Unlike the previous top models that use very complicated network architectures, we first demonstrate that carefully designing sequential inference models based on chain LSTMs can outperform all previous models.Based on this, we further show that by explicitly considering recursive architectures in both local inference modeling and inference composition, we achieve additional improvement.Particularly, incorporating syntactic parsing information contributes to our best result-it further improves the performance even when added to the already very strong model. Qian Chen 0003, Xiaodan Zhu 0001, Zhen-Hua Ling, Si Wei, Hui Jiang 0001, Diana Inkpen |
ACL (1) | 6 |
| 2017 | Location detection and disambiguation from twitter messages
Diana Inkpen, Anna Farzindar, Farzaneh Kazemi, Diman Ghazi |
J. Intell. Inf. Syst. | 1 |
| 2015 | Content-Based Recommender System Enriched with Wordnet Synsets
Haifa Alharthi, Diana Inkpen |
CICLing (2) | 2 |
| 2015 | Inferring Aspect-Specific Opinion Structure in Product Reviews Using Co-training
Dave Carter, Diana Inkpen |
CICLing (2) | 2 |
| 2015 | Detecting Emotion Stimuli in Emotion-Bearing Sentences
Diman Ghazi, Diana Inkpen, Stan Szpakowicz |
CICLing (2) | 2 |
| 2015 | Detecting and Disambiguating Locations Mentioned in Twitter Messages
Diana Inkpen, Anna Farzindar, Farzaneh Kazemi, Diman Ghazi |
CICLing (2) | 1 |
| 2014 | Prior and contextual emotion of words in sentential context
Diman Ghazi, Diana Inkpen, Stan Szpakowicz |
Comput. Speech Lang. | 2 |
| 2013 | Approaches of Anonymisation of an SMS Corpus
Namrata Patel, Pierre Accorsi, Diana Inkpen, Cédric Lopez, Mathieu Roche |
CICLing (1) | 3 |
| 2013 | Can I Hear You? Sentiment Analysis on Medical Forums
Tanveer Ali, David Schramm, Marina Sokolova, Diana Inkpen |
IJCNLP | 4 |
| 2013 | Computational Approaches to the Analysis of Emotion in Text
Diana Inkpen, Carlo Strapparava |
Comput. Intell. | 1 |
| 2013 | A Bootstrapping Method for Extracting Paraphrases of Emotion Expressions from TextsabstractBecause paraphrasing is one of the crucial tasks in natural language understanding and generation, this paper introduces a novel technique to extract paraphrases for emotion terms, from nonparallel corpora. We present a bootstrapping technique for identifying paraphrases, starting with a small number of seeds. WordNet Affect emotion words are used as seeds. The bootstrapping approach learns extraction patterns for six classes of emotions. We use annotated blogs and other data sets as texts from which to extract paraphrases, based on the highest scoring extraction patterns. The results include lexical and morphosyntactic paraphrases, that we evaluate with human judges. Fazel Keshtkar, Diana Inkpen |
Comput. Intell. | 2 |
| 2012 | Segmentation Similarity and Agreement
Chris Fournier, Diana Inkpen |
HLT-NAACL | 2 |
| 2012 | Getting More from Segmentation Evaluation
Martin Scaiano, Diana Inkpen |
HLT-NAACL | 2 |
| 2012 | A hierarchical approach to mood classification in blogsabstractAbstract In this article, we explore the task of mood classification for blog postings. We propose a novel approach that uses the hierarchy of possible moods to achieve better results than a standard machine learning approach. We also show that using sentiment orientation features improves the performance of classification. We used the Livejournal blog corpus as a data set to train and evaluate our method. We present extensive error analysis and discuss the difficulty of the task. Fazel Keshtkar, Diana Inkpen |
Nat. Lang. Eng. | 2 |
| 2011 | A Pattern-Based Model for Generating Text to Express Emotion
Fazel Keshtkar, Diana Inkpen |
ACII (2) | 2 |
| 2011 | Exploiting the systematic review protocol for classification of medical abstracts
Oana Frunza, Diana Inkpen, Stan Matwin, William Klement, Peter O'Blenis |
Artif. Intell. Medicine | 2 |
| 2011 | Letter: Performance of SVM and Bayesian classifiers on the systematic review classification taskabstractWe are grateful to Professor Cohen for his letter and clarification of the support vector machine algorithm (SVM) results (in press). We agree that the results he supplies fill a gap in our paper. We could not have performed this comparison in our paper, as the first version was written prior to the publication of his own article,1 which in any case, as Dr Cohen points out, did not include the SVM results in terms of within-groups sum of squares (WSS). We would like, however, to be cautious about the broader conclusion from the results in table 1 of Dr Cohen's letter. While they are indeed superior to our factorized version of the complement naïve Bayes (FCNB) approach, they do not necessarily indicate the general superiority of the SVM classifier over the Bayesian methods. Our implementation was an extension of the classical naïve Bayes classifier for the kind of imbalanced data likely to be encountered when classifying abstracts for a systematic review. New exciting developments in the area of Bayesian text classification, such as discriminative multinomial naïve Bayes2 and latent Dirichlet allocation,3 are more than likely to improve our results significantly. This is the topic of our current research. Moreover, we would like to bring up another aspect related to the use of SVM versus Bayesian approaches as regards the running time of the algorithms. It is well known that SVM is significantly slower than the Bayesian methods. While this may be acceptable for the data used in Cohen1 and in Cohen et al4 and in our paper5 which have the order of 103 abstracts, it may not be the case for datasets orders of larger magnitude. However, datasets with the order of 104 abstracts are, as far as we know, common in the practice of systematic review production. When working recently6 with such large datasets we attempted to use SVM, but the running times of the train/test protocols on the computers available to us were unacceptably long. Finally, we want to comment briefly on the performance of FCNB/weight engineering (WE) on the Opioids dataset. As this dataset has a very high imbalance (very low inclusion rate), it is encouraging to see that the FCNB/WE method, which, as we discuss in our paper, has been engineered specifically to work well with such imbalanced data, indeed performs better than the standard SVM. We agree with Professor Cohen that application of the text mining methods developed specifically for imbalanced data is an interesting topic of research (see eg, Zhuang and Dai7). None. Not commissioned; internally peer reviewed. Stan Matwin, Alexandre Kouznetsov, Diana Inkpen, Oana Frunza, Peter O'Blenis |
J. Am. Medical Informatics Assoc. | 3 |
| 2011 | A Machine Learning Approach for Identifying Disease-Treatment Relations in Short TextsabstractThe Machine Learning (ML) field has gained its momentum in almost any domain of research and just recently has become a reliable tool in the medical domain. The empirical domain of automatic learning is used in tasks such as medical decision support, medical imaging, protein-protein interaction, extraction of medical knowledge, and for overall patient management care. ML is envisioned as a tool by which computer-based systems can be integrated in the healthcare field in order to get a better, more efficient medical care. This paper describes a ML-based methodology for building an application that is capable of identifying and disseminating healthcare information. It extracts sentences from published medical papers that mention diseases and treatments, and identifies semantic relations that exist between diseases and treatments. Our evaluation results for these tasks show that the proposed methodology obtains reliable outcomes that could be integrated in an application to be used in the medical care domain. The potential value of this paper stands in the ML settings that we propose and in the fact that we outperform previous results on the same data set. Oana Frunza, Diana Inkpen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2010 | Identification of Translationese: A Machine Learning Approach
Iustina Ilisei, Diana Inkpen, Gloria Corpas Pastor, Ruslan Mitkov |
CICLing | 2 |
| 2010 | Romanian Zero Pronoun Distribution: A Comparative Study
Claudiu Mihaila, Iustina Ilisei, Diana Inkpen |
LREC | 3 |
| 2010 | A new algorithm for reducing the workload of experts in performing systematic reviewsabstractOBJECTIVE: To determine whether a factorized version of the complement naïve Bayes (FCNB) classifier can reduce the time spent by experts reviewing journal articles for inclusion in systematic reviews of drug class efficacy for disease treatment. DESIGN: The proposed classifier was evaluated on a test collection built from 15 systematic drug class reviews used in previous work. The FCNB classifier was constructed to classify each article as containing high-quality, drug class-specific evidence or not. Weight engineering (WE) techniques were added to reduce underestimation for Medical Subject Headings (MeSH)-based and Publication Type (PubType)-based features. Cross-validation experiments were performed to evaluate the classifier's parameters and performance. MEASUREMENTS: Work saved over sampling (WSS) at no less than a 95% recall was used as the main measure of performance. RESULTS: The minimum workload reduction for a systematic review for one topic, achieved with a FCNB/WE classifier, was 8.5%; the maximum was 62.2% and the average over the 15 topics was 33.5%. This is 15.0% higher than the average workload reduction obtained using a voting perceptron-based automated citation classification system. CONCLUSION: The FCNB/WE classifier is simple, easy to implement, and produces significantly better results in reducing the workload than previously achieved. The results support it being a useful algorithm for machine-learning-based automation of systematic reviews of drug class efficacy for disease treatment. Stan Matwin, Alexandre Kouznetsov, Diana Inkpen, Oana Frunza, Peter O'Blenis |
J. Am. Medical Informatics Assoc. | 3 |
| 2009 | Real-word spelling correction using Google web 1Tn-gram data setabstractWe present a method for correcting real-word spelling errors using the Google Web 1T n-gram data set and a normalized and modified version of the Longest Common Subsequence (LCS) string matching algorithm. Our method is focused mainly on how to improve the correction recall (the fraction of errors corrected) while keeping the correction precision (the fraction of suggestions that are correct) as high as possible. Evaluation results on a standard data set show that our method performs very well. Aminul Islam 0001, Diana Inkpen |
CIKM | 2 |
| 2009 | Real-Word Spelling Correction using Google Web 1T 3-grams
Aminul Islam 0001, Diana Inkpen |
EMNLP | 2 |
| 2009 | Inducing translations from officially published materials in Canadian government websites
Qibo Zhu, Diana Inkpen, Ash Asudeh |
MTSummit | 2 |
| 2008 | Combining Multiple Models for Speech Information Retrieval
Muath Alzghool, Diana Inkpen |
LREC | 2 |
| 2008 | Using the Complexity of the Distribution of Lexical Elements as a Feature in Authorship Attribution
Leanne Spracklin, Diana Inkpen, Amiya Nayak |
LREC | 2 |
| 2008 | Semantic text similarity using corpus-based word similarity and string similarityabstractWe present a method for measuring the semantic similarity of texts using a corpus-based measure of semantic word similarity and a normalized and modified version of the Longest Common Subsequence (LCS) string matching algorithm. Existing methods for computing text similarity have focused mainly on either large documents or individual words. We focus on computing the similarity between two sentences or two short paragraphs. The proposed method can be exploited in a variety of applications involving textual knowledge representation and knowledge discovery. Evaluation results on two different data sets show that our method outperforms several competing methods. Aminul Islam 0001, Diana Inkpen |
ACM Trans. Knowl. Discov. Data | 2 |
| 2008 | Applications of corpus-based semantic similarity and word segmentation to database schema matching
Aminul Islam 0001, Diana Inkpen, Iluju Kiringa |
VLDB J. | 2 |
| 2007 | A Generalized Approach to Word Segmentation Using Maximum Length Descending Frequency and Entropy Rate
Aminul Islam 0001, Diana Inkpen, Iluju Kiringa |
CICLing | 2 |
| 2007 | Near-Synonym Choice in an Intelligent Thesaurus
Diana Inkpen |
HLT-NAACL | 1 |
| 2007 | Automatic extraction of translations from web-based bilingual materials
Qibo Zhu, Diana Inkpen, Ash Asudeh |
Mach. Transl. | 2 |
| 2006 | Semi-Supervised Learning of Partial Cognates Using Bilingual BootstrappingabstractPartial cognates are pairs of words in two languages that have the same meaning in some, but not all contexts. Detecting the actual meaning of a partial cognate in context can be useful for Machine Translation tools and for Computer-Assisted Language Learning tools. In this paper we propose a supervised and a semi-supervised method to disambiguate partial cognates between two languages: French and English. The methods use only automatically-labeled data; therefore they can be applied for other pairs of languages as well. We also show that our methods perform well when using corpora from different domains. Oana Frunza, Diana Inkpen |
ACL | 2 |
| 2006 | Second Order Co-occurrence PMI for Determining the Semantic Similarity of Words
Aminul Islam 0001, Diana Inkpen |
LREC | 2 |
| 2006 | Investigating Cross-Language Speech Retrieval for a Spontaneous Conversational Speech Collection
Diana Inkpen, Muath Alzghool, Gareth J. F. Jones, Douglas W. Oard |
HLT-NAACL | 1 |
| 2006 | Sentiment Classification of Movie Reviews Using Contextual Valence ShiftersabstractWe present two methods for determining the sentiment expressed by a movie review. The semantic orientation of a review can be positive, negative, or neutral. We examine the effect of valence shifters on classifying the reviews. We examine three types of valence shifters: negations, intensifiers, and diminishers. Negations are used to reverse the semantic polarity of a particular term, while intensifiers and diminishers are used to increase and decrease, respectively, the degree to which a term is positive or negative. The first method classifies reviews based on the number of positive and negative terms they contain. We use the General Inquirer to identify positive and negative terms, as well as negation terms, intensifiers, and diminishers. We also use positive and negative terms from other sources, including a dictionary of synonym differences and a very large Web corpus. To compute corpus‐based semantic orientation values of terms, we use their association scores with a small group of positive and negative terms. We show that extending the term‐counting method with contextual valence shifters improves the accuracy of the classification. The second method uses a Machine Learning algorithm, Support Vector Machines. We start with unigram features and then add bigrams that consist of a valence shifter and another word. The accuracy of classification is very high, and the valence shifter bigrams slightly improve it. The features that contribute to the high accuracy are the words in the lists of positive and negative terms. Previous work focused on either the term‐counting method or the Machine Learning method. We show that combining the two methods achieves better results than either method alone. Alistair Kennedy, Diana Inkpen |
Comput. Intell. | 2 |
| 2006 | Building and Using a Lexical Knowledge Base of Near-Synonym DifferencesabstractChoosing the wrong word in a machine translation or natural language generation system can convey unwanted connotations, implications, or attitudes. The choice between near-synonyms such as error, mistake, slip, and blunder—words that share the same core meaning, but differ in their nuances—can be made only if knowledge about their differences is available. We present a method to automatically acquire a new type of lexical resource: a knowledge base of near-synonym differences. We develop an unsupervised decision-list algorithm that learns extraction patterns from a special dictionary of synonym differences. The patterns are then used to extract knowledge from the text of the dictionary. The initial knowledge base is later enriched with information from other machine-readable dictionaries. Information about the collocational behavior of the near-synonyms is acquired from free text. The knowledge base is used by Xenon, a natural language generation system that shows how the new lexical resource can be used to choose the best near-synonym in specific situations. Diana Inkpen, Graeme Hirst |
Comput. Linguistics | 1 |
| 2003 | Automatic Sense Disambiguation of the Near-Synonyms in a Dictionary Entry
Diana Inkpen, Graeme Hirst |
CICLing | 1 |
| 2001 | Experiments on Extracting Knowledge from a Machine-Readable Dictionary of Synonym Differences (Invited Talk)
Diana Inkpen, Graeme Hirst |
CICLing | 1 |