Cécile Paris

dblp:p/CParis · DBLP profile ↗
← Back
22ranked-venue papers in the field
4as first author
6since 2021 · last 2026
0000-0003-3816-0176ORCID · corroborated

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 13 (3 first)Data Mining & Knowledge Discovery · 5Knowledge Engineering, Semantic Web & Information Systems · 3 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Multi-Sourced, Multi-Agent Evidence Retrieval for Fact-Checking
abstract
Misinformation spreading over the Internet poses a significant threat to both societies and individuals, necessitating robust and scalable fact-checking that relies on retrieving accurate and trustworthy evidence. Previous methods rely on semantic and social-contextual patterns learned from training data, which limits their generalization to new data distributions. Recently, Retrieval Augmented Generation (RAG) based methods have been proposed to utilize the reasoning capability of LLMs with retrieved grounding evidence documents. However, these methods largely rely on textual similarity for evidence retrieval and struggle to retrieve evidence that captures multi-hop semantic relations within rich document contents. These limitations lead to overlooking subtle factual correlations between the evidence and the claims to be fact-checked during evidence retrieval, thus causing inaccurate veracity predictions.
Shuzhi Gong, Richard O. Sinnott, Jianzhong Qi 0001, Cécile Paris, Preslav Nakov, Zhuohan Xie
SIGIR4
2026 Question Answering Fit for Purpose: A Perspective From Natural Language Processing and User Modeling
abstract
Providing appropriate answers to questions is necessary in many situations, not just in the conversational AI systems we see and develop today. Research in Natural Language Processing (NLP) and User Modeling (UM) have investigated this topic for decades, starting in the era of ''symbolic AI''. While NLP in general was needed for the whole interaction (understanding the question and answering it), Natural Language Generation (NLG) was particularly concerned with providing good and coherent answers appropriate for the information need and the intended audience, which is where UM also played a role. At that time, information to include in the answers typically came from knowledge bases or data bases. Information Retrieval (IR) then was concerned with retrieving the documents (and later websites) most relevant to a query. As the amount of data and number of documents increased, information needs from users became increasingly complex. As a result, it seemed that combining advances in both NLP and IR was required. And of course, now, research often spans these two fields.
Cécile Paris
SIGIR1
2023 Fake News Detection Through Temporally Evolving User Interactions
Shuzhi Gong, Richard O. Sinnott, Jianzhong Qi 0001, Cécile Paris
PAKDD (4)4
2023 SciHarvester: Searching Scientific Documents for Numerical Values
abstract
A challenge for search technologies is to support scientific literature surveys that present overviews of the reported numerical values documented for specific physical properties. We present SciHarvester, a system tailored to address this problem for agronomic science. It provides an interface to search PubAg documents, allowing complex queries involving restrictions on numerical values. SciHarvester identifies relevant documents and generates overview of reported parameter values. The system allows interrogation of the results to explain the system's performance. Our evaluations demonstrate the promise of incorporating information extraction techniques with the use of neural scoring mechanisms.
Maciej Rybinski, Stephen Wan 0001, Sarvnaz Karimi, Cécile Paris, Brian Jin, Neil I. Huth, Peter J. Thorburn, Dean P. Holzworth
SIGIR4
2023 Dynalogue: A Transformer-Based Dialogue System with Dynamic Attention
abstract
Businesses face a range of cyber risks, both external threats and internal vulnerabilities that continue to evolve over time. As cyber attacks continue to increase in complexity and sophistication, more organisations will experience them. For this reason, it is important that organisations seek timely consultancy from cyber professionals so that they can respond to and recover from cyber attacks as quickly as possible. However, huge surges in cyber attacks have long left cyber professionals short of what is required to cover the security needs. This problem is getting worse when an increasing number of people choose to work from home during the pandemic because this situation usually yields extra communication cost.
Rongjunchen Zhang, Tingmin Wu, Xiao Chen 0002, Sheng Wen, Surya Nepal, Cécile Paris, Yang Xiang 0001
WWW6
2022 Explainable machine learning in cybersecurity: A survey
abstract
Machine learning (ML) techniques are increasingly important in cybersecurity, as they can quickly analyse and identify different types of threats from millions of events. In spite of the increasing number of possible applications of ML, successful adoption of ML models in cybersecurity still highly relies on the explainability of those models that are used for making predictions. Explanations that support ML model outputs are crucial in cybersecurity-oriented ML applications because people need to get more information from the model than just binary output for analysis. The explainable models help ML developers solve the “trust” problem for a security application prediction in a faithful way: validating model behaviours, diagnosing misclassifications and sometimes automatically patching errors in the target models. Therefore, explainable ML for cybersecurity has become a necessary and important research branch. In this paper, we present the topic of explainable ML in cybersecurity through two general types of explanations: (1) ante hoc explanation, and (2) post hoc explanation, with their methodologies. We systematically review and categorise the state-of-the-art research, and provide comparative studies to help researchers find the optimal solutions to specific problems. We further list open issues in this field to facilitate future studies. This survey will benefit diverse groups of readers from both academia and industries, who want to effectively use ML to solve cybersecurity challenges.
Feixue Yan, Sheng Wen, Surya Nepal, Cécile Paris, Yang Xiang 0001
Int. J. Intell. Syst.4
2020 Social Media Relevance Filtering Using Perplexity-Based Positive-Unlabelled Learning
Sunghwan Mac Kim, Stephen Wan 0001, Cécile Paris, Andreas Dünser
ICWSM3
2020 Less Is More: Rejecting Unreliable Reviews for Product Question Answering
Xiuzhen Zhang 0001, Jey Han Lau, Jeffrey Chan, Cécile Paris
ECML/PKDD (3)5
2020 Leveraging Sentiment Distributions to Distinguish Figurative From Literal Health Reports on Twitter
abstract
Harnessing data from social media to monitor health events is a promising avenue for public health surveillance. A key step is the detection of reports of a disease (referred to as ‘health mention classification’) amongst tweets that mention disease words. Prior work shows that figurative usage of disease words may prove to be challenging for health mention classification. Since the experience of a disease is associated with a negative sentiment, we present a method that utilises sentiment information to improve health mention classification. Specifically, our classifier for health mention classification combines pre-trained contextual word representations with sentiment distributions of words in the tweet. For our experiments, we extend a benchmark dataset of tweets for health mention classification, adding over 14k manually annotated tweets across diseases. We also additionally annotate each tweet with a label that indicates if the disease words are used in a figurative sense. Our classifier outperforms current SOTA approaches in detecting both health-related and figurative tweets that mention disease words. We also show that tweets containing disease words are mentioned figuratively more often than in a health-related context, proving to be challenging for classifiers targeting health-related tweets.
Rhys Biddle, Aditya Joshi 0001, Shaowu Liu, Cécile Paris, Guandong Xu
WWW4
2020 A survey of recent methods on deriving topics from Twitter: algorithm to evaluation
Robertus Nugroho, Cécile Paris, Surya Nepal, Jian Yang 0001, Weiliang Zhao
Knowl. Inf. Syst.2
2019 Discovering Relevant Reviews for Answering Product-Related Queries
abstract
With the increasing popularity of e-commerce, the number of product-related queries generated by customers is growing. Answering these queries manually in real time is infeasible, and so automatic question-answering systems can be immensely helpful. Product queries are, however, very different from open-domain questions: they tend to be product-specific and the answers they demand can be very subjective. Previous research suggests that reviews are a valuable resource for answering product queries, but a key challenge is the language mismatch between user queries and reviews. To address this, we propose two neural models that discover relevant reviews for answering product queries. We demonstrate that our best model produces strong performance, outperforming state-of-the-art systems by consistently finding the most relevant reviews for product queries.
Jey Han Lau, Xiuzhen Zhang 0001, Jeffrey Chan, Cécile Paris
ICDM5
2019 Automatic Recognition of Student Engagement Using Deep Learning and Facial Expression
Omid Mohamad Nezami, Mark Dras, Leonard G. C. Hamey, Debbie Richards 0001, Stephen Wan 0001, Cécile Paris
ECML/PKDD (3)6
2018 A Government-Run Online Community to Support Recipients of Welfare Payments
abstract
With the ubiquitous presence of smart phones and the availability of easy-to-use applications, there is an increase in the number of online services. A growing number of people now search for information and interact online. They expect to see services available and accessible online. To meet citizens’ expectations, governments have also increased their online presence. However, information and services are not the only reasons people go online. People also build their social circle online, seeking support and empathy, looking for someone with whom they can talk and who can understand their situation and worries. Online communities (and social networks in general) have been shown to have the potential to provide social and emotional peer-support. Our work aimed at determining whether online communities could be deployed in the public administration domain, in particular to support people receiving welfare payments, with similar benefits. We hypothesized that an online community could provide such support to disadvantaged citizens. Toward testing this hypothesis, after a user requirements analysis and some preparatory work, we designed and developed an online community for a specific target group of welfare recipients, as a collaboration between CSIRO and the Australian Department of Human Services. The community was deployed for one year. In this paper, we briefly explain our aims and the work that went into preparing for the community. We introduce the portal and the support it offered. We then report our observations and findings about both the informational and emotional support participants received, through an analysis of the comments posted in the community, and whether this support was perceived as welcome and useful.
Cécile Paris, Surya Nepal, Amanda Dennett
Int. J. Cooperative Inf. Syst.1
2017 Exploiting Users' Rating Behaviour to Enhance the Robustness of Social Recommendation
Zizhu Zhang, Weiliang Zhao, Jian Yang 0001, Surya Nepal, Cécile Paris
WISE (2)5
2015 Understanding Public Emotional Reactions on Twitter
Stephen Wan 0001, Cécile Paris
ICWSM2
2015 Time-Sensitive Topic Derivation in Twitter
Robertus Nugroho, Weiliang Zhao, Jian Yang 0001, Cécile Paris, Surya Nepal, Yan Mei
WISE (1)4
2014 Gamification for Online Communities: A Case Study for Delivering Government Services
abstract
Gamification, the idea of inserting game dynamics into portals or social networks, has recently evolved as an approach to encourage active participation in online communities. For an online community to start and proceed on to a sustainable operation, it is important that members are encouraged to contribute positively and frequently. We decided to introduce gamification in an online community that we designed and developed with the Australian Government's Department of Human Services to support welfare recipients transitioning from one payment to another. We first defined a formal model of gamification and a gamification design process. In instantiating our model to the online community, we realised that our context applied a number of constraints on the gamification elements that could be introduced. In this paper, we outline the design and implementation of a gamification model for online communities and its instantiation into our context, with its specific requirements. While we cannot comment on the success of gamification to drive user engagement in our context (for lack of the possibility of a controlled experiment), we found our implementation of badges-based gamification a helpful way to provide a useful abstraction on the life of the community, providing feedback enabling us to monitor and analyze the community. We thus show how feedback provided by such gamification data has a potential to be useful to community providers to better understand the community needs and addressing them appropriately to maintain a level of engagement in the community.
Sanat Kumar Bista, Surya Nepal, Cécile Paris, Nathalie Colineau
Int. J. Cooperative Inf. Syst.3
2012 Differences in Language and Style Between Two Social Media Communities
Cécile Paris, Paul Thomas 0001, Stephen Wan 0001
ICWSM1
2011 Expressing conditions in tailored brochures for public administration
abstract
Citizen-focused documents in Public Administration devote considerable effort to the expression of conditions. These conditions are commonly expressed as statements of eligibility requirements for the programs being described, but they manifest themselves in other places as well, such as in feedback to readers in tailored informational brochures and as input fields on program application forms. This paper discusses how administrative conditions can be represented in a manner that supports both the eligibility reasoning required for the generation of citizen-tailored documents and also the automated generation of condition expressions in a variety of forms. The paper pays particular attention to the question of how a generation mechanism can allow authors to override the default forms of automated expression when necessary. The discussion is based on a prototype tailored delivery application whose knowledge base is implemented in OWL DL and whose output is constructed using Myriad, a platform for tailored document planning and formatting.
Nathalie Colineau, Cécile Paris, Keith Vander Linden
ACM Symposium on Document Engineering2
2010 Focused and aggregated search: a perspective from natural language generation
Cécile Paris, Stephen Wan 0001, Paul Thomas 0001
Inf. Retr.1
2010 Supporting browsing-specific information needs: Introducing the Citation-Sensitive In-Browser Summariser
Stephen Wan 0001, Cécile Paris, Robert Dale
J. Web Semant.2
2000 Automatically Summarising Web Sites - Is There A Way Around It?
abstract
The challenge of automatically summarising Web pages and sites is a great one.However, currently there is no solution which offers an easy way to produce unbiased, coherent , and contentfull summaries of Web sites.In this work we suggest a new approach, which relies on the structure of hypertext and the way people describe information in it.As a proof-of-concept, we applied our approach to the problem of tailoring coherent snippets for search results.In this paper we describe the approach as it is implemented in our system, InCommonSense, and present results from a large scale evaluation of the snippets produced.We conclude by suggesting other applications that could make use of this technique.
Einat Amitay, Cécile Paris
CIKM2