Marie-Francine Moens

dblp:m/MarieFrancineMoens · also Sien Moens · DBLP profile ↗
← Back
52ranked-venue papers in the field
12as first author
8since 2021 · last 2024
0000-0002-3732-9323ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 36 (11 first)Data Mining & Knowledge Discovery · 8 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 2Database Systems & Data Management · 1Big Data, Cloud & Distributed Data Systems · 1
YearPublicationVenuePosition
2024 Align MacridVAE: Multimodal Alignment for Disentangled Recommendations
Ignacio Avas, Liesbeth Allein, Katrien Laenen, Marie-Francine Moens
ECIR (1)4
2024 Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
Omar Hamed, Souhail Bakkali, Matthew B. Blaschko, Marie-Francine Moens, Jordy Van Landeghem
ICDAR (4)4
2024 DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
Jordy Van Landeghem, Subhajit Maity, Ayan Banerjee 0002, Matthew B. Blaschko, Marie-Francine Moens, Josep Lladós 0001, Sanket Biswas
ICDAR (4)5
2023 Preventing profiling for ethical fake news detection
abstract
A news article's online audience provides useful insights about the article's identity. However, fake news classifiers using such information risk relying on profiling. In response to the rising demand for ethical AI, we present a profiling-avoiding algorithm that leverages Twitter users during model optimisation while excluding them when an article's veracity is evaluated. For this, we take inspiration from the social sciences and introduce two objective functions that maximise correlation between the article and its spreaders, and among those spreaders. We applied our profiling-avoiding algorithm to three popular neural classifiers and obtained results on fake news data discussing a variety of news topics. The positive impact on prediction performance demonstrates the soundness of the proposed objective functions to integrate social context in text-based classifiers. Moreover, statistical visualisation and dimension reduction techniques show that the user-inspired classifiers better discriminate between unseen fake and true news in their latent spaces. Our study serves as a stepping stone to resolve the underexplored issue of profiling-dependent decision-making in user-informed fake news detection.
Liesbeth Allein, Marie-Francine Moens, Domenico Perrotta
Inf. Process. Manag.2
2023 Investigating better context representations for generative question answering
Sumam Francis, Marie-Francine Moens
Inf. Retr. J.2
2022 Guest editorial: special issue on ECIR 2021
abstract
The 43rd European Conference on Information Retrieval, ECIR 2021, was supposed to take place as an in-person conference in Lucca, Italy. Due to the COVID-19 pandemic, ECIR 2021 was held entirely online from March 28 to April 1, 2021. The conference programme contained full paper presentations, poster presentations, system demonstrations, eight tutorials, five workshops, an industry event, a doctoral consortium, a reproducibility track, a panel on open access publishing and several online social events.
Djoerd Hiemstra, Marie-Francine Moens
Inf. Retr. J.2
2021 How Do Simple Transformations of Text and Image Features Impact Cosine-Based Semantic Match?
Guillem Collell, Marie-Francine Moens
ECIR (1)2
2021 Time-aware evidence ranking for fact-checking
abstract
Truth can vary over time. Fact-checking decisions on claim veracity should therefore take into account temporal information of both the claim and supporting or refuting evidence. In this work, we investigate the hypothesis that the timestamp of a Web page is crucial to how it should be ranked for a given claim. We delineate four temporal ranking methods that constrain evidence ranking differently and simulate hypothesis-specific evidence rankings given the evidence timestamps as gold standard. Evidence ranking in three fact-checking models is ultimately optimized using a learning-to-rank loss function. Our study reveals that time-aware evidence ranking not only surpasses relevance assumptions based purely on semantic similarity or position in a search results list, but also improves veracity predictions of time-sensitive claims in particular.
Liesbeth Allein, Isabelle Augenstein, Marie-Francine Moens
J. Web Semant.3
2020 A Comparative Study of Outfit Recommendation Methods with a Focus on Attention-based Fusion
Katrien Laenen, Marie-Francine Moens
Inf. Process. Manag.2
2019 Can Image Captioning Help Passage Retrieval in Multimodal Question Answering?
Shurong Sheng, Katrien Laenen, Marie-Francine Moens
ECIR (2)3
2018 User Profiling through Deep Multimodal Fusion
abstract
User profiling in social media has gained a lot of attention due to its varied set of applications in advertising, marketing, recruiting, and law enforcement. Among the various techniques for user modeling, there is fairly limited work on how to merge multiple sources or modalities of user data - such as text, images, and relations - to arrive at more accurate user profiles. In this paper, we propose a deep learning approach that extracts and fuses information across different modalities. Our hybrid user profiling framework utilizes a shared representation between modalities to integrate three sources of data at the feature level, and combines the decision of separate networks that operate on each combination of data sources at the decision level. Our experimental results on more than 5K Facebook users demonstrate that our approach outperforms competing approaches for inferring age, gender and personality traits of social media users. We get highly accurate results with AUC values of more than 0.9 for the task of age prediction and 0.95 for the task of gender prediction.
Golnoosh Farnadi, Jie Tang 0001, Martine De Cock, Marie-Francine Moens
WSDM4
2018 Web Search of Fashion Items with Multimodal Querying
abstract
In this paper, we introduce a novel multimodal fashion search paradigm where e-commerce data is searched with a multimodal query composed of both an image and text. In this setting, the query image shows a fashion product that the user likes and the query text allows to change certain product attributes to fit the product to the user's desire. Multimodal search gives users the means to clearly express what they are looking for. This is in contrast to current e-commerce search mechanisms, which are cumbersome and often fail to grasp the customer's needs. Multimodal search requires intermodal representations of visual and textual fashion attributes which can be mixed and matched to form the user's desired product, and which have a mechanism to indicate when a visual and textual fashion attribute represent the same concept. With a neural network, we induce a common, multimodal space for visual and textual fashion attributes where their inner product measures their semantic similarity. We build a multimodal retrieval model which operates on the obtained intermodal representations and which ranks images based on their relevance to a multimodal query. We demonstrate that our model is able to retrieve images that both exhibit the necessary query image attributes and satisfy the query texts. Moreover, we show that our model substantially outperforms two state-of-the-art retrieval models adapted to multimodal fashion search.
Katrien Laenen, Susana Zoghbi, Marie-Francine Moens
WSDM3
2017 Fast and Flexible Top-k Similarity Search on Large Networks
abstract
Similarity search is a fundamental problem in network analysis and can be applied in many applications, such as collaborator recommendation in coauthor networks, friend recommendation in social networks, and relation prediction in medical information networks. In this article, we propose a sampling-based method using random paths to estimate the similarities based on both common neighbors and structural contexts efficiently in very large homogeneous or heterogeneous information networks. We give a theoretical guarantee that the sampling size depends on the error-bound ε, the confidence level (1-δ), and the path length T of each random walk. We perform an extensive empirical study on a Tencent microblogging network of 1,000,000,000 edges. We show that our algorithm can return top- k similar vertices for any vertex in a network 300× faster than the state-of-the-art methods. We develop a prototype system of recommending similar authors to demonstrate the effectiveness of our method.
Jing Zhang 0001, Jie Tang 0001, Cong Ma 0001, Hanghang Tong, Yu Jing, Juan-Zi Li, Walter Luyten, Marie-Francine Moens
ACM Trans. Inf. Syst.8
2016 C-BiLDA extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content
Geert Heyman, Ivan Vulic, Marie-Francine Moens
Data Min. Knowl. Discov.3
2016 Latent Dirichlet allocation for linking user-generated content and e-commerce data
Susana Zoghbi, Ivan Vulic, Marie-Francine Moens
Inf. Sci.3
2015 Scalable adaptive label propagation in Grappa
abstract
Nodes of a social graph often represent entities with specific labels, denoting properties such as age-group or gender. Design of algorithms to assign labels to unlabeled nodes by leveraging node-proximity and a-priori labels of seed nodes is of significant interest. A semi-supervised approach to solve this problem is termed "LPA-Label Propagation Algorithm" where labels of a subset of nodes are iteratively propagated through the network to infer yet unknown node labels. While LPA for node labelling is extremely fast and simple, it works well only with an assumption of node-homophily — connected nodes are connected because they must deserve a similar label — which can often be a misnomer. In this paper we propose a novel algorithm "Adaptive Label Propagation" that dynamically adapts to the underlying characteristics of homophily, heterophily, or otherwise, of the connections of the network, and applies suitable label propagation strategies accordingly. Moreover, our adaptive label propagation approach is scalable as demonstrated by its implementation in Grappa, a distributed shared-memory system. Our experiments on social graphs from Facebook, YouTube, Live Journal, Orkut and Netlog demonstrate that our approach not only improves the labelling accuracy but also computes results for millions of users within a few seconds.
Golnoosh Farnadi, Zeinab Mahdavifar, Ivan Keller, Jacob Nelson 0001, Ankur Teredesai, Marie-Francine Moens, Martine De Cock
IEEE BigData6
2015 Monolingual and Cross-Lingual Information Retrieval Models Based on (Bilingual) Word Embeddings
abstract
We propose a new unified framework for monolingual (MoIR) and cross-lingual information retrieval (CLIR) which relies on the induction of dense real-valued word vectors known as word embeddings (WE) from comparable data. To this end, we make several important contributions: (1) We present a novel word representation learning model called Bilingual Word Embeddings Skip-Gram (BWESG) which is the first model able to learn bilingual word embeddings solely on the basis of document-aligned comparable data; (2) We demonstrate a simple yet effective approach to building document embeddings from single word embeddings by utilizing models from compositional distributional semantics. BWESG induces a shared cross-lingual embedding vector space in which both words, queries, and documents may be presented as dense real-valued vectors; (3) We build novel ad-hoc MoIR and CLIR models which rely on the induced word and document embeddings and the shared cross-lingual embedding space; (4) Experiments for English and Dutch MoIR, as well as for English-to-Dutch and Dutch-to-English CLIR using benchmarking CLEF 2001-2003 collections and queries demonstrate the utility of our WE-based MoIR and CLIR models. The best results on the CLEF collections are obtained by the combination of the WE-based approach and a unigram language model. We also report on significant improvements in ad-hoc IR tasks of our WE-based framework over the state-of-the-art framework for learning text representations from comparable data based on latent Dirichlet allocation (LDA).
Ivan Vulic, Marie-Francine Moens
SIGIR2
2015 Probabilistic topic modeling in multilingual settings: An overview of its methodology and applications
Ivan Vulic, Wim De Smet, Jie Tang 0001, Marie-Francine Moens
Inf. Process. Manag.4
2015 Global machine learning for spatial ontology population
Parisa Kordjamshidi, Marie-Francine Moens
J. Web Semant.2
2014 Learning to bridge colloquial and formal language applied to linking and search of E-Commerce data
abstract
We study the problem of linking information between different idiomatic usages of the same language, for example, colloquial and formal language. We propose a novel probabilistic topic model called multi-idiomatic LDA (MiLDA). Its modeling principles follow the intuition that certain words are shared between two idioms of the same language, while other words are non-shared, that is, idiom-specific. We demonstrate the ability of our model to learn relations between cross-idiomatic topics in a dataset containing product descriptions and reviews. We intrinsically evaluate our model by the perplexity measure. Following that, as an extrinsic evaluation, we present the utility of the new MiLDA topic model in a recently proposed IR task of linking Pinterest pins (given in colloquial English on the users' side) to online webshops (given in formal English on the retailers' side). We show that our multi-idiomatic model outperforms the standard monolingual LDA model and the pure bilingual LDA model both in terms of perplexity and MAP scores in the IR task.
Ivan Vulic, Susana Zoghbi, Marie-Francine Moens
SIGIR3
2014 Multilingual probabilistic topic modeling and its applications in web mining and search
abstract
Multilingual topic models are a fairly novel group of unsupervised, language-independent and generative machine learning models. This tutorial covers all key aspects of their probabilistic framework and demonstrates how to easily integrate these models into frameworks for cross-lingual and multilingual Web mining and search.
Marie-Francine Moens, Ivan Vulic
WSDM1
2013 Automatic Labeling of Forums Using Bloom's Taxonomy
Vanessa Echeverría, Juan-Carlos Gomez 0001, Marie-Francine Moens
ADMA (1)3
2013 Monolingual and Cross-Lingual Probabilistic Topic Models and Their Applications in Information Retrieval
Marie-Francine Moens, Ivan Vulic
ECIR1
2013 A Unified Framework for Monolingual and Cross-Lingual Relevance Modeling Based on Probabilistic Topic Models
Ivan Vulic, Marie-Francine Moens
ECIR2
2013 Representations for multi-document event clustering
Wim De Smet, Marie-Francine Moens
Data Min. Knowl. Discov.2
2013 Cross-language information retrieval models based on latent topic models trained with document-aligned comparable corpora
Ivan Vulic, Wim De Smet, Marie-Francine Moens
Inf. Retr.3
2012 The downside of markup: examining the harmful effects of CSS and javascript on indexing today's web
abstract
The continued development and maturation of advanced HTML features such as Cascading style sheets (CSS), Javascript, and AJAX, as well as their widespread adoption by browsers, has enabled web pages to flourish with sophistication and interactivity. Unfortunately, this presents challenges to the web search community, as a web page's representation in the browser (i.e., what users see) can diverge dramatically from its raw HTML content (i.e., what search engines index and retrieve). For example, interactive pages may contain content in regions that are not visible before a user action, such as focusing a tab, but which are nonetheless still contained within the raw HTML. We study this divergence by comparing raw HTML to its fully rendered form across a number of metrics spanning presentation, geometry, and content, using a large, representative sample of popular web pages. We find that a large divergence currently exists, and we show via a historical analysis that this divergence has grown more pronounced over the last decade. The general finding of our study is that continuing to index the web via simple HTML parsing will diminish the effectiveness of retrieval on the modern web, and that the IR community should work toward more sophisticated web page processing in indexing technology.
Karl Gyllstrom, Carsten Eickhoff, Arjen P. de Vries, Marie-Francine Moens
CIKM4
2012 Plink-LDA: Using Link as Prior Information in Topic Modeling
Huan Xia, Juan-Zi Li, Jie Tang 0001, Marie-Francine Moens
DASFAA (1)4
2012 EmSe: Supporting Children's Information Needs within a Hospital Environment
Leif Azzopardi, Douglas Dowie, Sergio Duarte Torres, Carsten Eickhoff, Richard Glassey, Karl Gyllstrom, Djoerd Hiemstra, Franciska de Jong, Frea Kruisinga, Kelly Ann Marshall, Marie-Francine Moens, Tamara Polajnar, Frans van der Sluis, Arjen P. de Vries
ECIR11
2012 Highly discriminative statistical features for email classification
Juan-Carlos Gomez 0001, Erik Boiy, Marie-Francine Moens
Knowl. Inf. Syst.3
2011 Examining the "leftness" property of Wikipedia categories
abstract
Wikipedia's rich category structure has helped make it one of the largest semantic taxonomies in existence, a property that has been central to much recent research. However, Wikipedia's category representation is simplistic: an article contains a single list of categories, with no data about their relative importance. We investigate the ordering of category lists to determine how a category's position in the list correlates with its relevance to the article and overall significance. We identify a number of interesting connections between a category's position and its persistence within the article, age, popularity, size, and descriptiveness.
Karl Gyllstrom, Marie-Francine Moens
CIKM2
2011 Clash of the Typings - Finding Controversies and Children's Topics Within Queries
Karl Gyllstrom, Marie-Francine Moens
ECIR2
2011 Knowledge Transfer across Multilingual Corpora via Latent Topics
Wim De Smet, Jie Tang 0001, Marie-Francine Moens
PAKDD (1)3
2011 Introduction to the special issue on question answering
Marie-Francine Moens, Patrick Saint-Dizier
Inf. Process. Manag.1
2011 Knowledge and reasoning for question answering: Research perspectives
Patrick Saint-Dizier, Marie-Francine Moens
Inf. Process. Manag.2
2011 A survey on question answering technology from an information retrieval perspective
Oleksandr Kolomiyets, Marie-Francine Moens
Inf. Sci.2
2010 Wisdom of the ages: toward delivering the children's web with the link-based agerank algorithm
abstract
Though children frequently use web search engines to learn, interact, and be entertained, modern web search engines are poorly suited to children's needs, requiring relatively complex querying and filtering of results in order to find pages oriented to young audiences. To address this limitation, we designed AgeRank, a link-based algorithm that ranks web pages according their appropriateness for young audiences. We show its effectiveness through a multipart evaluation that demonstrates AgeRank to be accurate in page-labeling, widely-spanning in page coverage, and with high potential to improve children's search. As a fast, scalable, and effective algorithm, AgeRank can be adopted by search engines seeking to more effectively address the needs of young users, or easily fitted to complementary machine-learning based classification approaches.
Karl Gyllstrom, Marie-Francine Moens
CIKM2
2010 A picture is worth a thousand search results: finding child-oriented multimedia results with collAge
abstract
We present a simple and effective approach to complement search results for children’s web queries with child-oriented multimedia results, such as coloring pages and music sheets. Our approach determines appropriate media types for a query by searching Google’s database of frequent queries for co-occurrences of a query’s terms (e.g., “dinosaurs”) with preselected multimedia terms (e.g., “coloring pages”). We show the effectiveness of this approach through an online user evaluation.
Karl Gyllstrom, Marie-Francine Moens
SIGIR2
2010 Linking content in unstructured sources
abstract
This tutorial focuses on the task of automated information linking in text and multimedia sources. In any task where information is fused from different sources, this linking is a necessary step. To solve the problem we borrow methods from computational linguistics, computer vision and data mining. Although the main focus is on finding equivalence relations in the sources, the tutorial opens views on the recognition of other relation types.
Marie-Francine Moens
WWW1
2009 Information Extraction and Linking in a Retrieval Context
Marie-Francine Moens, Djoerd Hiemstra
ECIR1
2009 A machine learning approach to sentiment analysis in multilingual Web texts
Erik Boiy, Marie-Francine Moens
Inf. Retr.2
2008 Finding the Best Picture: Cross-Media Retrieval of Content
Koen Deschacht, Marie-Francine Moens
ECIR2
2007 Cross-Document Entity Tracking
Roxana Angheluta, Marie-Francine Moens
ECIR2
2007 Extraction of Folksonomies from Noisy Texts
Wim De Smet, Marie-Francine Moens
ICWSM2
2007 Summarizing court decisions
Marie-Francine Moens
Inf. Process. Manag.1
2006 Rpref: a generalization of Bpref towards graded relevance judgments
abstract
We present rpref ; our generalization of the bpref evaluation metric for assessing the quality of search engine results, given graded rather than binary user relevance judgments.
Jan De Beer, Marie-Francine Moens
SIGIR2
2005 Generic technologies for single- and multi-document summarization
Marie-Francine Moens, Roxana Angheluta, Jos Dumortier
Inf. Process. Manag.1
2001 Generic Topic Segmentation of Document Texts
abstract
Topic segmentation is an important initial step in many text-based tasks. A hierarchical representation of a texts topics is useful in retrieval and allows judging relevancy at different levels of detail. This short paper describes research on generic algorithms for topic detection and segmentation that are applicable on texts of heterogeneous types and domains.
Marie-Francine Moens, Rik De Busser
SIGIR1
2000 Text categorization: the assignment of subject descriptors to magazine articles
Marie-Francine Moens, Jos Dumortier
Inf. Process. Manag.1
1999 Abstracting of Legal Cases: The Potential of Clustering Based on the Selection of Representative Objects
abstract
The SALOMON project automatically summarizes Belgian criminal cases in order to improve access to the large number of existing and future court decisions. SALOMON extracts text units from the case text to form a case summary. Such a case summary facilitates the rapid determination of the relevance of the case or may be employed in text search. An important part of the research concerns the development of techniques for automatic recognition of representative text paragraphs (or sentences) in texts of unrestricted domains. These techniques are employed to eliminate redundant material in the case texts, and to identify informative text paragraphs which are relevant to include in the case summary. An evaluation of a test set of 700 criminal cases demonstrates that the algorithms have an application potential for automatic indexing, abstracting, and text linking.
Marie-Francine Moens, Caroline Uyttendaele, Jos Dumortier
J. Am. Soc. Inf. Sci.1
1998 Automatic Abstracting of Magazine Articles: The Creation of 'Highlight' Abstracts
Marie-Francine Moens, Jos Dumortier
SIGIR1
1997 Automatic text structuring and categorization as a first step in summarizing legal cases
Marie-Francine Moens, Caroline Uyttendaele
Inf. Process. Manag.1