VLDB 2026 Research / reviewers in the wild / expert
Marie-Francine Moens
dblp:m/MarieFrancineMoens · also Sien Moens
· DBLP profile ↗
52ranked-venue papers in the field
12as first author
8since 2021 · last 2024
0000-0002-3732-9323ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 36 (11 first)Data Mining & Knowledge Discovery · 8 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 2Database Systems & Data Management · 1Big Data, Cloud & Distributed Data Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Align MacridVAE: Multimodal Alignment for Disentangled Recommendations
Ignacio Avas, Liesbeth Allein, Katrien Laenen, Marie-Francine Moens |
ECIR (1) | 4 |
| 2024 | Multimodal Adaptive Inference for Document Image Classification with Anytime Early Exiting
Omar Hamed, Souhail Bakkali, Matthew B. Blaschko, Marie-Francine Moens, Jordy Van Landeghem |
ICDAR (4) | 4 |
| 2024 | DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
Jordy Van Landeghem, Subhajit Maity, Ayan Banerjee 0002, Matthew B. Blaschko, Marie-Francine Moens, Josep Lladós 0001, Sanket Biswas |
ICDAR (4) | 5 |
| 2023 | Preventing profiling for ethical fake news detectionabstractA news article's online audience provides useful insights about the article's identity. However, fake news classifiers using such information risk relying on profiling. In response to the rising demand for ethical AI, we present a profiling-avoiding algorithm that leverages Twitter users during model optimisation while excluding them when an article's veracity is evaluated. For this, we take inspiration from the social sciences and introduce two objective functions that maximise correlation between the article and its spreaders, and among those spreaders. We applied our profiling-avoiding algorithm to three popular neural classifiers and obtained results on fake news data discussing a variety of news topics. The positive impact on prediction performance demonstrates the soundness of the proposed objective functions to integrate social context in text-based classifiers. Moreover, statistical visualisation and dimension reduction techniques show that the user-inspired classifiers better discriminate between unseen fake and true news in their latent spaces. Our study serves as a stepping stone to resolve the underexplored issue of profiling-dependent decision-making in user-informed fake news detection. Liesbeth Allein, Marie-Francine Moens, Domenico Perrotta |
Inf. Process. Manag. | 2 |
| 2023 | Investigating better context representations for generative question answering
Sumam Francis, Marie-Francine Moens |
Inf. Retr. J. | 2 |
| 2022 | Guest editorial: special issue on ECIR 2021abstractThe 43rd European Conference on Information Retrieval, ECIR 2021, was supposed to take place as an in-person conference in Lucca, Italy. Due to the COVID-19 pandemic, ECIR 2021 was held entirely online from March 28 to April 1, 2021. The conference programme contained full paper presentations, poster presentations, system demonstrations, eight tutorials, five workshops, an industry event, a doctoral consortium, a reproducibility track, a panel on open access publishing and several online social events. Djoerd Hiemstra, Marie-Francine Moens |
Inf. Retr. J. | 2 |
| 2021 | How Do Simple Transformations of Text and Image Features Impact Cosine-Based Semantic Match?
Guillem Collell, Marie-Francine Moens |
ECIR (1) | 2 |
| 2021 | Time-aware evidence ranking for fact-checkingabstractTruth can vary over time. Fact-checking decisions on claim veracity should therefore take into account temporal information of both the claim and supporting or refuting evidence. In this work, we investigate the hypothesis that the timestamp of a Web page is crucial to how it should be ranked for a given claim. We delineate four temporal ranking methods that constrain evidence ranking differently and simulate hypothesis-specific evidence rankings given the evidence timestamps as gold standard. Evidence ranking in three fact-checking models is ultimately optimized using a learning-to-rank loss function. Our study reveals that time-aware evidence ranking not only surpasses relevance assumptions based purely on semantic similarity or position in a search results list, but also improves veracity predictions of time-sensitive claims in particular. Liesbeth Allein, Isabelle Augenstein, Marie-Francine Moens |
J. Web Semant. | 3 |
| 2020 | A Comparative Study of Outfit Recommendation Methods with a Focus on Attention-based Fusion
Katrien Laenen, Marie-Francine Moens |
Inf. Process. Manag. | 2 |
| 2019 | Can Image Captioning Help Passage Retrieval in Multimodal Question Answering?
Shurong Sheng, Katrien Laenen, Marie-Francine Moens |
ECIR (2) | 3 |
| 2018 | User Profiling through Deep Multimodal FusionabstractUser profiling in social media has gained a lot of attention due to its varied set of applications in advertising, marketing, recruiting, and law enforcement. Among the various techniques for user modeling, there is fairly limited work on how to merge multiple sources or modalities of user data - such as text, images, and relations - to arrive at more accurate user profiles. In this paper, we propose a deep learning approach that extracts and fuses information across different modalities. Our hybrid user profiling framework utilizes a shared representation between modalities to integrate three sources of data at the feature level, and combines the decision of separate networks that operate on each combination of data sources at the decision level. Our experimental results on more than 5K Facebook users demonstrate that our approach outperforms competing approaches for inferring age, gender and personality traits of social media users. We get highly accurate results with AUC values of more than 0.9 for the task of age prediction and 0.95 for the task of gender prediction. Golnoosh Farnadi, Jie Tang 0001, Martine De Cock, Marie-Francine Moens |
WSDM | 4 |
| 2018 | Web Search of Fashion Items with Multimodal QueryingabstractIn this paper, we introduce a novel multimodal fashion search paradigm where e-commerce data is searched with a multimodal query composed of both an image and text. In this setting, the query image shows a fashion product that the user likes and the query text allows to change certain product attributes to fit the product to the user's desire. Multimodal search gives users the means to clearly express what they are looking for. This is in contrast to current e-commerce search mechanisms, which are cumbersome and often fail to grasp the customer's needs. Multimodal search requires intermodal representations of visual and textual fashion attributes which can be mixed and matched to form the user's desired product, and which have a mechanism to indicate when a visual and textual fashion attribute represent the same concept. With a neural network, we induce a common, multimodal space for visual and textual fashion attributes where their inner product measures their semantic similarity. We build a multimodal retrieval model which operates on the obtained intermodal representations and which ranks images based on their relevance to a multimodal query. We demonstrate that our model is able to retrieve images that both exhibit the necessary query image attributes and satisfy the query texts. Moreover, we show that our model substantially outperforms two state-of-the-art retrieval models adapted to multimodal fashion search. Katrien Laenen, Susana Zoghbi, Marie-Francine Moens |
WSDM | 3 |
| 2017 | Fast and Flexible Top-k Similarity Search on Large NetworksabstractSimilarity search is a fundamental problem in network analysis and can be applied in many applications, such as collaborator recommendation in coauthor networks, friend recommendation in social networks, and relation prediction in medical information networks. In this article, we propose a sampling-based method using random paths to estimate the similarities based on both common neighbors and structural contexts efficiently in very large homogeneous or heterogeneous information networks. We give a theoretical guarantee that the sampling size depends on the error-bound ε, the confidence level (1-δ), and the path length T of each random walk. We perform an extensive empirical study on a Tencent microblogging network of 1,000,000,000 edges. We show that our algorithm can return top- k similar vertices for any vertex in a network 300× faster than the state-of-the-art methods. We develop a prototype system of recommending similar authors to demonstrate the effectiveness of our method. Jing Zhang 0001, Jie Tang 0001, Cong Ma 0001, Hanghang Tong, Yu Jing, Juan-Zi Li, Walter Luyten, Marie-Francine Moens |
ACM Trans. Inf. Syst. | 8 |
| 2016 | C-BiLDA extracting cross-lingual topics from non-parallel texts by distinguishing shared from unshared content
Geert Heyman, Ivan Vulic, Marie-Francine Moens |
Data Min. Knowl. Discov. | 3 |
| 2016 | Latent Dirichlet allocation for linking user-generated content and e-commerce data
Susana Zoghbi, Ivan Vulic, Marie-Francine Moens |
Inf. Sci. | 3 |
| 2015 | Scalable adaptive label propagation in GrappaabstractNodes of a social graph often represent entities with specific labels, denoting properties such as age-group or gender. Design of algorithms to assign labels to unlabeled nodes by leveraging node-proximity and a-priori labels of seed nodes is of significant interest. A semi-supervised approach to solve this problem is termed "LPA-Label Propagation Algorithm" where labels of a subset of nodes are iteratively propagated through the network to infer yet unknown node labels. While LPA for node labelling is extremely fast and simple, it works well only with an assumption of node-homophily — connected nodes are connected because they must deserve a similar label — which can often be a misnomer. In this paper we propose a novel algorithm "Adaptive Label Propagation" that dynamically adapts to the underlying characteristics of homophily, heterophily, or otherwise, of the connections of the network, and applies suitable label propagation strategies accordingly. Moreover, our adaptive label propagation approach is scalable as demonstrated by its implementation in Grappa, a distributed shared-memory system. Our experiments on social graphs from Facebook, YouTube, Live Journal, Orkut and Netlog demonstrate that our approach not only improves the labelling accuracy but also computes results for millions of users within a few seconds. Golnoosh Farnadi, Zeinab Mahdavifar, Ivan Keller, Jacob Nelson 0001, Ankur Teredesai, Marie-Francine Moens, Martine De Cock |
IEEE BigData | 6 |
| 2015 | Monolingual and Cross-Lingual Information Retrieval Models Based on (Bilingual) Word EmbeddingsabstractWe propose a new unified framework for monolingual (MoIR) and cross-lingual information retrieval (CLIR) which relies on the induction of dense real-valued word vectors known as word embeddings (WE) from comparable data. To this end, we make several important contributions: (1) We present a novel word representation learning model called Bilingual Word Embeddings Skip-Gram (BWESG) which is the first model able to learn bilingual word embeddings solely on the basis of document-aligned comparable data; (2) We demonstrate a simple yet effective approach to building document embeddings from single word embeddings by utilizing models from compositional distributional semantics. BWESG induces a shared cross-lingual embedding vector space in which both words, queries, and documents may be presented as dense real-valued vectors; (3) We build novel ad-hoc MoIR and CLIR models which rely on the induced word and document embeddings and the shared cross-lingual embedding space; (4) Experiments for English and Dutch MoIR, as well as for English-to-Dutch and Dutch-to-English CLIR using benchmarking CLEF 2001-2003 collections and queries demonstrate the utility of our WE-based MoIR and CLIR models. The best results on the CLEF collections are obtained by the combination of the WE-based approach and a unigram language model. We also report on significant improvements in ad-hoc IR tasks of our WE-based framework over the state-of-the-art framework for learning text representations from comparable data based on latent Dirichlet allocation (LDA). Ivan Vulic, Marie-Francine Moens |
SIGIR | 2 |
| 2015 | Probabilistic topic modeling in multilingual settings: An overview of its methodology and applications
Ivan Vulic, Wim De Smet, Jie Tang 0001, Marie-Francine Moens |
Inf. Process. Manag. | 4 |
| 2015 | Global machine learning for spatial ontology population
Parisa Kordjamshidi, Marie-Francine Moens |
J. Web Semant. | 2 |
| 2014 | Learning to bridge colloquial and formal language applied to linking and search of E-Commerce dataabstractWe study the problem of linking information between different idiomatic usages of the same language, for example, colloquial and formal language. We propose a novel probabilistic topic model called multi-idiomatic LDA (MiLDA). Its modeling principles follow the intuition that certain words are shared between two idioms of the same language, while other words are non-shared, that is, idiom-specific. We demonstrate the ability of our model to learn relations between cross-idiomatic topics in a dataset containing product descriptions and reviews. We intrinsically evaluate our model by the perplexity measure. Following that, as an extrinsic evaluation, we present the utility of the new MiLDA topic model in a recently proposed IR task of linking Pinterest pins (given in colloquial English on the users' side) to online webshops (given in formal English on the retailers' side). We show that our multi-idiomatic model outperforms the standard monolingual LDA model and the pure bilingual LDA model both in terms of perplexity and MAP scores in the IR task. Ivan Vulic, Susana Zoghbi, Marie-Francine Moens |
SIGIR | 3 |
| 2014 | Multilingual probabilistic topic modeling and its applications in web mining and searchabstractMultilingual topic models are a fairly novel group of unsupervised, language-independent and generative machine learning models. This tutorial covers all key aspects of their probabilistic framework and demonstrates how to easily integrate these models into frameworks for cross-lingual and multilingual Web mining and search. Marie-Francine Moens, Ivan Vulic |
WSDM | 1 |
| 2013 | Automatic Labeling of Forums Using Bloom's Taxonomy
Vanessa Echeverría, Juan-Carlos Gomez 0001, Marie-Francine Moens |
ADMA (1) | 3 |
| 2013 | Monolingual and Cross-Lingual Probabilistic Topic Models and Their Applications in Information Retrieval
Marie-Francine Moens, Ivan Vulic |
ECIR | 1 |
| 2013 | A Unified Framework for Monolingual and Cross-Lingual Relevance Modeling Based on Probabilistic Topic Models
Ivan Vulic, Marie-Francine Moens |
ECIR | 2 |
| 2013 | Representations for multi-document event clustering
Wim De Smet, Marie-Francine Moens |
Data Min. Knowl. Discov. | 2 |
| 2013 | Cross-language information retrieval models based on latent topic models trained with document-aligned comparable corpora
Ivan Vulic, Wim De Smet, Marie-Francine Moens |
Inf. Retr. | 3 |
| 2012 | The downside of markup: examining the harmful effects of CSS and javascript on indexing today's webabstractThe continued development and maturation of advanced HTML features such as Cascading style sheets (CSS), Javascript, and AJAX, as well as their widespread adoption by browsers, has enabled web pages to flourish with sophistication and interactivity. Unfortunately, this presents challenges to the web search community, as a web page's representation in the browser (i.e., what users see) can diverge dramatically from its raw HTML content (i.e., what search engines index and retrieve). For example, interactive pages may contain content in regions that are not visible before a user action, such as focusing a tab, but which are nonetheless still contained within the raw HTML. We study this divergence by comparing raw HTML to its fully rendered form across a number of metrics spanning presentation, geometry, and content, using a large, representative sample of popular web pages. We find that a large divergence currently exists, and we show via a historical analysis that this divergence has grown more pronounced over the last decade. The general finding of our study is that continuing to index the web via simple HTML parsing will diminish the effectiveness of retrieval on the modern web, and that the IR community should work toward more sophisticated web page processing in indexing technology. Karl Gyllstrom, Carsten Eickhoff, Arjen P. de Vries, Marie-Francine Moens |
CIKM | 4 |
| 2012 | Plink-LDA: Using Link as Prior Information in Topic Modeling
Huan Xia, Juan-Zi Li, Jie Tang 0001, Marie-Francine Moens |
DASFAA (1) | 4 |
| 2012 | EmSe: Supporting Children's Information Needs within a Hospital Environment
Leif Azzopardi, Douglas Dowie, Sergio Duarte Torres, Carsten Eickhoff, Richard Glassey, Karl Gyllstrom, Djoerd Hiemstra, Franciska de Jong, Frea Kruisinga, Kelly Ann Marshall, Marie-Francine Moens, Tamara Polajnar, Frans van der Sluis, Arjen P. de Vries |
ECIR | 11 |
| 2012 | Highly discriminative statistical features for email classification
Juan-Carlos Gomez 0001, Erik Boiy, Marie-Francine Moens |
Knowl. Inf. Syst. | 3 |
| 2011 | Examining the "leftness" property of Wikipedia categoriesabstractWikipedia's rich category structure has helped make it one of the largest semantic taxonomies in existence, a property that has been central to much recent research. However, Wikipedia's category representation is simplistic: an article contains a single list of categories, with no data about their relative importance. We investigate the ordering of category lists to determine how a category's position in the list correlates with its relevance to the article and overall significance. We identify a number of interesting connections between a category's position and its persistence within the article, age, popularity, size, and descriptiveness. Karl Gyllstrom, Marie-Francine Moens |
CIKM | 2 |
| 2011 | Clash of the Typings - Finding Controversies and Children's Topics Within Queries
Karl Gyllstrom, Marie-Francine Moens |
ECIR | 2 |
| 2011 | Knowledge Transfer across Multilingual Corpora via Latent Topics
Wim De Smet, Jie Tang 0001, Marie-Francine Moens |
PAKDD (1) | 3 |
| 2011 | Introduction to the special issue on question answering
Marie-Francine Moens, Patrick Saint-Dizier |
Inf. Process. Manag. | 1 |
| 2011 | Knowledge and reasoning for question answering: Research perspectives
Patrick Saint-Dizier, Marie-Francine Moens |
Inf. Process. Manag. | 2 |
| 2011 | A survey on question answering technology from an information retrieval perspective
Oleksandr Kolomiyets, Marie-Francine Moens |
Inf. Sci. | 2 |
| 2010 | Wisdom of the ages: toward delivering the children's web with the link-based agerank algorithmabstractThough children frequently use web search engines to learn, interact, and be entertained, modern web search engines are poorly suited to children's needs, requiring relatively complex querying and filtering of results in order to find pages oriented to young audiences. To address this limitation, we designed AgeRank, a link-based algorithm that ranks web pages according their appropriateness for young audiences. We show its effectiveness through a multipart evaluation that demonstrates AgeRank to be accurate in page-labeling, widely-spanning in page coverage, and with high potential to improve children's search. As a fast, scalable, and effective algorithm, AgeRank can be adopted by search engines seeking to more effectively address the needs of young users, or easily fitted to complementary machine-learning based classification approaches. Karl Gyllstrom, Marie-Francine Moens |
CIKM | 2 |
| 2010 | A picture is worth a thousand search results: finding child-oriented multimedia results with collAgeabstractWe present a simple and effective approach to complement search results for children’s web queries with child-oriented multimedia results, such as coloring pages and music sheets. Our approach determines appropriate media types for a query by searching Google’s database of frequent queries for co-occurrences of a query’s terms (e.g., “dinosaurs”) with preselected multimedia terms (e.g., “coloring pages”). We show the effectiveness of this approach through an online user evaluation. Karl Gyllstrom, Marie-Francine Moens |
SIGIR | 2 |
| 2010 | Linking content in unstructured sourcesabstractThis tutorial focuses on the task of automated information linking in text and multimedia sources. In any task where information is fused from different sources, this linking is a necessary step. To solve the problem we borrow methods from computational linguistics, computer vision and data mining. Although the main focus is on finding equivalence relations in the sources, the tutorial opens views on the recognition of other relation types. Marie-Francine Moens |
WWW | 1 |
| 2009 | Information Extraction and Linking in a Retrieval Context
Marie-Francine Moens, Djoerd Hiemstra |
ECIR | 1 |
| 2009 | A machine learning approach to sentiment analysis in multilingual Web texts
Erik Boiy, Marie-Francine Moens |
Inf. Retr. | 2 |
| 2008 | Finding the Best Picture: Cross-Media Retrieval of Content
Koen Deschacht, Marie-Francine Moens |
ECIR | 2 |
| 2007 | Cross-Document Entity Tracking
Roxana Angheluta, Marie-Francine Moens |
ECIR | 2 |
| 2007 | Extraction of Folksonomies from Noisy Texts
Wim De Smet, Marie-Francine Moens |
ICWSM | 2 |
| 2007 | Summarizing court decisions
Marie-Francine Moens |
Inf. Process. Manag. | 1 |
| 2006 | Rpref: a generalization of Bpref towards graded relevance judgmentsabstractWe present rpref ; our generalization of the bpref evaluation metric for assessing the quality of search engine results, given graded rather than binary user relevance judgments. Jan De Beer, Marie-Francine Moens |
SIGIR | 2 |
| 2005 | Generic technologies for single- and multi-document summarization
Marie-Francine Moens, Roxana Angheluta, Jos Dumortier |
Inf. Process. Manag. | 1 |
| 2001 | Generic Topic Segmentation of Document TextsabstractTopic segmentation is an important initial step in many text-based tasks. A hierarchical representation of a texts topics is useful in retrieval and allows judging relevancy at different levels of detail. This short paper describes research on generic algorithms for topic detection and segmentation that are applicable on texts of heterogeneous types and domains. Marie-Francine Moens, Rik De Busser |
SIGIR | 1 |
| 2000 | Text categorization: the assignment of subject descriptors to magazine articles
Marie-Francine Moens, Jos Dumortier |
Inf. Process. Manag. | 1 |
| 1999 | Abstracting of Legal Cases: The Potential of Clustering Based on the Selection of Representative ObjectsabstractThe SALOMON project automatically summarizes Belgian criminal cases in order to improve access to the large number of existing and future court decisions. SALOMON extracts text units from the case text to form a case summary. Such a case summary facilitates the rapid determination of the relevance of the case or may be employed in text search. An important part of the research concerns the development of techniques for automatic recognition of representative text paragraphs (or sentences) in texts of unrestricted domains. These techniques are employed to eliminate redundant material in the case texts, and to identify informative text paragraphs which are relevant to include in the case summary. An evaluation of a test set of 700 criminal cases demonstrates that the algorithms have an application potential for automatic indexing, abstracting, and text linking. Marie-Francine Moens, Caroline Uyttendaele, Jos Dumortier |
J. Am. Soc. Inf. Sci. | 1 |
| 1998 | Automatic Abstracting of Magazine Articles: The Creation of 'Highlight' Abstracts
Marie-Francine Moens, Jos Dumortier |
SIGIR | 1 |
| 1997 | Automatic text structuring and categorization as a first step in summarizing legal cases
Marie-Francine Moens, Caroline Uyttendaele |
Inf. Process. Manag. | 1 |