Gerard de Melo

dblp:86/1747 · DBLP profile ↗
← Back
38ranked-venue papers in the field
7as first author
9since 2021 · last 2023
0000-0002-2930-2059ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 20 (5 first)Data Mining & Knowledge Discovery · 8 (1 first)Database Systems & Data Management · 4 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 4Other / Interdisciplinary · 2
YearPublicationVenuePosition
2023 Robust NLP for Finance (RobustFin)
abstract
Natural language processing (NLP) technologies have been widely applied in business domains such as e-commerce and customer service, but their adoption in the financial sector has been constrained by industry-specific performance standards and regulatory restrictions. This challenge has created new opportunities for core research in related areas. Recent advancements in NLP, such as the advent of large language models, has encouraged adoption in the finance sector. However, compared to other domains, finance has stricter requirements for robustness, explainability, and generalizability. Given this background, we propose to organize the first Robust NLP for Finance (RobustFin) workshop at KDD '23 to encourage the study of and research on robustness and explainability technologies with regard to financial NLP. The goal of the workshop is to extend the applications of NLP in finance, while motivating further research in robust NLP.
Sameena Shah, Xiaodan Zhu 0001, Gerard de Melo, Armineh Nourbakhsh, Xiaomo Liu, Charese Smiley, Zhiyu Chen 0002
KDD3
2023 Fifth Knowledge-aware and Conversational Recommender Systems Workshop (KaRS)
abstract
Recommender systems have become ubiquitous in daily life, but their limitations in interacting with human users have become evident. Deep learning approaches have led to the development of data-driven algorithms that identify connections between users and items, but they often miss a critical actor in the loop - the end-user. Knowledge-based approaches are gaining attention due to the availability of knowledge-graphs, such as DBpedia and Wikidata, which provide semantics-aware information on different knowledge domains. These approaches are being used for recommendation and challenges such as knowledge graph embeddings, hybrid recommendation, and interpretable recommendation. Moreover, the emergence of neural-symbolic systems, which combine data-driven and symbolic methods, can significantly improve recommendation systems. A growing number of research papers on such topics demonstrate the growing interest and research potential of these systems. Furthermore, content features become crucial when interaction requires it. The development of conversational recommender systems presents new challenges, as they require multi-turn dialogues between users and systems, blurring the line between recommendation and retrieval. Evaluation of these systems goes beyond simple accuracy metrics and is hampered by the limited availability of datasets. While research and development into conversational recommender systems has been less prominent in the past, recent literature shows growing interest and potential for these systems.
Vito Walter Anelli, Pierpaolo Basile, Gerard de Melo, Francesco M. Donini, Antonio Ferrara 0001, Cataldo Musto, Fedelucio Narducci, Azzurra Ragone, Markus Zanker
RecSys3
2022 D-HYPR: Harnessing Neighborhood Modeling and Asymmetry Preservation for Digraph Representation Learning
abstract
Digraph Representation Learning (DRL) aims to learn representations for directed homogeneous graphs (digraphs). Prior work in DRL is largely constrained (e.g., limited to directed acyclic graphs), or has poor generalizability across tasks (e.g., evaluated solely on one task). Most Graph Neural Networks (GNNs) exhibit poor performance on digraphs due to the neglect of modeling neighborhoods and preserving asymmetry. In this paper, we address these notable challenges by leveraging hyperbolic collaborative learning from multi-ordered and partitioned neighborhoods, and regularizers inspired by socio-psychological factors. Our resulting formalism, Digraph Hyperbolic Networks (D-HYPR) -- albeit conceptually simple -- generalizes to digraphs where cycles and non-transitive relations are common, and is applicable to multiple downstream tasks including node classification, link presence prediction, and link property prediction. In order to assess the effectiveness of D-HYPR, extensive evaluations were performed across 8 real-world digraph datasets involving 21 prior techniques. D-HYPR statistically significantly outperforms the current state of the art. We release our code at https://github.com/hongluzhou/dhypr
Honglu Zhou, Advith Chegu, Samuel S. Sohn, Zuohui Fu, Gerard de Melo, Mubbasir Kapadia
CIKM5
2022 Temporal Event Reasoning Using Multi-source Auxiliary Learning Objectives
Xin Dong 0010, Tanay Kumar Saha, Ke Zhang 0013, Joel R. Tetreault, Alejandro Jaimes, Gerard de Melo
ECIR (2)6
2022 Fourth Knowledge-aware and Conversational Recommender Systems Workshop (KaRS)
abstract
In the last few years, a renewed interest of the research community in conversational recommender systems (CRSs) has been emerging. This is likely due to the massive proliferation of Digital Assistants (DAs) such as Amazon Alexa, Siri, or Google Assistant that are revolutionizing the way users interact with machines. DAs allow users to execute a wide range of actions through an interaction mostly based on natural language utterances. However, although DAs are able to complete tasks such as sending texts, making phone calls, or playing songs, they still remain at an early stage in terms of their recommendation capabilities via a conversation. In addition, we have been witnessing the advent of increasingly precise and powerful recommendation algorithms and techniques able to effectively assess users’ tastes and predict information that may be of interest to them. Most of these approaches rely on the collaborative paradigm (often exploiting machine learning techniques) and neglect the huge amount of knowledge, both structured and unstructured, describing the domain of interest of a recommendation engine. Although very effective in predicting relevant items, collaborative approaches miss some very interesting features that go beyond the accuracy of results and move in the direction of providing novel and diverse results as well as generating explanations for recommended items. Knowledge-aware side information becomes crucial when a conversational interaction is implemented, in particular for preference elicitation, explanation, and critiquing steps.
Vito Walter Anelli, Pierpaolo Basile, Gerard de Melo, Francesco M. Donini, Antonio Ferrara 0001, Cataldo Musto, Fedelucio Narducci, Azzurra Ragone, Markus Zanker
RecSys3
2022 Path Language Modeling over Knowledge Graphsfor Explainable Recommendation
abstract
To facilitate human decisions with credible suggestions, personalized recommender systems should have the ability to generate corresponding explanations while making recommendations. Knowledge graphs (KG), which contain comprehensive information about users and products, are widely used to enable this. By reasoning over a KG in a node-by-node manner, existing explainable models provide a KG-grounded path for each user-recommended item. Such paths serve as an explanation and reflect the historical behavior pattern of the user. However, not all items can be reached following the connections within the constructed KG under finite hops. Hence, previous approaches are constrained by a recall bias in terms of existing connectivity of KG structures. To overcome this, we propose a novel Path Language Modeling Recommendation (PLM-Rec) framework, learning a language model over KG paths consisting of entities and edges. Through path sequence decoding, PLM-Rec unifies recommendation and explanation in a single step and fulfills them simultaneously. As a result, PLM-Rec not only captures the user behaviors but also eliminates the restriction to pre-existing KG connections, thereby alleviating the aforementioned recall bias. Moreover, the proposed technique makes it possible to conduct explainable recommendation even when the KG is sparse or possesses a large number of relations. Experiments and extensive ablation studies on three Amazon e-commerce datasets demonstrate the effectiveness and explainability of the PLM-Rec framework.
Shijie Geng, Zuohui Fu, Juntao Tan, Yingqiang Ge, Gerard de Melo, Yongfeng Zhang 0003
WWW5
2021 Popcorn: Human-in-the-loop Popularity Debiasing in Conversational Recommender Systems
abstract
Recent conversational recommender systems (CRS) provide a promising solution to accurately capture a user's preferences by communicating with users in natural language to interactively guide them while pro-actively eliciting their current interests. Previous research on this mainly focused on either learning a supervised model with semantic features extracted from the user's responses, or training a policy network to control the dialogue state. However, none of them has considered the issue of popularity bias in a CRS. This paper proposes a human-in-the-loop popularity debiasing framework that integrates real-time semantic understanding of open-ended user utterances as well as historical records, while also effectively managing the dialogue with the user. This allows the CRS to balance the recommendation performance as well as the item popularity so as to avoid the well-known "long-tail'' effect. We demonstrate the effectiveness of our approach via experiments on two conversational recommendation datasets, and the results confirm that our proposed approach achieves high-accuracy recommendation while mitigating popularity bias.
Zuohui Fu, Yikun Xian, Shijie Geng, Gerard de Melo, Yongfeng Zhang 0003
CIKM4
2021 EXACTA: Explainable Column Annotation
abstract
Column annotation, the process of annotating tabular columns with labels, plays a fundamental role in digital marketing data governance. It has a direct impact on how customers manage their data and facilitates compliance with regulations, restrictions, and policies applicable to data use. Despite substantial gains in accuracy brought by recent deep learning-driven column annotation methods, their incapability of explaining why columns are matched with particular target labels has drawn concern, due to the black-box nature of deep neural networks. Such explainability is of particular importance in industrial marketing scenarios, where data stewards need to quickly verify and calibrate the annotation results to ascertain the correctness of downstream applications. This work sheds new light on the explainable column annotation problem, the first of its kind column annotation task. To achieve this, we propose a new approach called EXACTA, which conducts multi-hop knowledge graph reasoning using inverse reinforcement learning to find a path from a column to a potential target label while ensuring both annotation performance and explainability. We experiment on four benchmarks, both publicly available and real-world ones, and undertake a comprehensive analysis on the explainability. The results suggest that our method not only provides competitive annotation performance compared with existing deep learning-based models, but more importantly, produces faithfully explainable paths for annotated columns to facilitate human examination.
Yikun Xian, Handong Zhao, Tak Yeon Lee, Sungchul Kim, Ryan Rossi, Zuohui Fu, Gerard de Melo, S. Muthukrishnan 0001
KDD7
2021 HOOPS: Human-in-the-Loop Graph Reasoning for Conversational Recommendation
abstract
There is increasing recognition of the need for human-centered AI that learns from human feedback. However, most current AI systems focus more on the model design, but less on human participation as part of the pipeline. In this work, we propose a Human-in-the-Loop (HitL) graph reasoning paradigm and develop a corresponding dataset named HOOPS for the task of KG-driven conversational recommendation. Specifically, we first construct a KG interpreting diverse user behaviors and identify pertinent attribute entities for each user--item pair. Then we simulate the conversational turns reflecting the human decision making process of choosing suitable items tracing the KG structures transparently. We also provide a benchmark method with reported performance on the dataset to ascertain the feasibility of HitL graph reasoning for recommendation using our developed dataset, and show that it provides novel opportunities for the research community.
Zuohui Fu, Yikun Xian, Yaxin Zhu, Zelong Li 0001, Gerard de Melo, Yongfeng Zhang 0003
SIGIR6
2020 CAFE: Coarse-to-Fine Neural Symbolic Reasoning for Explainable Recommendation
abstract
Recent research explores incorporating knowledge graphs (KG) into e-commerce recommender systems, not only to achieve better recommendation performance, but more importantly to generate explanations of why particular decisions are made. This can be achieved by explicit KG reasoning, where a model starts from a user node, sequentially determines the next step, and walks towards an item node of potential interest to the user. However, this is challenging due to the huge search space, unknown destination, and sparse signals over the KG, so informative and effective guidance is needed to achieve a satisfactory recommendation quality. To this end, we propose a CoArse-to-FinE neural symbolic reasoning approach (CAFE). It first generates user profiles as coarse sketches of user behaviors, which subsequently guide a path-finding process to derive reasoning paths for recommendations as fine-grained predictions. User profiles can capture prominent user behaviors from the history, and provide valuable signals about which kinds of path patterns are more likely to lead to potential items of interest for the user. To better exploit the user profiles, an improved path-finding algorithm called Profile-guided Path Reasoning (PPR) is also developed, which leverages an inventory of neural symbolic reasoning modules to effectively and efficiently find a batch of paths over a large-scale KG. We extensively experiment on four real-world benchmarks and observe substantial gains in the recommendation performance compared with state-of-the-art methods.
Yikun Xian, Zuohui Fu, Handong Zhao, Yingqiang Ge, Xu Chen 0017, Qiaoying Huang, Shijie Geng, Zhou Qin 0001, Gerard de Melo, S. Muthukrishnan 0001, Yongfeng Zhang 0003
CIKM9
2020 Explainable Link Prediction for Emerging Entities in Knowledge Graphs
Rajarshi Bhowmik, Gerard de Melo
ISWC (1)2
2020 Leveraging Adversarial Training in Self-Learning for Cross-Lingual Text Classification
abstract
In cross-lingual text classification, one seeks to exploit labeled data from one language to train a text classification model that can then be applied to a completely different language. Recent multilingual representation models have made it much easier to achieve this. Still, there may still be subtle differences between languages that are neglected when doing so. To address this, we present a semi- supervised adversarial training process that minimizes the maximal loss for label-preserving input perturbations. The resulting model then serves as a teacher to induce labels for unlabeled target lan- guage samples that can be used during further adversarial training, allowing us to gradually adapt our model to the target language. Compared with a number of strong baselines, we observe signifi- cant gains in effectiveness on document and intent classification for a diverse set of languages.
Xin Dong 0010, Yaxin Zhu, Zuohui Fu, Dongkuan Xu, Sen Yang 0002, Gerard de Melo
SIGIR7
2020 Fairness-Aware Explainable Recommendation over Knowledge Graphs
abstract
There has been growing attention on fairness considerations recently, especially in the context of intelligent decision making systems. For example, explainable recommendation systems may suffer from both explanation bias and performance disparity. We show that inactive users may be more susceptible to receiving unsatisfactory recommendations due to their insufficient training data, and that their recommendations may be biased by the training records of active users due to the nature of collaborative filtering, which leads to unfair treatment by the system. In this paper, we analyze different groups of users according to their level of activity, and find that bias exists in recommendation performance between different groups. Empirically, we find that such performance gap is caused by the disparity of data distribution, specifically the knowledge graph path distribution in this work. We propose a fairness constrained approach via heuristic re-ranking to mitigate this unfairness problem in the context of explainable recommendation over knowledge graphs. We experiment on several real-world datasets with state-of-the-art knowledge graph-based explainable recommendation algorithms. The promising results show that our algorithm is not only able to provide high-quality explainable recommendations, but also reduces the recommendation unfairness in several aspects.
Zuohui Fu, Yikun Xian, Ruoyuan Gao, Jieyu Zhao 0004, Qiaoying Huang, Yingqiang Ge, Shijie Geng, Chirag Shah 0001, Yongfeng Zhang 0003, Gerard de Melo
SIGIR11
2020 Illustrate Your Story: Enriching Text with Images
abstract
Human perception is known to be predominantly visual. As modern web infrastructure promoted the storage of media, the web-data paradigm shifted from text-only documents to those containing text and images. A multitude of blog posts, news articles, and social media posts exist on the Internet today as examples of multimodal stories. The manual alignment of images and text in a story is time-consuming and labor intensive. We present a web application for automatically selecting relevant images from an album and placing them in suitable contexts within a body of text. The application solves a global optimization problem that maximizes the coherence of text paragraphs and image descriptors, and allows for exploring the underlying image descriptors and similarity metrics. Experiments show that our method can align images with texts with high semantic fit, and to user satisfaction.
Sreyasi Nag Chowdhury, William Cheng, Gerard de Melo, Simon Razniewski, Gerhard Weikum
WSDM3
2020 What Sparks Joy: The AffectVec Emotion Database
abstract
Affective analysis of textual data is instrumental in understanding human communication in the modern era of social media. A number of resources have been proposed in attempts to characterize the emotions tied to words in a text. In this work, we show that we can obtain a database that goes beyond the common binary scores for emotion classification provided by past work. Instead, we harness the power of Big Data by using neural vector space models trained with large-scale supervision from co-occurrence patterns. We modify the vector space to better account for emotional associations, which then enables us to induce AffectVec, a new emotion database providing graded emotion intensity scores for English language words with regard to a fine-grained inventory of over 200 different emotion categories. Our experiments show that AffectVec outperforms existing emotion lexicons by substantial margins in intrinsic evaluations as well as for affective text classification.
Shahab Raji, Gerard de Melo
WWW2
2019 Reinforcement Knowledge Graph Reasoning for Explainable Recommendation
abstract
Recent advances in personalized recommendation have sparked great interest in the exploitation of rich structured information provided by knowledge graphs. Unlike most existing approaches that only focus on leveraging knowledge graphs for more accurate recommendation, we aim to conduct explicit reasoning with knowledge for decision making so that the recommendations are generated and supported by an interpretable causal inference procedure. To this end, we propose a method called Policy-Guided Path Reasoning (PGPR), which couples recommendation and interpretability by providing actual paths in a knowledge graph. Our contributions include four aspects. We first highlight the significance of incorporating knowledge graphs into recommendation to formally define and interpret the reasoning process. Second, we propose a reinforcement learning (RL) approach featured by an innovative soft reward strategy, user-conditional action pruning and a multi-hop scoring function. Third, we design a policy-guided graph search algorithm to efficiently and effectively sample reasoning paths for recommendation. Finally, we extensively evaluate our method on several large-scale real-world benchmark datasets, obtaining favorable results compared with state-of-the-art methods.
Yikun Xian, Zuohui Fu, S. Muthukrishnan 0001, Gerard de Melo, Yongfeng Zhang 0003
SIGIR4
2019 Be Concise and Precise: Synthesizing Open-Domain Entity Descriptions from Facts
abstract
Despite being vast repositories of factual information, cross-domain knowledge graphs, such as Wikidata and the Google Knowledge Graph, only sparsely provide short synoptic descriptions for entities. Such descriptions that briefly identify the most discernible features of an entity provide readers with a near-instantaneous understanding of what kind of entity they are being presented. They can also aid in tasks such as named entity disambiguation, ontological type determination, and answering entity queries. Given the rapidly increasing numbers of entities in knowledge graphs, a fully automated synthesis of succinct textual descriptions from underlying factual information is essential. To this end, we propose a novel fact-to-sequence encoder-decoder model with a suitable copy mechanism to generate concise and precise textual descriptions of entities. In an in-depth evaluation, we demonstrate that our method significantly outperforms state-of-the-art alternatives.
Rajarshi Bhowmik, Gerard de Melo
WWW2
2018 Five Shades of Untruth: Finer-Grained Classification of Fake News
abstract
Prior work on algorithmic truth assessment on unreliable content, has mostly pursued binary classifiers - factual vs. fake - and disregarded the finer shades of untruth. On the other hand, manual analysis of questionable content has proposed a more fine-grained classification: distinguishing between hoaxes, irony and propaganda, or the six-way rating by the PolitiFact community. In this paper, we present a principled approach to capture these finer shades in automatically assessing and classifying news articles and claims. We systematically explore a variety of signals from both news and social media, and give an analysis of the underlying features.
Yafang Wang, Gerard de Melo, Gerhard Weikum
ASONAM3
2018 Social Media vs. News Media: Analyzing Real-World Events from Different Perspectives
Yafang Wang, Zeyuan Cui, Shijun Liu, Gerard de Melo
DEXA (2)6
2018 Visualizing Multi-document Semantics via Open Domain Information Extraction
Yongpan Sheng, Zenglin Xu, Yafang Wang, Zhonghui You, Gerard de Melo
ECML/PKDD (3)7
2018 Co-PACRR: A Context-Aware Neural IR Model for Ad-hoc Retrieval
abstract
Neural IR models, such as DRMM and PACRR, have achieved strong results by successfully capturing relevance matching signals. We argue that the context of these matching signals is also important. Intuitively, when extracting, modeling, and combining matching signals, one would like to consider the surrounding text(local context) as well as other signals from the same document that can contribute to the overall relevance score. In this work, we highlight three potential shortcomings caused by not considering context information and propose three neural ingredients to address them: a disambiguation component, cascade k-max pooling, and a shuffling combination layer. Incorporating these components into the PACRR model yields Co-PACER, a novel context-aware neural IR model. Extensive comparisons with established models on TREC Web Track data confirm that the proposed model can achieve superior search results. In addition, an ablation analysis is conducted to gain insights into the impact of and interactions between different components. We release our code to enable future comparisons.
Kai Hui 0001, Andrew Yates, Klaus Berberich, Gerard de Melo
WSDM4
2016 Summary Generation for Temporal Extractions
Yafang Wang, Zhaochun Ren, Martin Theobald, Maximilian Dylla, Gerard de Melo
DEXA (1)5
2016 ShapeExplorer: Querying and Exploring Shapes using Visual Knowledge
Tong Ge, Yafang Wang, Gerard de Melo, Zengguang Hao, Andrei Sharf, Baoquan Chen
EDBT3
2016 Heuristics for Connecting Heterogeneous Knowledge via FrameBase
Jacobo Rouces, Gerard de Melo, Katja Hose
ESWC2
2016 WebBrain: Joint Neural Learning of Large-Scale Commonsense Knowledge
Niket Tandon, Charles Hariman, Gerard de Melo
ISWC (1)4
2015 Knowlywood: Mining Activity Knowledge From Hollywood Narratives
abstract
Despite the success of large knowledge bases, one kind of knowledge that has not received attention so far is that of human activities. An example of such an activity is proposing to someone (to get married). For the computer, knowing that this involves two adults, often but not necessarily a woman and a man, that it often takes place in some romantic location, that it typically involves flowers or jewelry, and that it is usually followed by kissing, is a valuable asset for tasks like natural language dialog, scene understanding, or video search.
Niket Tandon, Gerard de Melo, Abir De, Gerhard Weikum
CIKM2
2015 Scalable Learning Technologies for Big Data Mining
Gerard de Melo, Aparna S. Varde
DASFAA (2)1
2015 FrameBase: Representing N-Ary Relations Using Semantic Frames
Jacobo Rouces, Gerard de Melo, Katja Hose
ESWC2
2014 PIKM 2014: The 7th ACM Workshop for Ph.D. Students in Information and Knowledge Management
abstract
PIKM workshop offers to Ph.D. students the possibility to bring their work to an international and interdisciplinary research community, and create a network of young researchers to exchange and develop new and promising ideas. Similarly to the CIKM, PIKM workshop covers a wide range of topics in the areas of databases, information retrieval and knowledge management.
Gerard de Melo, Mouna Kacimi, Aparna S. Varde
CIKM1
2014 Embedding NomLex-BR nominalizations into OpenWordnet-PT
abstract
This paper presents NomLex-BR, a lexical resource describing Brazilian Portuguese nominalizations, and its integration with OpenWordnet-PT.We first describe the original English NOMLEX lexical resource and how we used it to bootstrap a Portuguese version.Subsequently, we describe how this lexicon can be embedded into OpenWordnet-PT, which facilitates its use and helps spot-checking both the bigger integrated resource and the original lexicon.Lastly, we outline some of the other, more substantial work that we plan to engage for the project of using linguistic insights for knowledge representation in Portuguese.
Alexandre Rademaker, Valeria de Paiva, Gerard de Melo, Livy Real
GWC3
2014 OpenWordNet-PT: A Project Report
abstract
This paper presents OpenWordNet-PT, a freely available open-source wordnet for Portuguese, with its latest developments and practical uses.We provide a detailed description of the RDF representation developed for OpenWordnet-PT.We highlight our efforts to extend the coverage of our resource and add nominalization relations connecting nouns and verbs.Finally, we present several real-world applications where OpenWordnet-PT was put to use, including a large-scale high-throughput sentiment analysis system.
Alexandre Rademaker, Valeria de Paiva, Gerard de Melo, Livy Real, Maíra Gatti de Bayser
GWC3
2014 WebChild: harvesting and organizing commonsense knowledge from the web
abstract
This paper presents a method for automatically constructing a large commonsense knowledge base, called WebChild, from Web contents. WebChild contains triples that connect nouns with adjectives via fine-grained relations like hasShape, hasTaste, evokesEmotion, etc. The arguments of these assertions, nouns and adjectives, are disambiguated by mapping them onto their proper WordNet senses. Our method is based on semi-supervised Label Propagation over graphs of noisy candidate assertions. We automatically derive seeds from WordNet and by pattern matching from Web text collections. The Label Propagation algorithm provides us with domain sets and range sets for 19 different relations, and with confidence-ranked assertions between WordNet senses. Large-scale experiments demonstrate the high accuracy (more than 80 percent) and coverage (more than four million fine grained disambiguated assertions) of WebChild.
Niket Tandon, Gerard de Melo, Fabian M. Suchanek, Gerhard Weikum
WSDM2
2014 Taxonomic data integration from multilingual Wikipedia editions
Gerard de Melo, Gerhard Weikum
Knowl. Inf. Syst.1
2013 Searching the Web of Data
Gerard de Melo, Katja Hose
ECIR1
2012 LINDA: distributed web-of-data-scale entity matching
abstract
Linked Data has emerged as a powerful way of interconnecting structured data on the Web. However, the cross-linkage between Linked Data sources is not as extensive as one would hope for. In this paper, we formalize the task of automatically creating "sameAs" links across data sources in a globally consistent manner. Our algorithm, presented in a multi-core as well as a distributed version, achieves this link generation by accounting for joint evidence of a match. Experiments confirm that our system scales beyond 100 million entities and delivers highly accurate results despite the vast heterogeneity and daunting scale.
Christoph Böhm 0001, Gerard de Melo, Felix Naumann, Gerhard Weikum
CIKM2
2010 MENTA: inducing multilingual taxonomies from wikipedia
abstract
In recent years, a number of projects have turned to Wikipedia to establish large-scale taxonomies that describe orders of magnitude more entities than traditional manually built knowledge bases. So far, however, the multilingual nature of Wikipedia has largely been neglected. This paper investigates how entities from all editions of Wikipedia as well as WordNet can be integrated into a single coherent taxonomic class hierarchy. We rely on linking heuristics to discover potential taxonomic relationships, graph partitioning to form consistent equivalence classes of entities, and a Markov chain-based ranking approach to construct the final taxonomy. This results in MENTA (Multilingual Entity Taxonomy), a resource that describes 5.4 million entities and is presumably the largest multilingual lexical knowledge base currently available.
Gerard de Melo, Gerhard Weikum
CIKM1
2009 Towards a universal wordnet by learning from combined evidence
abstract
Lexical databases are invaluable sources of knowledge about words and their meanings, with numerous applications in areas like NLP, IR, and AI. We propose a methodology for the automatic construction of a large-scale multilingual lexical database where words of many languages are hierarchically organized in terms of their meanings and their semantic relations to other words. This resource is bootstrapped from WordNet, a well-known English-language resource. Our approach extends WordNet with around 1.5 million meaning links for 800,000 words in over 200 languages, drawing on evidence extracted from a variety of resources including existing (monolingual) wordnets, (mostly bilingual) translation dictionaries, and parallel corpora. Graph-based scoring functions and statistical learning techniques are used to iteratively integrate this information and build an output graph. Experiments show that this wordnet has a high level of precision and coverage, and that it can be useful in applied tasks such as cross-lingual text classification.
Gerard de Melo, Gerhard Weikum
CIKM1
2007 Multilingual Text Classification Using Ontologies
Gerard de Melo, Stefan Siersdorfer
ECIR1