VLDB 2026 Research / reviewers in the wild / expert
Martine De Cock
dblp:79/4249
· DBLP profile ↗
97ranked-venue papers
10as first author
8since 2021 · last 2025
0000-0001-7917-0771ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 8 first-author · 5 since 2021Databases, data management, data science and information retrieval · 23 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 1 first-authorTheory of computation · 8Software engineering, systems software and programming languages · 6Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 since 2021Security and privacy · 4 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Enhancing Privacy in the Early Detection of Sexual Predators Through Federated Learning and Differential PrivacyabstractThe increased screen time and isolation caused by the COVID-19 pandemic have led to a significant surge in cases of online grooming, which is the use of strategies by predators to lure children into sexual exploitation. Previous efforts to detect grooming in industry and academia have involved accessing and monitoring private conversations through centrally-trained models or sending private conversations to a global server. In this work, we implement a privacy-preserving pipeline for the early detection of sexual predators. We leverage federated learning and differential privacy in order to create safer online spaces for children while respecting their privacy. We investigate various privacy-preserving implementations and discuss their benefits and shortcomings. Our extensive evaluation using real-world data proves that privacy and utility can coexist with only a slight reduction in utility. Khaoula Chehbouni, Martine De Cock, Gilles Caporossi, Afaf Taïk, Reihaneh Rabbany, Golnoosh Farnadi |
AAAI | 2 |
| 2024 | CaPS: Collaborative and Private Synthetic Data Generation from Distributed SourcesabstractData is the lifeblood of the modern world, forming a fundamental part of AI, decision-making, and research advances. With increase in interest in data, governments have taken important steps towards a regulated data world, drastically impacting data sharing and data usability and resulting in massive amounts of data confined within the walls of organizations. While synthetic data generation (SDG) is an appealing solution to break down these walls and enable data sharing, the main drawback of existing solutions is the assumption of a trusted aggregator for generative model training. Given that many data holders may not want to, or be legally allowed to, entrust a central entity with their raw data, we propose a framework for collaborative and private generation of synthetic tabular data from distributed data holders. Our solution is general, applicable to any marginal-based SDG, and provides input privacy by replacing the trusted aggregator with secure multi-party computation (MPC) protocols and output privacy via differential privacy (DP). We demonstrate the applicability and scalability of our approach for the state-of-the-art select-measure-generate SDG algorithms MWEM+PGM and AIM. Sikha Pentyala, Mayana Pereira, Martine De Cock |
ICML | 3 |
| 2024 | Privacy-Preserving Membership Queries for Federated Anomaly DetectionabstractIn this work, we propose a new privacy-preserving membership query protocol that lets a centralized entity privately query datasets held by one or more other parties to check if they contain a given element. This protocol, based on elliptic curve-based ElGamal and oblivious key-value stores, ensures that those 'data-augmenting' parties only have to send their encrypted data to the centralized entity once, making the protocol particularly efficient when the centralized entity repeatedly queries the same sets of data. We apply this protocol to detect anomalies in cross-silo federations. Data anomalies across such cross-silo federations are challenging to detect because (1) the centralized entities have little knowledge of the actual users, (2) the data-augmenting entities do not have a global view of the system, and (3) privacy concerns and regulations prevent pooling all the data. Our protocol allows for anomaly detection even in strongly separated distributed systems while protecting users' privacy. Specifically, we propose a cross-silo federated architecture in which a centralized entity (the backbone) has labeled data to train a machine learning model for detecting anomalous instances. The other entities in the federation are data-augmenting clients (the user-facing entities) who collaborate with the centralized entity to extract feature values to improve the utility of the model. These feature values are computed using our privacy-preserving membership query protocol. The model can be trained with an off-the-shelf machine learning algorithm that provides differential privacy to prevent it from memorizing instances from the training data, thereby providing output privacy. However, it is not straightforward to also efficiently provide input privacy, which ensures that none of the entities in the federation ever see the data of other entities in an unencrypted form. We demonstrate the effectiveness of our approach in the financial domain, motivated by the PETs Prize Challenge, which is a collaborative effort between the US and UK governments to combat international fraudulent transactions. We show that the private queries significantly increase the precision and recall of the otherwise centralized system and argue that this improvement translates to other use cases as well. Jelle Vos, Sikha Pentyala, Steven Golob, Ricardo Maia 0001, Dean F. Kelley, Zekeriya Erkin, Martine De Cock, Anderson Nascimento |
Proc. Priv. Enhancing Technol. | 7 |
| 2023 | Privacy-Preserving Fair Item Ranking
Jia Ao Sun, Sikha Pentyala, Martine De Cock, Golnoosh Farnadi |
ECIR (2) | 3 |
| 2023 | Secure Multi-Party Computation for Personalized Human Activity Recognition
David Melanson, Ricardo Maia 0001, Hee-Seok Kim, Anderson C. A. Nascimento, Martine De Cock |
Neural Process. Lett. | 5 |
| 2022 | Privacy-preserving training of tree ensembles over continuous data
Samuel Adams, Chaitali Choudhary, Martine De Cock, Rafael Dowsley, David Melanson, Anderson C. A. Nascimento, Davis Railsback, Jianwei Shen 0002 |
Proc. Priv. Enhancing Technol. | 3 |
| 2021 | Privacy-Preserving Feature Selection with Secure Multiparty ComputationabstractExisting work on privacy-preserving machine learning with Secure Multiparty Computation (MPC) is almost exclusively focused on model training and on inference with trained models, thereby overlooking the important data pre-processing stage. In this work, we propose the first MPC based protocol for private feature selection based on the filter method, which is independent of model training, and can be used in combination with any MPC protocol to rank features. We propose an efficient feature scoring protocol based on Gini impurity to this end. To demonstrate the feasibility of our approach for practical data science, we perform experiments with the proposed MPC protocols for feature selection in a commonly used machine-learning-as-a-service configuration where computations are outsourced to multiple servers, with semi-honest and with malicious adversaries. Regarding effectiveness, we show that secure feature selection with the proposed protocols improves the accuracy of classifiers on a variety of real-world data sets, without leaking information about the feature values or even which features were selected. Regarding efficiency, we document runtimes ranging from several seconds to an hour for our protocols to finish, depending on the size of the data set and the security settings. Xiling Li, Rafael Dowsley, Martine De Cock |
ICML | 3 |
| 2021 | Privacy-Preserving Video Classification with Convolutional Neural NetworksabstractMany video classification applications require access to personal data, thereby posing an invasive security risk to the users’ privacy. We propose a privacy-preserving implementation of single-frame method based video classification with convolutional neural networks that allows a party to infer a label from a video without necessitating the video owner to disclose their video to other entities in an unencrypted manner. Similarly, our approach removes the requirement of the classifier owner from revealing their model parameters to outside entities in plaintext. To this end, we combine existing Secure Multi-Party Computation (MPC) protocols for private image classification with our novel MPC protocols for oblivious single-frame selection and secure label aggregation across frames. The result is an end-to-end privacy-preserving video classification pipeline. We evaluate our proposed solution in an application for private human emotion recognition. Our results across a variety of security settings, spanning honest and dishonest majority configurations of the computing parties, and for both passive and active adversaries, demonstrate that videos can be classified with state-of-the-art accuracy, and without leaking sensitive user information. Sikha Pentyala, Rafael Dowsley, Martine De Cock |
ICML | 3 |
| 2020 | Malicious DNS Tunneling Detection in Real-Traffic DNS DataabstractWhile originally not intended for data transfer, the Domain Name System (DNS) is currently used to this end anyway, in a process called DNS tunneling (DNST). Malicious users exploit DNST for data exfiltration from infected machines, posing a critical security threat. We train and evaluate state-of-the-art convolutional neural network, random forest, and ensemble classifiers to detect tunneling in DNS traffic. Finally, we assess the classifiers' performance and robustness by exposing them to one day of real-traffic data. Danielle Lambion, Michael Josten, Femi G. Olumofin, Martine De Cock |
IEEE BigData | 4 |
| 2019 | Hardening DGA Classifiers Utilizing IVAPabstractDomain Generation Algorithms (DGAs) are used by malware to generate a deterministic set of domains, usually by utilizing a pseudo-random seed. A malicious botmaster can establish connections between their command-and-control center (C&C) and any malware-infected machines by registering domains that will be DGA-generated given a specific seed, rendering traditional domain blacklisting ineffective. Given the nature of this threat, the real-time detection of DGA domains based on incoming DNS traffic is highly important. The use of neural network machine learning (ML) models for this task has been well-studied, but there is still substantial room for improvement. In this paper, we propose to use Inductive Venn-Abers predictors (IVAPs) to calibrate the output of existing ML models for DGA classification. The IVAP is a computationally efficient procedure which consistently improves the predictive accuracy of classifiers at the expense of not offering predictions for a small subset of inputs and consuming an additional amount of training data. Charles Grumer, Jonathan Peck, Femi G. Olumofin, Anderson C. A. Nascimento, Martine De Cock |
IEEE BigData | 5 |
| 2019 | Privacy-Preserving Classification of Personal Text Messages with Secure Multi-Party ComputationabstractClassification of personal text messages has many useful applications in surveillance, e-commerce, and mental health care, to name a few. Giving applications access to personal texts can easily lead to (un)intentional privacy violations. We propose the first privacy-preserving solution for text classification that is provably secure. Our method, which is based on Secure Multiparty Computation (SMC), encompasses both feature extraction from texts, and subsequent classification with logistic regression and tree ensembles. We prove that when using our secure text classification method, the application does not learn anything about the text, and the author of the text does not learn anything about the text classification model used by the application beyond what is given by the classification result itself. We perform end-to-end experiments with an application for detecting hate speech against women and immigrants, demonstrating excellent runtime results without loss of accuracy. Devin Reich, Ariel Todoki, Rafael Dowsley, Martine De Cock, Anderson C. A. Nascimento |
NeurIPS | 4 |
| 2019 | Efficient and Private Scoring of Decision Trees, Support Vector Machines and Logistic Regression Models Based on Pre-ComputationabstractMany data-driven personalized services require that private data of users is scored against a trained machine learning model. In this paper we propose a novel protocol for privacy-preserving classification of decision trees, a popular machine learning model in these scenarios. Our solutions is composed out of building blocks, namely a secure comparison protocol, a protocol for obliviously selecting inputs, and a protocol for multiplication. By combining some of the building blocks for our decision tree classification protocol, we also improve previously proposed solutions for classification of support vector machines and logistic regression models. Our protocols are information theoretically secure and, unlike previously proposed solutions, do not require modular exponentiations. We show that our protocols for privacy-preserving classification lead to more efficient results from the point of view of computational and communication complexities. We present accuracy and runtime results for seven classification benchmark datasets from the UCI repository. Martine De Cock, Rafael Dowsley, Caleb Horst, Rajendra S. Katti, Anderson C. A. Nascimento, Wing-Sea Poon, Stacey Truex |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2018 | Privacy-Preserving Linear Regression for Brain-Computer Interface ApplicationsabstractMany machine learning (ML) applications rely on large amounts of personal data for training and inference. Among the most intimate exploited data sources is electroencephalogram (EEG) data. The emergence of consumer -grade, low-cost brain -computer interfaces (BCIs) and corresponding software development kits' is bringing the use of BCI within reach of application developers. The access that BCI applications have to neural signals rightly raises privacy concerns. Application developers can easily gain knowledge beyond the professed scope from unprotected EEG signals, including passwords, ATM PINs, and other personal data. The challenge is how to engage in meaningful ML with EEG data while protecting the privacy of users. Anisha Agarwal, Rafael Dowsley, Nicholas D. McKinney, Dongrui Wu, Chin-Teng Lin, Martine De Cock, Anderson C. A. Nascimento |
IEEE BigData | 6 |
| 2018 | Privacy-Preserving User Profiling with Facebook LikesabstractThe content generated by users on social media is rich in personal information that can be mined to construct accurate user profiles, and subsequently used for tailored advertising or other personalized services. Facebook has recently come under scrutiny after a third party gained access to the data of millions of users and mined it to construct psychographical profiles, which were allegedly used to influence voters in elections. As part of a possible solution to avoid data breaches while still being able to perform meaningful machine learning (ML) on social media data, we propose a privacy-preserving algorithm for k-nearest neighbor (kNN) [1] , one of the oldest ML methods, used traditionally in collaborative filtering recommender systems. Sanchya Bhagat, Keerthanaa Saminathan, Anisha Agarwal, Rafael Dowsley, Martine De Cock, Anderson C. A. Nascimento |
IEEE BigData | 5 |
| 2018 | Privacy-Preserving Scoring of Tree Ensembles: A Novel Framework for AI in HealthcareabstractMachine Learning (ML) techniques now impact a wide variety of domains. Highly regulated industries such as healthcare and finance have stringent compliance and data governance policies around data sharing. Advances in secure multiparty computation (SMC) for privacy-preserving machine learning (PPML) can help transform these regulated industries by allowing ML computations over encrypted data with personally identifiable information (PII). Yet very little of SMC-based PPML has been put into practice so far. In this paper we present the very first framework for privacy-preserving classification of tree ensembles with application in healthcare. We first describe the underlying cryptographic protocols that enable a healthcare organization to send encrypted data securely to a ML scoring service and obtain encrypted class labels without the scoring service actually seeing that input in the clear. We then describe the deployment challenges we solved to integrate these protocols in a cloud based scalable risk-prediction platform with multiple ML models for healthcare AI. Included are system internals, and evaluations of our deployment for supporting physicians to drive better clinical outcomes in an accurate, scalable, and provably secure manner. To the best of our knowledge, this is the first such applied framework with SMC-based privacy-preserving machine learning for healthcare. Kyle Fritchman, Keerthanaa Saminathan, Rafael Dowsley, Tyler Hughes, Martine De Cock, Anderson C. A. Nascimento, Ankur Teredesai |
IEEE BigData | 5 |
| 2018 | An Evaluation of DGA ClassifiersabstractDomain Generation Algorithms (DGAs) are a popular technique used by contemporary malware for command-and-control (C&C) purposes. Such malware utilizes DGAs to create a set of domain names that, when resolved, provide information necessary to establish a link to a C&C server. Automated discovery of such domain names in real-time DNS traffic is critical for network security as it allows to detect infection, and, in some cases, take countermeasures to disrupt the communication and identify infected machines. Detection of the specific DGA malware family provides the administrator valuable information about the kind of infection and steps that need to be taken. In this paper we compare and evaluate machine learning methods that classify domain names as benign or DGA, and label the latter according to their malware family. Unlike previous work, we select data for test and training sets according to observation time and known seeds. This allows us to assess the robustness of the trained classifiers for detecting domains generated by the same families at a different time or when seeds change. Our study includes tree ensemble models based on human-engineered features and deep neural networks that learn features automatically from domain names. We find that all state-of-the-art classifiers are significantly better at catching domain names from malware families with a time-dependent seed compared to time-invariant DGAs. In addition, when applying the trained classifiers on a day of real traffic, we find that many domain names unjustifiably are flagged as malicious, thereby revealing the shortcomings of relying on a standard whitelist for training a production grade DGA detection system. Raaghavi Sivaguru, Chhaya Choudhary, Vadym Tymchenko, Anderson C. A. Nascimento, Martine De Cock |
IEEE BigData | 6 |
| 2018 | Character Level based Detection of DGA Domain NamesabstractRecently several different deep learning architectures have been proposed that take a string of characters as the raw input signal and automatically derive features for text classification. Few studies are available that compare the effectiveness of these approaches for character based text classification with each other. In this paper we perform such an empirical comparison for the important cybersecurity problem of DGA detection: classifying domain names as either benign vs. produced by malware (i.e., by a Domain Generation Algorithm). Training and evaluating on a dataset with 2M domain names shows that there is surprisingly little difference between various convolutional neural network (CNN) and recurrent neural network (RNN) based architectures in terms of accuracy, prompting a preference for the simpler architectures, since they are faster to train and to score, and less prone to overfitting. Anderson C. A. Nascimento, Martine De Cock |
IJCNN | 5 |
| 2018 | Dictionary Extraction and Detection of Algorithmically Generated Domain Names in Passive DNS Traffic
Mayana Pereira, Shaun Coleman, Martine De Cock, Anderson C. A. Nascimento |
RAID | 4 |
| 2018 | User Profiling through Deep Multimodal FusionabstractUser profiling in social media has gained a lot of attention due to its varied set of applications in advertising, marketing, recruiting, and law enforcement. Among the various techniques for user modeling, there is fairly limited work on how to merge multiple sources or modalities of user data - such as text, images, and relations - to arrive at more accurate user profiles. In this paper, we propose a deep learning approach that extracts and fuses information across different modalities. Our hybrid user profiling framework utilizes a shared representation between modalities to integrate three sources of data at the feature level, and combines the decision of separate networks that operate on each combination of data sources at the decision level. Our experimental results on more than 5K Facebook users demonstrate that our approach outperforms competing approaches for inferring age, gender and personality traits of social media users. We get highly accurate results with AUC values of more than 0.9 for the task of age prediction and 0.95 for the task of gender prediction. Golnoosh Farnadi, Jie Tang 0001, Martine De Cock, Marie-Francine Moens |
WSDM | 3 |
| 2018 | Modeling multi-valued biological interaction networks using fuzzy answer set programmingabstractFuzzy Answer Set Programming (FASP) is an extension of the popular Answer Set Programming (ASP) paradigm that allows for modeling and solving combinatorial search problems in continuous domains. The recent development of practical solvers for FASP has enabled its applicability to real-world problems. In this paper, we investigate the application of FASP in modeling the dynamics of Gene Regulatory Networks (GRNs). A commonly used simplifying assumption to model the dynamics of GRNs is to assume only Boolean levels of activation of each node. Our work extends this Boolean network formalism by allowing multi-valued activation levels . We show how FASP can be used to model the dynamics of such networks. We experimentally assess the efficiency of our method using real biological networks found in the literature, as well as on randomly-generated synthetic networks. The experiments demonstrate the applicability and usefulness of our proposed method to find network attractors. Mushthofa, Steven Schockaert, Ling-Hong Hung, Kathleen Marchal, Martine De Cock |
Fuzzy Sets Syst. | 5 |
| 2018 | Modelling incomplete information in Boolean games using possibilistic logic
Sofie De Clercq, Steven Schockaert, Ann Nowé, Martine De Cock |
Int. J. Approx. Reason. | 4 |
| 2017 | Exact and heuristic methods for solving Boolean games
Sofie De Clercq, Kim Bauters, Steven Schockaert, Mihail Mihaylov, Ann Nowé, Martine De Cock |
Auton. Agents Multi Agent Syst. | 6 |
| 2017 | Repairing inconsistent answer set programs using rules of thumb: A gene regulatory networks case study
Elie Merhej, Steven Schockaert, Martine De Cock |
Int. J. Approx. Reason. | 3 |
| 2017 | Soft quantification in statistical relational learning
Golnoosh Farnadi, Stephen H. Bach, Marie-Francine Moens, Lise Getoor, Martine De Cock |
Mach. Learn. | 5 |
| 2016 | VirtualIdentity: Privacy preserving user profilingabstractUser profiling from user generated content (UGC) is a common practice that supports the business models of many social media companies. Existing systems require that the UGC is fully exposed to the module that constructs the user profiles. In this paper we show that it is possible to build user profiles without ever accessing the user's original data, and without exposing the trained machine learning models for user profiling - which are the intellectual property of the company - to the users of the social media site. We present VirtualIdentity, an application that uses secure multi-party cryptographic protocols to detect the age, gender and personality traits of users by classifying their user-generated text and personal pictures with trained support vector machine models in a privacy preserving manner. Wing-Sea Poon, Golnoosh Farnadi, Caleb Horst, Kebra Thompson, Michael Nickels, Anderson C. A. Nascimento, Martine De Cock |
ASONAM | 8 |
| 2016 | Formalizing Commitment-Based Deals in Boolean GamesabstractBoolean games (BGs) are a strategic framework in which agents' goals are described using propositional logic. Despite the popularity of BGs, the problem of how agents can coordinate with others to (at least partially) achieve their goals has hardly received any attention. However, negotiation protocols that have been developed outside the setting of BGs can be adopted for this purpose, provided that we can formalize (i) how agents can make commitments and (ii) how deals between coalitions of agents can be identified given a set of active commitments. In this paper, we focus on these two aims. First, we show how agents can formulate commitments that are in accordance with their goals, and what it means for the commitments of an agent to be consistent. Second, we formalize deals in terms of coalitions who can achieve their goals without help from others. We show that verifying the consistency of a set of commitments of one agent is ΠP2-complete while checking the existence of a deal in a set of mutual commitments is Σp2 Sofie De Clercq, Steven Schockaert, Ann Nowé, Martine De Cock |
ECAI | 4 |
| 2016 | Computing attractors of multi-valued Gene Regulatory Networks using Fuzzy Answer Set ProgrammingabstractFuzzy Answer Set Programming (FASP) extends the popular Answer Set Programming (ASP) paradigm to modeling and solving combinatorial search problems in continuous domains. The recent development of FASP solvers has turned FASP into a practical tool for solving real-world problems. In this paper, we propose the use of FASP for modeling the dynamics of Gene Regulatory Networks (GRNs), an important kind of biological network. A commonly used simplifying assumption to model the dynamics of GRNs is to assume only Boolean levels of activation of each node. ASP has been used to model such Boolean networks. Our work extends this Boolean network formalism by allowing multi-valued activation levels. We show how FASP can be used to model the dynamics of such networks. We also experimentally assess the plausibility of our method using real biological networks found in the literature. Mushthofa, Steven Schockaert, Martine De Cock |
FUZZ-IEEE | 3 |
| 2016 | Solving stable matching problems using answer set programmingabstractAbstract Since the introduction of the stable marriage problem (SMP) by Gale and Shapley (1962), several variants and extensions have been investigated. While this variety is useful to widen the application potential, each variant requires a new algorithm for finding the stable matchings. To address this issue, we propose an encoding of the SMP using answer set programming (ASP), which can straightforwardly be adapted and extended to suit the needs of specific applications. The use of ASP also means that we can take advantage of highly efficient off-the-shelf solvers. To illustrate the flexibility of our approach, we show how our ASP encoding naturally allows us to select optimal stable matchings, i.e. matchings that are optimal according to some user-specified criterion. To the best of our knowledge, our encoding offers the first exact implementation to find sex-equal, minimum regret, egalitarian or maximum cardinality stable matchings for SMP instances in which individuals may designate unacceptable partners and ties between preferences are allowed. Sofie De Clercq, Steven Schockaert, Martine De Cock, Ann Nowé |
Theory Pract. Log. Program. | 3 |
| 2016 | Computational personality recognition in social mediaabstractA variety of approaches have been recently proposed to automatically infer users’ personality from their user generated content in social media. Approaches differ in terms of the machine learning algorithms and the feature sets used, type of utilized footprint, and the social media environment used to collect the data. In this paper, we perform a comparative analysis of state-of-the-art computational personality recognition methods on a varied set of social media ground truth data from Facebook, Twitter and YouTube. We answer three questions: (1) Should personality prediction be treated as a multi-label prediction task (i.e., all personality traits of a given user are predicted at once), or should each trait be identified separately? (2) Which predictive features work well across different on-line environments? and (3) What is the decay in accuracy when porting models trained in one social media environment to another? Golnoosh Farnadi, Geetha Sitaraman, Shanu Sushmita, Fabio Celli, Michal Kosinski, David Stillwell, Sergio Davalos, Marie-Francine Moens, Martine De Cock |
User Model. User Adapt. Interact. | 9 |
| 2015 | Scalable adaptive label propagation in GrappaabstractNodes of a social graph often represent entities with specific labels, denoting properties such as age-group or gender. Design of algorithms to assign labels to unlabeled nodes by leveraging node-proximity and a-priori labels of seed nodes is of significant interest. A semi-supervised approach to solve this problem is termed "LPA-Label Propagation Algorithm" where labels of a subset of nodes are iteratively propagated through the network to infer yet unknown node labels. While LPA for node labelling is extremely fast and simple, it works well only with an assumption of node-homophily — connected nodes are connected because they must deserve a similar label — which can often be a misnomer. In this paper we propose a novel algorithm "Adaptive Label Propagation" that dynamically adapts to the underlying characteristics of homophily, heterophily, or otherwise, of the connections of the network, and applies suitable label propagation strategies accordingly. Moreover, our adaptive label propagation approach is scalable as demonstrated by its implementation in Grappa, a distributed shared-memory system. Our experiments on social graphs from Facebook, YouTube, Live Journal, Orkut and Netlog demonstrate that our approach not only improves the labelling accuracy but also computes results for millions of users within a few seconds. Golnoosh Farnadi, Zeinab Mahdavifar, Ivan Keller, Jacob Nelson 0001, Ankur Teredesai, Marie-Francine Moens, Martine De Cock |
IEEE BigData | 7 |
| 2015 | Fuzzy Rough Set Prototype Selection for RegressionabstractInstance selection methods are a class of preprocessing techniques that have been widely studied in machine learning to remove redundant or noisy instances from a training set. The main focus of such prior efforts has been on the selection of suitable training instances to perform a classification task for crisp class labels. In this paper, we propose a novel instance selection technique termed Fuzzy Rough Set Prototype Selection for Regression (FRPS-R) for solving regression problems, where the outcome is continuous. We use concepts from fuzzy rough set theory and extend the currently well-known fuzzy rough set prototype selection technique to model the quality of all available elements and then use a wrapper approach to select an optimal subset of high-quality instances; thereby generalizing the idea. Our experimental evaluation shows that the application of our proposed instance selection technique can significantly improve the predictive performance of the weighted k-nearest neighbor regression algorithm, in particular when noise is present in the original training set. Sarah Vluymans, Yvan Saeys, Chris Cornelis, Ankur Teredesai, Martine De Cock |
FUZZ-IEEE | 5 |
| 2015 | Multilateral Negotiation in Boolean Games with Incomplete Information Using Generalized Possibilistic Logic
Sofie De Clercq, Steven Schockaert, Ann Nowé, Martine De Cock |
IJCAI | 4 |
| 2015 | Statistical Relational Learning with Soft Quantifiers
Golnoosh Farnadi, Stephen H. Bach, Marjon Blondeel, Marie-Francine Moens, Lise Getoor, Martine De Cock |
ILP | 6 |
| 2015 | Solving Disjunctive Fuzzy Answer Set Programs
Mushthofa, Steven Schockaert, Martine De Cock |
LPNMR | 3 |
| 2015 | On the relationship between fuzzy autoepistemic logic and fuzzy modal logics of belief
Marjon Blondeel, Tommaso Flaminio, Steven Schockaert, Lluís Godo, Martine De Cock |
Fuzzy Sets Syst. | 5 |
| 2015 | Characterizing and extending answer set semantics using possibility theoryabstractAbstract Answer Set Programming (ASP) is a popular framework for modelling combinatorial problems. However, ASP cannot be used easily for reasoning about uncertain information. Possibilistic ASP (PASP) is an extension of ASP that combines possibilistic logic and ASP. In PASP a weight is associated with each rule, whereas this weight is interpreted as the certainty with which the conclusion can be established when the body is known to hold. As such, it allows us to model and reason about uncertain information in an intuitive way. In this paper we present new semantics for PASP in which rules are interpreted as constraints on possibility distributions. Special models of these constraints are then identified as possibilistic answer sets. In addition, since ASP is a special case of PASP in which all the rules are entirely certain, we obtain a new characterization of ASP in terms of constraints on possibility distributions. This allows us to uncover a new form of disjunction, called weak disjunction, that has not been previously considered in the literature. In addition to introducing and motivating the semantics of weak disjunction, we also pinpoint its computational complexity. In particular, while the complexity of most reasoning tasks coincides with standard disjunctive ASP, we find that brave reasoning for programs with weak disjunctions is easier. Kim Bauters, Steven Schockaert, Martine De Cock, Dirk Vermeir |
Theory Pract. Log. Program. | 3 |
| 2014 | Computing fuzzy rough approximations in large scale information systemsabstractRough set theory is a popular and powerful machine learning tool. It is especially suitable for dealing with information systems that exhibit inconsistencies, i.e. objects that have the same values for the conditional attributes but a different value for the decision attribute. In line with the emerging granular computing paradigm, rough set theory groups objects together based on the indiscernibility of their attribute values. Fuzzy rough set theory extends rough set theory to data with continuous attributes, and detects degrees of inconsistency in the data. Key to this is turning the indiscernibility relation into a gradual relation, acknowledging that objects can be similar to a certain extent. In very large datasets with millions of objects, computing the gradual indiscernibility relation (or in other words, the soft granules) is very demanding, both in terms of runtime and in terms of memory. It is however required for the computation of the lower and upper approximations of concepts in the fuzzy rough set analysis pipeline. Current non-distributed implementations in R are limited by memory capacity. For example, we found that a state of the art non-distributed implementation in R could not handle 30,000 rows and 10 attributes on a node with 62GB of memory. This is clearly insufficient to scale fuzzy rough set analysis to massive datasets. In this paper we present a parallel and distributed solution based on Message Passing Interface (MPI) to compute fuzzy rough approximations in very large information systems. Our results show that our parallel approach scales with problem size to information systems with millions of objects. To the best of our knowledge, no other parallel and distributed solutions have been proposed so far in the literature for this problem. Hasan Asfoor, Rajagopalan Srinivasan, Gayathri Vasudevan, Nele Verbiest, Chris Cornelis, Matthew E. Tolentino, Ankur Teredesai, Martine De Cock |
IEEE BigData | 8 |
| 2014 | A finite-valued solver for disjunctive fuzzy answer set programsabstractFuzzy Answer Set Programming (FASP) is a declarative programming paradigm which extends the flexibility and expressiveness of classical Answer Set Programming (ASP), with the aim of modeling continuous application domains. In contrast to the availability of efficient ASP solvers, there have been few attempts at implementing FASP solvers. In this paper, we propose an implementation of FASP based on a reduction to classical ASP. We also develop a prototype implementation of this method. To the best of our knowledge, this is the first solver for disjunctive FASP programs. Moreover, we experimentally show that our solver performs well in comparison to an existing solver (under reasonable assumptions) for the more restrictive class of normal FASP programs. Mushthofa, Steven Schockaert, Martine De Cock |
ECAI | 3 |
| 2014 | Decentralized Computation of Pareto Optimal Pure Nash Equilibria of Boolean Games with Privacy Concerns abstractIn Boolean games, agents try to reach a goal formulated as a Boolean formula. These games are attractive because of their compact representations. However, few methods are available to compute the solutions and they are either limited or do not take privacy or communication concerns into account. In this paper we propose the use of an algorithm related to reinforcement learning to address this problem. Our method is decentralized in the sense that agents try to achieve their goals without knowledge of the other agents’ goals. We prove that this is a sound method to compute a Pareto optimal pure Nash equilibrium for an interesting class of Boolean games. Experimental results are used to investigate the performance of the algorithm. Sofie De Clercq, Kim Bauters, Steven Schockaert, Mihail Mihaylov, Martine De Cock, Ann Nowé |
ICAART (2) | 5 |
| 2014 | Possibilistic Boolean Games: Strategic Reasoning under Incomplete Information
Sofie De Clercq, Steven Schockaert, Martine De Cock, Ann Nowé |
JELIA | 3 |
| 2014 | Using Answer Set Programming for Solving Boolean Games
Sofie De Clercq, Kim Bauters, Steven Schockaert, Martine De Cock, Ann Nowé |
KR | 4 |
| 2014 | ASP-G: an ASP-based method for finding attractors in genetic regulatory networksabstractMOTIVATION: Boolean network models are suitable to simulate GRNs in the absence of detailed kinetic information. However, reducing the biological reality implies making assumptions on how genes interact (interaction rules) and how their state is updated during the simulation (update scheme). The exact choice of the assumptions largely determines the outcome of the simulations. In most cases, however, the biologically correct assumptions are unknown. An ideal simulation thus implies testing different rules and schemes to determine those that best capture an observed biological phenomenon. This is not trivial because most current methods to simulate Boolean network models of GRNs and to compute their attractors impose specific assumptions that cannot be easily altered, as they are built into the system. RESULTS: To allow for a more flexible simulation framework, we developed ASP-G. We show the correctness of ASP-G in simulating Boolean network models and obtaining attractors under different assumptions by successfully recapitulating the detection of attractors of previously published studies. We also provide an example of how performing simulation of network models under different settings help determine the assumptions under which a certain conclusion holds. The main added value of ASP-G is in its modularity and declarativity, making it more flexible and less error-prone than traditional approaches. The declarative nature of ASP-G comes at the expense of being slower than the more dedicated systems but still achieves a good efficiency with respect to computational time. AVAILABILITY AND IMPLEMENTATION: The source code of ASP-G is available at http://bioinformatics.intec.ugent.be/kmarchal/Supplementary_Information_Musthofa_2014/asp-g.zip. Mushthofa, Gustavo Torres, Yves Van de Peer, Kathleen Marchal, Martine De Cock |
Bioinform. | 5 |
| 2014 | Fuzzy autoepistemic logic and its relation to fuzzy answer set programming
Marjon Blondeel, Steven Schockaert, Martine De Cock, Dirk Vermeir |
Fuzzy Sets Syst. | 3 |
| 2014 | Semantics for possibilistic answer set programs: Uncertain rules versus rules with uncertain conclusions
Kim Bauters, Steven Schockaert, Martine De Cock, Dirk Vermeir |
Int. J. Approx. Reason. | 3 |
| 2014 | Complexity of fuzzy answer set programming under Łukasiewicz semantics
Marjon Blondeel, Steven Schockaert, Dirk Vermeir, Martine De Cock |
Int. J. Approx. Reason. | 4 |
| 2014 | Using the crowd for readability predictionabstractAbstract While human annotation is crucial for many natural language processing tasks, it is often very expensive and time-consuming. Inspired by previous work on crowdsourcing, we investigate the viability of using non-expert labels instead of gold standard annotations from experts for a machine learning approach to automatic readability prediction. In order to do so, we evaluate two different methodologies to assess the readability of a wide variety of text material: A more traditional setup in which expert readers make readability judgments and a crowdsourcing setup for users who are not necessarily experts. To this purpose two assessment tools were implemented: a tool where expert readers can rank a batch of texts based on readability, and a lightweight crowdsourcing tool, which invites users to provide pairwise comparisons. To validate this approach, readability assessments for a corpus of written Dutch generic texts were gathered. By collecting multiple assessments per text, we explicitly wanted to level out readers' background knowledge and attitude. Our findings show that the assessments collected through both methodologies are highly consistent and that crowdsourcing is a viable alternative to expert labeling. This is a good news as crowdsourcing is more lightweight to use and can have access to a much wider audience of potential annotators. By performing a set of basic machine learning experiments using a feature set that mainly encodes basic lexical and morpho-syntactic information, we further illustrate how the collected data can be used to perform text comparisons or to assign an absolute readability score to an individual text. We do not focus on optimising the algorithms to achieve the best possible results for the learning tasks, but carry them out to illustrate the various possibilities of our data sets. The results on different data sets, however, show that our system outperforms the readability formulas and a baseline language modelling approach. We conclude that readability assessment by comparing texts is a polyvalent methodology, which can be adapted to specific domains and target audiences if required. Orphée De Clercq, Véronique Hoste, Bart Desmet, Philip van Oosten, Martine De Cock, Lieve Macken |
Nat. Lang. Eng. | 5 |
| 2013 | The Microsoft Academic Search challenges at KDD Cup 2013abstractMicrosoft Academic Search is a free search engine specific to scholarly material. It currently covers more than 50 million publications and over 19 million authors across a variety of domains. One of the main challenges in correctly indexing this material is author name ambiguity and the resulting noise in author profiles. KDD Cup 2013 invited participants to tackle this problem in 2 ways: (1) by automatically determining which papers in an author profile are truly written by a given author, and (2) by identifying which author profiles need to be merged because they belong to the same author. This paper presents a brief account of the contest and the lessons learned. Martine De Cock, Senjuti Basu Roy, Swapna Savvana, Vani Mandava, Brian Dalessandro, Claudia Perlich, Will Cukierski, Benjamin Hamner |
IEEE BigData | 1 |
| 2013 | Five Languages Are Better Than One: An Attempt to Bypass the Data Acquisition Bottleneck for WSD
Els Lefever, Véronique Hoste, Martine De Cock |
CICLing (1) | 3 |
| 2013 | Solving satisfiability in fuzzy logics by mixing CMA-ESabstractSatisfiability in propositional logic is well researched and many approaches to checking and solving exist. In infinite-valued or fuzzy logics, however, there have only recently been attempts at developing methods for solving satisfiability. In this paper, we propose new benchmark problems and analyse the function landscape of different problem classes, focussing our analysis on plateaus. Based on this study, we develop Mixing CMA-ES (M-CMA-ES), an extension to CMA-ES that is well suited to solving problems with many large plateaus. We empirically show the relation between certain function landscape properties and M-CMA-ES performance. Tim Brys, Madalina M. Drugan, Peter A. N. Bosman, Martine De Cock, Ann Nowé |
GECCO | 4 |
| 2013 | Towards a Deeper Understanding of Nonmonotonic Reasoning with Degrees
Marjon Blondeel, Steven Schockaert, Dirk Vermeir, Martine De Cock |
IJCAI | 4 |
| 2013 | Expressiveness of communication in answer set programmingabstractAbstract Answer set programming (ASP) is a form of declarative programming that allows to succinctly formulate and efficiently solve complex problems. An intuitive extension of this formalism is communicating ASP, in which multiple ASP programs collaborate to solve the problem at hand. However, the expressiveness of communicating ASP has not been thoroughly studied. In this paper, we present a systematic study of the additional expressiveness offered by allowing ASP programs to communicate. First, we consider a simple form of communication where programs are only allowed to ask questions to each other. For the most part, we deliberately consider only simple programs, i.e. programs for which computing the answer sets is in P. We find that the problem of deciding whether a literal is in some answer set of a communicating ASP program using simple communication is NP-hard. In other words, due to the ability of these simple ASP programs to communicate and collaborate, we move up a step in the polynomial hierarchy. Second, we modify the communication mechanism to also allow us to focus on a sequence of communicating programs, where each program in the sequence may successively remove some of the remaining models. This mimics a network of leaders, where the first leader has the first say and may remove models that he or she finds unsatisfactory. Using this particular communication mechanism allows us to capture the entire polynomial hierarchy. This means, in particular, that communicating ASP could be used to solve problems that are above the second level of polynomial hierarchy, such as some forms of abductive reasoning as well as PSPACE-complete problems such as STRIPS planning. Kim Bauters, Steven Schockaert, Jeroen Janssen, Dirk Vermeir, Martine De Cock |
Theory Pract. Log. Program. | 5 |
| 2013 | Enhancing the trust-based recommendation process with explicit distrustabstractWhen a Web application with a built-in recommender offers a social networking component which enables its users to form a trust network, it can generate more personalized recommendations by combining user ratings with information from the trust network. These are the so-called trust-enhanced recommendation systems. While research on the incorporation of trust for recommendations is thriving, the potential of explicitly stated distrust remains almost unexplored. In this article, we introduce a distrust-enhanced recommendation algorithm which has its roots in Golbeck's trust-based weighted mean. Through experiments on a set of reviews from Epinions.com, we show that our new algorithm outperforms its standard trust-only counterpart with respect to accuracy, thereby demonstrating the positive effect that explicit distrust can have on trust-based recommendations. Patricia Victor, Nele Verbiest, Chris Cornelis, Martine De Cock |
ACM Trans. Web | 4 |
| 2012 | Possible and Necessary Answer Sets of Possibilistic Answer Set ProgramsabstractAnswer set programming (ASP) and possibility theory can be combined to form possibilistic answer set programming (PASP), a framework for non-monotonic reasoning under uncertainty. Existing proposals view answer sets of PASP programs as weighted epistemic states, in which the strength by which different literals are believed to hold may vary. In contrast, in this paper we propose an approach in which epistemic states remain Boolean, but some epistemic states may be considered more plausible than others. A PASP program is then a representation of an incomplete description of these epistemic states where certainties are associated with each rule which is interpreted in terms of a necessity measure. The main contribution of this paper is the introduction of a new semantics for PASP as well as a study of the resulting complexity. Kim Bauters, Steven Schockaert, Martine De Cock, Dirk Vermeir |
ICTAI | 3 |
| 2012 | Discovering Missing Wikipedia Inter-language Links by means of Cross-lingual Word Sense Disambiguation
Els Lefever, Véronique Hoste, Martine De Cock |
LREC | 3 |
| 2012 | A core language for fuzzy answer set programming
Jeroen Janssen, Steven Schockaert, Dirk Vermeir, Martine De Cock |
Int. J. Approx. Reason. | 4 |
| 2012 | Reducing fuzzy answer set programming to model finding in fuzzy logicsabstractAbstract In recent years, answer set programming (ASP) has been extended to deal with multivalued predicates. The resulting formalismsallow for the modeling of continuous problems as elegantly as ASP allows for the modeling of discrete problems, by combining thestable model semantics underlying ASP with fuzzy logics. However, contrary to the case of classical ASP where manyefficient solvers have been constructed, to date there is no efficient fuzzy ASP solver. A well-knowntechnique for classical ASP consists of translating an ASP programPto a propositional theory whose models exactlycorrespond to the answer sets ofP. In this paper, we show how this idea can be extended to fuzzy ASP, paving the wayto implement efficient fuzzy ASP solvers that can take advantage of existing fuzzy logic reasoners. Jeroen Janssen, Dirk Vermeir, Steven Schockaert, Martine De Cock |
Theory Pract. Log. Program. | 4 |
| 2011 | Fuzzy Autoepistemic Logic: Reflecting about Knowledge of Truth Degrees
Marjon Blondeel, Steven Schockaert, Martine De Cock, Dirk Vermeir |
ECSQARU | 3 |
| 2011 | Communicating ASP and the Polynomial Hierarchy
Kim Bauters, Steven Schockaert, Dirk Vermeir, Martine De Cock |
LPNMR | 4 |
| 2011 | Practical aggregation operators for gradual trust and distrust
Patricia Victor, Chris Cornelis, Martine De Cock, Enrique Herrera-Viedma |
Fuzzy Sets Syst. | 3 |
| 2010 | Born to trade: A genetically evolved keyword bidder for sponsored searchabstractIn sponsored search auctions, advertisers choose a set of keywords based on products they wish to market. They bid for advertising slots that will be displayed on the search results page when a user submits a query containing the keywords that the advertiser selected. Deciding how much to bid is a real challenge: if the bid is too low with respect to the bids of other advertisers, the ad might not get displayed in a favorable position; a bid that is too high on the other hand might not be profitable either, since the attracted number of conversions might not be enough to compensate for the high cost per click. In this paper we propose a genetically evolved keyword bidding strategy that decides how much to bid for each query based on historical data such as the position obtained on the previous day. In light of the fact that our approach does not implement any particular expert knowledge on keyword auctions, it did remarkably well in the Trading Agent Competition at IJCAI2009. Michael Munsey, Jonathan Veilleux, Sindhura Bikkani, Ankur Teredesai, Martine De Cock |
IEEE Congress on Evolutionary Computation | 5 |
| 2010 | Possibilistic Answer Set Programming Revisited
Kim Bauters, Steven Schockaert, Martine De Cock, Dirk Vermeir |
UAI | 3 |
| 2010 | Ranking Approaches for Microblog SearchabstractRanking microblogs, such as tweets, as search results for a query is challenging, among other things because of the sheer amount of microblogs that are being generated in real time, as well as the short length of each individual microblog. In this paper, we describe several new strategies for ranking microblogs in a real-time search engine. Evaluating these ranking strategies is non-trivial due to the lack of a publicly available ground truth validation dataset. We have therefore developed a framework to obtain such validation data, as well as evaluation measures to assess the accuracy of the proposed ranking strategies. Our experiments demonstrate that it is beneficial for microblog search engines to take into account social network properties of the authors of microblogs in addition to properties of the microblog itself. Rinkesh Nagmoti, Ankur Teredesai, Martine De Cock |
Web Intelligence | 3 |
| 2010 | Clustering web people search results using fuzzy ants
Els Lefever, Timur Fayruzov, Véronique Hoste, Martine De Cock |
Inf. Sci. | 4 |
| 2010 | Reasoning about fuzzy temporal information from the web: towards retrieval of historical events
Steven Schockaert, Martine De Cock, Etienne E. Kerre |
Soft Comput. | 2 |
| 2009 | Modeling Protein Interaction Networks with Answer Set ProgrammingabstractIn this paper we propose the use of answer set programming (ASP) to model protein interaction networks. We argue that this declarative formalism rivals the popular boolean networks in terms of ease of use, while at the same time being more expressive. As we demonstrate for the particularcase of a fission yeast network, all information present in a boolean network, as well as relevant background assumptions,can be expressed explicitly in an answer set program. Moreover, readily available answer set solvers can then be used to find the stable states of the network. Timur Fayruzov, Martine De Cock, Chris Cornelis, Dirk Vermeir |
BIBM | 2 |
| 2009 | A Comparative Analysis of Trust-Enhanced Recommenders for Controversial Items
Patricia Victor, Chris Cornelis, Martine De Cock, Ankur Teredesai |
ICWSM | 3 |
| 2009 | Spatial reasoning in a fuzzy region connection calculus
Steven Schockaert, Martine De Cock, Etienne E. Kerre |
Artif. Intell. | 2 |
| 2009 | Linguistic feature analysis for protein interaction extractionabstractBACKGROUND: The rapid growth of the amount of publicly available reports on biomedical experimental results has recently caused a boost of text mining approaches for protein interaction extraction. Most approaches rely implicitly or explicitly on linguistic, i.e., lexical and syntactic, data extracted from text. However, only few attempts have been made to evaluate the contribution of the different feature types. In this work, we contribute to this evaluation by studying the relative importance of deep syntactic features, i.e., grammatical relations, shallow syntactic features (part-of-speech information) and lexical features. For this purpose, we use a recently proposed approach that uses support vector machines with structured kernels. RESULTS: Our results reveal that the contribution of the different feature types varies for the different data sets on which the experiments were conducted. The smaller the training corpus compared to the test data, the more important the role of grammatical relations becomes. Moreover, deep syntactic information based classifiers prove to be more robust on heterogeneous texts where no or only limited common vocabulary is shared. CONCLUSION: Our findings suggest that grammatical relations play an important role in the interaction extraction task. Moreover, the net advantage of adding lexical and shallow syntactic features is small related to the number of added features. This implies that efficient classifiers can be built by using only a small fraction of the features that are typically being used in recent approaches. Timur Fayruzov, Martine De Cock, Chris Cornelis, Véronique Hoste |
BMC Bioinform. | 2 |
| 2009 | Gradual trust and distrust in recommender systems
Patricia Victor, Chris Cornelis, Martine De Cock, Paulo Pinheiro 0001 |
Fuzzy Sets Syst. | 3 |
| 2009 | Efficient Algorithms for Fuzzy Qualitative Temporal ReasoningabstractFuzzy qualitative temporal relations have been proposed to reason about events whose temporal boundaries are ill defined. Although the corresponding reasoning tasks are in the same complexity class as their crisp counterparts, in practice, the scalability of fuzzy temporal reasoners may be insufficient for applications that require a high expressivity and deal with a large number of events. On the other hand, transitivity rules can be used to make sound but incomplete inferences in polynomial time, utilizing a variant of Allen's path-consistency algorithm. The aim of this paper is to investigate how this polynomial time algorithm can be improved without altering its time complexity. To this end, we establish a characterization of 2-consistency of fuzzy temporal relations and provide transitivity rules that are significantly stronger than those resulting from straightforwardly generalizing transitivity rules for crisp temporal relations. We furthermore provide experimental evidence for the effectiveness of our improved algorithm. Steven Schockaert, Martine De Cock |
IEEE Trans. Fuzzy Syst. | 2 |
| 2008 | Modelling nearness and cardinal directions between fuzzy regionsabstractA significant part of real-world spatial information is affected by vagueness. For example, boundaries of non-administrative geographical regions tend to be ill-defined, while information about the nearness and relative orientation of two places is typically expressed through vague linguistic descriptions. In this paper, we propose a general framework to represent such information, using the concept of relatedness measures for fuzzy sets. Regions are represented as fuzzy sets in a two-dimensional Euclidean space, and nearness and relative orientation are expressed as fuzzy relations. To support fuzzy spatial reasoning, we derive transitivity rules and provide efficient techniques to deal with the complex interactions between nearness and cardinal directions. Steven Schockaert, Martine De Cock, Etienne E. Kerre |
FUZZ-IEEE | 2 |
| 2008 | Compiling Fuzzy Answer Set Programs to Fuzzy Propositional Theories
Jeroen Janssen, Stijn Heymans, Dirk Vermeir, Martine De Cock |
ICLP | 4 |
| 2008 | Temporal reasoning about fuzzy intervals
Steven Schockaert, Martine De Cock |
Artif. Intell. | 2 |
| 2008 | Location approximation for local search services using natural language hintsabstractLocal search services allow a user to search for businesses that satisfy a given geographical constraint. In contrast to traditional web search engines, current local search services rely heavily on static, structured data. Although this yields very accurate systems, it also implies a limited coverage, and limited support for using landmarks and neighborhood names in queries. To overcome these limitations, we propose to augment the structured information available to a local search service, based on the vast amount of unstructured and semi-structured data available on the web. This requires a computational framework to represent vague natural language information about the nearness of places, as well as the spatial extent of vague neighborhoods. In this paper, we propose such a framework based on fuzzy set theory, and show how natural language information can be translated into this framework. We provide experimental results that show the effectiveness of the proposed techniques, and demonstrate that local search based on natural language hints about the location of places with an unknown address, is feasible. Steven Schockaert, Martine De Cock, Etienne E. Kerre |
Int. J. Geogr. Inf. Sci. | 2 |
| 2008 | Fuzzy region connection calculus: Representing vague topological information
Steven Schockaert, Martine De Cock, Chris Cornelis, Etienne E. Kerre |
Int. J. Approx. Reason. | 2 |
| 2008 | Fuzzy region connection calculus: An interpretation based on closeness
Steven Schockaert, Martine De Cock, Chris Cornelis, Etienne E. Kerre |
Int. J. Approx. Reason. | 2 |
| 2008 | Fuzzifying Allen's Temporal Interval RelationsabstractWhen the time span of an event is imprecise, it can be represented by a fuzzy set, called a fuzzy time interval. In this paper, we propose a framework to represent, compute, and reason about temporal relationships between such events. Since our model is based on fuzzy orderings of time points, it is not only suitable to express precise relationships between imprecise events (ldquoRoosevelt died before the beginning of the Cold Warrdquo) but also imprecise relationships (ldquoRoosevelt died just before the beginning of the Cold Warrdquo). We show that, unlike previous models, our model is a generalization that preserves many of the properties of the 13 relations Allen introduced for crisp time intervals. Furthermore, we show how our model can be used for efficient fuzzy temporal reasoning by means of a transitivity table. Finally, we illustrate its use in the context of question answering systems. Steven Schockaert, Martine De Cock, Etienne E. Kerre |
IEEE Trans. Fuzzy Syst. | 2 |
| 2007 | Reasoning about vague topological informationabstractTopological information plays a fundamental role in the human perception of spatial configurations and is thereby one of the most prominent geographical features in natural language. As vagueness abounds in geography, flexible formalisms with the ability to capture vague topological information are often needed in practice. While such formalisms have already been introduced by various authors, complete reasoning procedures are usually not discussed. In this paper, we show how many interesting reasoning tasks, such as consistency checking and entailment checking, can be supported in a generalization of the well-known RCC-8 calculus. In particular, we present decision procedures based on linear programming, solving all reasoning tasks of interest. We furthermore show how deciding the consistency of vague topological information can be reduced to the consistency problem of the original RCC-8. Steven Schockaert, Martine De Cock |
CIKM | 2 |
| 2007 | Computing Fuzzy Answer Sets Using dlvhex
Davy Van Nieuwenborgh, Martine De Cock, Dirk Vermeir |
ICLP | 2 |
| 2007 | Qualitative Temporal Reasoning about Vague Events
Steven Schockaert, Martine De Cock, Etienne E. Kerre |
IJCAI | 2 |
| 2007 | Neighborhood restrictions in geographic IRabstractGeographic information retrieval (GIR) systems allow users to specify a geographic context, in addition to a more traditional query, enabling the system to pinpoint interesting search results whose relevancy is location-dependent. In particular local search services have become a widely used mechanism to find businesses, such as hotels, restaurants, and shops, which satisfy a geographical restriction. Unfortunately, many useful types of geographic restrictions are currently not supported in these systems, including restrictions that specify the neighborhood in which the business should be located. As the boundaries of city neighborhoods are not readily available, automated techniques to construct representations of the spatial extent of neighborhoods are required to support this kind of restrictions. In this paper, we propose such a technique, using fuzzy footprints to cope with the inherent vagueness of most neighborhood boundaries, and we provide experimental results that demonstrate the potential of our technique in a local search setting. Steven Schockaert, Martine De Cock |
SIGIR | 2 |
| 2007 | Selection ofWeb Services with Imprecise QoS ConstraintsabstractWhen several functionally equivalent web services are available to perform the same task, their Quality of Service (QoS) characteristics such as performance and reliability become important in the selection process. Consumers that specify their QoS requirements too strictly however, risk not finding any web services meeting their demands. Therefore, in this paper we allow QoS constraints to be described imprecisely as fuzzy sets. We compare the effectiveness of an intelligent web service selection algorithm that takes these imprecise QoS constraints into account with a baseline algorithm acting on precise QoS values. Martine De Cock, Sam Chung, Omar Hafeez |
Web Intelligence | 1 |
| 2007 | Clustering web search results using fuzzy antsabstractAlgorithms for clustering Web search results have to be efficient and robust. Furthermore they must be able to cluster a data set without using any kind of a priori information, such as the required number of clusters. Clustering algorithms inspired by the behavior of real ants generally meet these requirements. In this article we propose a novel approach to ant-based clustering, based on fuzzy logic. We show that it improves existing approaches and illustrates how our algorithm can be applied to the problem of Web search results clustering. © 2007 Wiley Periodicals, Inc. Int J Int Syst 22: 455–474, 2007. Steven Schockaert, Martine De Cock, Chris Cornelis, Etienne E. Kerre |
Int. J. Intell. Syst. | 2 |
| 2007 | Fuzzy Rough Sets: The Forgotten StepabstractTraditional rough set theory uses equivalence relations to compute lower and upper approximations of sets. The corresponding equivalence classes either coincide or are disjoint. This behaviour is lost when moving on to a fuzzy T-equivalence relation. However, none of the existing studies on fuzzy rough set theory tries to exploit the fact that an element can belong to some degree to several “soft similarity classes” at the same time. In this paper we show that taking this truly fuzzy characteristic into account may lead to new and interesting definitions of lower and upper approximations. We explore two of them in detail and we investigate under which conditions they differ from the commonly used definitions. Finally we show the possible practical relevance of the newly introduced approximations for query refinement. Martine De Cock, Chris Cornelis, Etienne E. Kerre |
IEEE Trans. Fuzzy Syst. | 1 |
| 2006 | Question Answering with Imperfect Temporal Information
Steven Schockaert, David Ahn, Martine De Cock, Etienne E. Kerre |
FQAS | 3 |
| 2006 | An Efficient Characterization of Fuzzy Temporal Interval RelationsabstractFuzzy temporal interval relations have been defined to support temporal knowledge representation and reasoning in the presence of vagueness. The most important impediment to use these fuzzy relations in real-world applications is the lack of a characterization that is both easy to implement and computationally efficient. In this paper, we provide such a characterization for the important class of piecewise linear fuzzy time intervals, which covers all types of fuzzy time intervals that we are likely to encounter in applications. Furthermore, we discuss a more elegant characterization for the special case of trapezoidally shaped fuzzy intervals. Steven Schockaert, Martine De Cock, Etienne E. Kerre |
FUZZ-IEEE | 2 |
| 2006 | Fuzzy Answer Set Programming
Davy Van Nieuwenborgh, Martine De Cock, Dirk Vermeir |
JELIA | 2 |
| 2006 | Fuzzy versus quantitative association rules: a fair data-driven comparisonabstractAs opposed to quantitative association rule mining, fuzzy association rule mining is said to prevent the overestimation of boundary cases, as can be shown by small examples. Rule mining, however, becomes interesting in large databases, where the problem of boundary cases is less apparent and can be further suppressed by using sensible partitioning methods. A data-driven approach is used to investigate if there is a significant difference between quantitative and fuzzy association rules in large databases. The influence of the choice of a particular triangular norm in this respect is also examined. Hannes Verlinde, Martine De Cock, Raymond T. Boute |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2005 | Elicitation of fuzzy association rules from positive and negative examples
Martine De Cock, Chris Cornelis, Etienne E. Kerre |
Fuzzy Sets Syst. | 1 |
| 2004 | Fuzzy rough sets: beyond the obviousabstractRough set theory was introduced in 1982. Soon it was combined with fuzzy set theory, giving rise to a hybrid model, involving fuzzy sets and fuzzy relations, which appears to be a natural, elegant generalization. In this paper we reveal that in the fuzzification process an important step seems to be overlooked. The most fascinating part is that this forgotten step arises from the true essence of fuzzy set theory: namely, that an element can belong to a given degree to more than one fuzzy set at the same time. Martine De Cock, Chris Cornelis, Etienne E. Kerre |
FUZZ-IEEE | 1 |
| 2004 | Efficient Approximate Reasoning with Positive and Negative Information
Chris Cornelis, Martine De Cock, Etienne E. Kerre |
KES | 2 |
| 2004 | Mining Positive and Negative Fuzzy Association Rules
Chris Cornelis, Martine De Cock, Etienne E. Kerre |
KES | 4 |
| 2004 | Fuzzy modifiers based on fuzzy relations
Martine De Cock, Etienne E. Kerre |
Inf. Sci. | 1 |
| 2003 | Intuitionistic fuzzy rough sets: at the crossroads of imperfect knowledgeabstractAbstract: Just like rough set theory, fuzzy set theory addresses the topic of dealing with imperfect knowledge. Recent investigations have shown how both theories can be combined into a more flexible, more expressive framework for modelling and processing incomplete information in information systems. At the same time, intuitionistic fuzzy sets have been proposed as an attractive extension of fuzzy sets, enriching the latter with extra features to represent uncertainty (on top of vagueness). Unfortunately, the various tentative definitions of the concept of an ‘intuitionistic fuzzy rough set’ that were raised in their wake are a far cry from the original objectives of rough set theory. We intend to fill an obvious gap by introducing a new definition of intuitionistic fuzzy rough sets, as the most natural generalization of Pawlak's original concept of rough sets. Chris Cornelis, Martine De Cock, Etienne E. Kerre |
Expert Syst. J. Knowl. Eng. | 2 |
| 2003 | On (un)suitable fuzzy relations to model approximate equality
Martine De Cock, Etienne E. Kerre |
Fuzzy Sets Syst. | 1 |
| 2003 | Why fuzzy -equivalence relations do not resolve the Poincaré paradox, and related issues
Martine De Cock, Etienne E. Kerre |
Fuzzy Sets Syst. | 1 |
| 2002 | Linguistic Hedges: a Quantifier Based Approach
Martine De Cock |
HIS | 1 |