Guillaume Cleuziou

dblp:61/545 · DBLP profile ↗
← Back
22ranked-venue papers
13as first author
3since 2021 · last 2024
0000-0002-2885-1152ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 11 first-author · 2 since 2021Databases, data management, data science and information retrieval · 9 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Theory of computation · 2 · 2 first-authorSoftware engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Information retrieval · 64% Data mining · 32% Data integration and cleaning · 5%
Artificial intelligence
1 paper
Knowledge representation and reasoning · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
retrieval models and ranking
0.212014
Query log driven web search results clustering · SIGIR 2014
Information retrieval › search engines
search result clustering
0.212014
Query log driven web search results clustering · SIGIR 2014
Knowledge, reasoning and agents › Knowledge representation and reasoning
ontology
0.112021
Multi-instance learning of pretopological spaces to model complex propagation phenomena: Application to lexical taxonomy learning · Artif. Intell. 2021
Data mining
clustering
0.112009
CoFKM: A Centralized Method for Multiple-View Clustering · ICDM 2009
Data mining › clustering
multi-view clustering
0.112009
CoFKM: A Centralized Method for Multiple-View Clustering · ICDM 2009
Data integration and cleaning
data fusion
0.012009
CoFKM: A Centralized Method for Multiple-View Clustering · ICDM 2009

Methods — techniques the papers use, named apart from their topics

pretopological spaces · 0.5multi-instance learning · 0.5dual c-means · 0.2clustering · 0.2fuzzy clustering · 0.1centralized fusion · 0.1
YearPublicationVenuePosition
2024 From Document to Program Embeddings: Can Distributional Hypothesis Really Be Used on Programming Languages?
abstract
Programming language processing is a field of increasing interest, as more and more models become available, either to address specific tasks or to acquire general knowledge which can then be fine-tuned on downstream tasks. All these models are based on architectures that come from the field of natural language processing, most of them being built on the distributional hypothesis from linguistics. Although this transition from one field to another appears to have occurred naturally, it is not so obvious to claim that this hypothesis will be appropriate for extracting semantics from programs. In this paper, we investigate to which extent a distributional hypothesis can be applied to code embedding. To this end, we first formulate various hypotheses adapted to the specific information contained in programming languages. We then provide a framework to evaluate the effectiveness of these hypotheses through the quality of the resulting embedding spaces. This framework is based on the doc2vec model as a generic language model, as its implementation of the original distributional hypothesis is easy to understand and to adapt to any new ones. Among other tools, we propose a new evaluation method based on program analogies, which measures how well the models capture the underlying structure and meaning of the code. We apply the proposed framework to a set of (distributional) hypotheses and show that we can rule out certain hypotheses in favor of others. Specifically, our study indicates that instruction-based hypotheses capture less semantic information than token-based ones. Furthermore, we observe that distributional hypotheses on tokens are effective in both source code, execution traces, and abstract syntax trees. Additionally, we find that the semantics captured on programs by these three hypotheses are of comparable levels and natures.
Thibaut Martinet, Guillaume Cleuziou, Matthieu Exbrayat, Frédéric Flouvat
ECAI2
2021 Learning student program embeddings using abstract execution traces
Guillaume Cleuziou, Frédéric Flouvat
EDM1
2021 Multi-instance learning of pretopological spaces to model complex propagation phenomena: Application to lexical taxonomy learning
Gaëtan Caillaut, Guillaume Cleuziou
Artif. Intell.2
2019 Learning Pretopological Spaces to Extract Ego-Centered Communities
Gaëtan Caillaut, Guillaume Cleuziou, Nicolas Dugué
PAKDD (2)2
2015 Learning Pretopological Spaces for Lexical Taxonomy Acquisition
Guillaume Cleuziou, Gaël Dias
ECML/PKDD (2)1
2015 Kernel methods for point symmetry-based clustering
Guillaume Cleuziou, José G. Moreno 0001
Pattern Recognit.1
2014 Query log driven web search results clustering
abstract
Different important studies in Web search results clustering have recently shown increasing performances motivated by the use of external resources. Following this trend, we present a new algorithm called Dual C-Means, which provides a theoretical background for clustering in different representation spaces. Its originality relies on the fact that external resources can drive the clustering process as well as the labeling task in a single step. To validate our hypotheses, a series of experiments are conducted over different standard datasets and in particular over a new dataset built from the TREC Web Track 2012 to take into account query logs information. The comprehensive empirical evaluation of the proposed approach demonstrates its significant advantages over traditional clustering and labeling techniques.
José G. Moreno 0001, Gaël Dias, Guillaume Cleuziou
SIGIR3
2014 Generalization of c-means for identifying non-disjoint clusters with overlap regulation
Chiheb-Eddine Ben N'cir, Guillaume Cleuziou, Nadia Essoussi
Pattern Recognit. Lett.2
2013 Osom: A method for building overlapping topological maps
Guillaume Cleuziou
Pattern Recognit. Lett.1
2011 A pretopological framework for the automatic construction of lexical-semantic structures from texts
abstract
We present in this paper a new approach for the automatic generation of lexical structures from texts. This tedious task is based on the strong hypothesis that simple statistical observations on textual usages can provide pieces of semantics about the lexicon. Using such "naive" observations only, we propose a (pre)-topological framework to formalize and combine various hypothesis on textual data usages and then to derive a structure similar to usual lexical knowledge basis such as WordNet. In addition we also consider the evaluation problem for obtained lexical structures ; a multi-level evaluation strategy is proposed that measures the fitting between a given reference structure and automatically generated structures on different point of views : intrinsic/structural and application-based points of view. The evaluation strategy is then used to quantify the contribution of the new structuring approach with respect to the corresponding solution proposed by (Sanderson et al. 2000) on two case studies that differs on the domain and the size of the lexicon.
Guillaume Cleuziou, Davide Buscaldi, Vincent Levorato, Gaël Dias
CIKM1
2011 Informative Polythetic Hierarchical Ephemeral Clustering
abstract
Ephemeral clustering has been studied for more than a decade, although with low user acceptance. According to us, this situation is mainly due to (1) an excessive number of generated clusters, which makes browsing difficult and (2) low quality labeling, which introduces imprecision within the search process. In this paper, our motivation is twofold. First, we propose to reduce the number of clusters of Web page results, but keeping all different query meanings. For that purpose, we propose a new polythetic methodology based on an informative similarity measure, the InfoSimba, and a new hierarchical clustering algorithm, the HISGK-means. Second, a theoretical background is proposed to define meaningful cluster labels embedded in the definition of the HISGK-means algorithm, which may elect as best label, words outside the given cluster. To confirm our intuitions, we propose a new evaluation framework, which shows that we are able to extract most of the important query meanings but generating much less clusters than state-of-the-art systems.
Gaël Dias, Guillaume Cleuziou, David Machado
Web Intelligence2
2009 CoFKM: A Centralized Method for Multiple-View Clustering
abstract
This paper deals with clustering for multi-view data, i.e. objects described by several sets of variables or proximity matrices. Many important domains or applications such as information retrieval, biology, chemistry and marketing are concerned by this problematic. The aim of this data mining research field is to search for clustering patterns that perform a consensus between the patterns from different views. This requires to merge information from each view by performing a fusion process that identifies the agreement between the views and solves the conflicts. Various fusion strategies can be applied, occurring either before, after or during the clustering process. We draw our inspiration from the existing algorithms based on a centralized strategy. We propose a fuzzy clustering approach that generalizes the three fusion strategies and outperforms the main existing multi-view clustering algorithm both on synthetic and real datasets.
Guillaume Cleuziou, Matthieu Exbrayat, Lionel Martin, Jacques-Henri Sublemontier
ICDM1
2008 Fully Unsupervised Graph-Based Discovery of General-Specific Noun Relationships from Web Corpora Frequency Counts
Gaël Dias, Raycho Mukelov, Guillaume Cleuziou
CoNLL3
2008 Mapping General-Specific Noun Relationships to WordNet Hypernym/Hyponym Relations
Gaël Dias, Raycho Mukelov, Guillaume Cleuziou
EKAW3
2008 An extended version of the k-means method for overlapping clustering
abstract
This paper deals with overlapping clustering, a trade off between crisp and fuzzy clustering. It has been motivated by recent applications in various domains such as information retrieval or biology. We show that the problem of finding a suitable coverage of data by overlapping clusters is not a trivial task. We propose a new objective criterion and the associated algorithm OKM that generalizes the k-means algorithm. Experiments show that overlapping clustering is a good alternative and indicate that OKM outperforms other existing methods.
Guillaume Cleuziou
ICPR1
2007 On the Impact of Lexical and Linguistic Features in Genre- and Domain-Based Categorization
Guillaume Cleuziou, Céline Poudat
CICLing1
2006 Structuring Natural Language Data by Learning Rewriting Rules
Guillaume Cleuziou, Lionel Martin, Christel Vrain
ILP1
2006 A Proximity Measure and a Clustering Method for Concept Extraction in an Ontology Building Perspective
Guillaume Cleuziou, Sylvie Billot, Stanislas Lew, Lionel Martin, Christel Vrain
ISMIS1
2004 PoBOC: An Overlapping Clustering Algorithm, Application to Rule-Based Classification and Textual Data
Guillaume Cleuziou, Lionel Martin, Christel Vrain
ECAI1
2004 DDOC: Overlapping Clustering of Words for Document Classification
Guillaume Cleuziou, Lionel Martin, Viviane Clavier, Christel Vrain
SPIRE1
2003 Genre and Domain Processing in an Information Retrieval Perspective
Céline Poudat, Guillaume Cleuziou
ICWE2
2003 Disjunctive Learning with a Soft-Clustering Method
Guillaume Cleuziou, Lionel Martin, Christel Vrain
ILP1