Robin Brochier

dblp:219/7173 · DBLP profile ↗
← Back
3ranked-venue papers
3as first author
1since 2021 · last 2021
0000-0002-6188-6509ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Data mining · 35% Information retrieval · 28% Knowledge graphs · 23%
Artificial intelligence
1 paper
Representation and self-supervised learning · 100%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge graphs
link prediction
0.512021
Predicting Links on Wikipedia with Anchor Text Information · SIGIR 2021
Information retrieval
web graph
0.512021
Predicting Links on Wikipedia with Anchor Text Information · SIGIR 2021
Machine learning › Representation and self-supervised learning › word representation
word embedding
0.412019
Global Vectors for Node Representations · WWW 2019
Data mining › structured data mining › graph mining
graph learning
0.412019
Global Vectors for Node Representations · WWW 2019
Data mining › structured data mining › graph mining
network embedding
0.412019
Global Vectors for Node Representations · WWW 2019
Web and social media mining
web mining
0.112021
Predicting Links on Wikipedia with Anchor Text Information · SIGIR 2021
Web and social media mining › user-generated content
wikipedia
0.112021
Predicting Links on Wikipedia with Anchor Text Information · SIGIR 2021
Information retrieval › document processing › document analysis
document representation
0.112019
Global Vectors for Node Representations · WWW 2019

Methods — techniques the papers use, named apart from their topics

skip-gram with negative sampling · 0.8matrix factorization · 0.8glove · 0.8transductive learning · 0.5sampling methodology · 0.5inductive learning · 0.5
YearPublicationVenuePosition
2021 Predicting Links on Wikipedia with Anchor Text Information
abstract
Wikipedia, the largest open-collaborative online encyclopedia, is a corpus of documents bound together by internal hyperlinks. These links form the building blocks of a large network whose structure contains important information on the concepts covered in this encyclopedia. The presence of a link between two articles, materialised by an anchor text in the source page pointing to the target page, can increase readers' understanding of a topic. However, the process of linking follows specific editorial rules to avoid both under-linking and over-linking. In this paper, we study the transductive and the inductive tasks of link prediction on several subsets of the English Wikipedia and identify some key challenges behind automatic linking based on anchor text information. We propose an appropriate evaluation sampling methodology and compare several algorithms. Moreover, we propose baseline models that provide a good estimation of the overall difficulty of the tasks.
Robin Brochier, Frédéric Béchet
SIGIR1
2020 Inductive Document Network Embedding with Topic-Word Attention
Robin Brochier, Adrien Guille, Julien Velcin
ECIR (1)1
2019 Global Vectors for Node Representations
abstract
Most network embedding algorithms consist in measuring co-occur-rences of nodes via random walks then learning the embeddings using Skip-Gram with Negative Sampling. While it has proven to be a relevant choice, there are alternatives, such as GloVe, which has not been investigated yet for network embedding. Even though SGNS better handles non co-occurrence than GloVe, it has a worse time-complexity. In this paper, we propose a matrix factorization approach for network embedding, inspired by GloVe, that better handles non co-occurrence with a competitive time-complexity. We also show how to extend this model to deal with networks where nodes are documents, by simultaneously learning word, node and document representations. Quantitative evaluations show that our model achieves state-of-the-art performance, while not being so sensitive to the choice of hyper-parameters. Qualitatively speaking, we show how our model helps exploring a network of documents by generating complementary network-oriented and content-oriented keywords.
Robin Brochier, Adrien Guille, Julien Velcin
WWW1